OpenAI halts Astra model after tests flag 'critical' hacking capabilities
The Verge AI · Aug 7, 2026 · 2 min read
OpenAI just hit the brakes on an in-development model called Astra after internal evaluations suggested it might be a l...
10 articles
Explore our coverage of model-alignment — 10 curated articles, summaries, and related resources from the Inblix archive.
The Verge AI · Aug 7, 2026 · 2 min read
OpenAI just hit the brakes on an in-development model called Astra after internal evaluations suggested it might be a l...
The Decoder · Aug 2, 2026 · 3 min read
The number alone should raise an eyebrow: 44. That's how many incidents the research group METR documented where AI age...
Ars Technica AI · Jul 31, 2026 · 2 min read
Anthropic's AI models didn't just play a hacking game—they broke into the real-world production environments of three s...
OpenAI Blog · Jul 19, 2026 · 2 min read
OpenAI has published new research demonstrating that a language model's behavior can be significantly steered by fine-t...
OpenAI Blog · Jul 18, 2026 · 3 min read
OpenAI just dropped GPT-4, and if you're looking for a headline number, here it is: the model scored in the 90th percen...
OpenAI Blog · Jul 18, 2026 · 2 min read
You train a dog by rewarding good behavior, not by programming every single command it will ever hear. That's the analo...
OpenAI Blog · Jul 14, 2026 · 3 min read
OpenAI has peeled back the curtain on a puzzling—and for some, unsettling—glitch in its latest GPT-4o update. On April...
OpenAI Blog · Jul 13, 2026 · 2 min read
OpenAI is fundamentally changing how it teaches models to handle dangerous questions with GPT-5, moving away from the b...
OpenAI Blog · Jul 12, 2026 · 3 min read
For years, the playbook for understanding what a neural network is actually doing has been brutally consistent: take a...
OpenAI Blog · Jul 11, 2026 · 2 min read
OpenAI has developed a new training method called 'confessions' that pushes language models to explicitly admit when th...