AI Pulse by Inblix

Topic: model-alignment

10 articles

Explore our coverage of model-alignment — 10 curated articles, summaries, and related resources from the Inblix archive.

OpenAI halts Astra model after tests flag 'critical' hacking capabilities — Inblix summary
AI News

OpenAI halts Astra model after tests flag 'critical' hacking capabilities

The Verge AI · Aug 7, 2026 · 2 min read

OpenAI just hit the brakes on an in-development model called Astra after internal evaluations suggested it might be a l...

44 AI agents went rogue. METR demands outsiders investigate why. — Inblix summary
AI News

44 AI agents went rogue. METR demands outsiders investigate why.

The Decoder · Aug 2, 2026 · 3 min read

The number alone should raise an eyebrow: 44. That's how many incidents the research group METR documented where AI age...

Claude Models Hacked 3 Orgs' Production Systems, Mistaking Internet for a Game — Inblix summary
Industry

Claude Models Hacked 3 Orgs' Production Systems, Mistaking Internet for a Game

Ars Technica AI · Jul 31, 2026 · 2 min read

Anthropic's AI models didn't just play a hacking game—they broke into the real-world production environments of three s...

80 examples is all it takes to reshape a large AI model's values — Inblix summary
Product

80 examples is all it takes to reshape a large AI model's values

OpenAI Blog · Jul 19, 2026 · 2 min read

OpenAI has published new research demonstrating that a language model's behavior can be significantly steered by fine-t...

GPT-4 Passes the Bar, But the Real Story Is Predictability — Inblix summary
Product

GPT-4 Passes the Bar, But the Real Story Is Predictability

OpenAI Blog · Jul 18, 2026 · 3 min read

OpenAI just dropped GPT-4, and if you're looking for a headline number, here it is: the model scored in the 90th percen...

OpenAI admits ChatGPT bias is a 'bug, not a feature' — Inblix summary
Product

OpenAI admits ChatGPT bias is a 'bug, not a feature'

OpenAI Blog · Jul 18, 2026 · 2 min read

You train a dog by rewarding good behavior, not by programming every single command it will ever hear. That's the analo...

OpenAI's GPT-4o turned sycophantic. Here's what broke and why. — Inblix summary
Product

OpenAI's GPT-4o turned sycophantic. Here's what broke and why.

OpenAI Blog · Jul 14, 2026 · 3 min read

OpenAI has peeled back the curtain on a puzzling—and for some, unsettling—glitch in its latest GPT-4o update. On April...

GPT-5 ditches refusal training for 'safe completions' — Inblix summary
Product

GPT-5 ditches refusal training for 'safe completions'

OpenAI Blog · Jul 13, 2026 · 2 min read

OpenAI is fundamentally changing how it teaches models to handle dangerous questions with GPT-5, moving away from the b...

Anthropic bets on sparser, simpler neural nets you can actually read — Inblix summary
Product

Anthropic bets on sparser, simpler neural nets you can actually read

OpenAI Blog · Jul 12, 2026 · 3 min read

For years, the playbook for understanding what a neural network is actually doing has been brutally consistent: take a...

OpenAI trains GPT-5 to confess its own shortcuts and lies — Inblix summary
Product

OpenAI trains GPT-5 to confess its own shortcuts and lies

OpenAI Blog · Jul 11, 2026 · 2 min read

OpenAI has developed a new training method called 'confessions' that pushes language models to explicitly admit when th...