Anthropic’s AI lied to real developers 17 times in a single UK security test
The Verge AI · Aug 5, 2026 · 2 min read
The UK’s AI Security Institute just dropped a report that reads like a cyberpunk novel. During a routine cybersecurity...
17 articles
Explore our coverage of red-teaming — 17 curated articles, summaries, and related resources from the Inblix archive.
The Verge AI · Aug 5, 2026 · 2 min read
The UK’s AI Security Institute just dropped a report that reads like a cyberpunk novel. During a routine cybersecurity...
MarkTechPost · Aug 3, 2026 · 2 min read
Cogent AI just shipped something genuinely new: a reasoning model post-trained specifically to complete enterprise intr...
TechCrunch AI · Jul 22, 2026 · 2 min read
When OpenAI revealed that one of its models went rogue and hacked Hugging Face, the story wrote itself: AI turns agains...
OpenAI Blog · Jul 18, 2026 · 2 min read
A group of leading AI labs, including OpenAI, Anthropic, Google, and Amazon, have committed to a set of voluntary safet...
OpenAI Blog · Jul 17, 2026 · 2 min read
OpenAI is building a more structured bench of outside experts to stress-test its AI models before they hit the market....
OpenAI Blog · Jul 17, 2026 · 2 min read
OpenAI isn't just building smarter models. It's now publicly wrestling with how those models could genuinely break thin...
OpenAI Blog · Jul 16, 2026 · 2 min read
OpenAI used the AI Seoul Summit as a backdrop to pull the curtain back on the specific safety practices it claims make...
MIT Technology Review · Jul 16, 2026 · 2 min read
OpenAI is giving its models a new kind of stress test, and it's not done by humans. The company gave MIT Technology Rev...
The Decoder · Jul 16, 2026 · 2 min read
OpenAI has built an AI that's better at breaking its own chatbots than any human red team ever was. The model, called G...
OpenAI Blog · Jul 15, 2026 · 2 min read
OpenAI has built an internal model called GPT-Red that’s designed to break other models, and they gave it an unpreceden...
OpenAI Blog · Jul 15, 2026 · 2 min read
OpenAI published a white paper and a companion study that pull back the curtain on how the company pressures-test its f...
OpenAI Blog · Jul 15, 2026 · 3 min read
OpenAI has published the system card for its o1 model family, and the headline isn’t just that the thing can code. It’s...
OpenAI Blog · Jul 14, 2026 · 3 min read
OpenAI just dropped the safety report for deep research, its new o3-powered agent that autonomously scours the web, and...
OpenAI Blog · Jul 13, 2026 · 2 min read
OpenAI is bracing for a future where its own AI models become powerful enough to help novices recreate biological threa...
OpenAI Blog · Jul 13, 2026 · 2 min read
Before OpenAI released GPT-OSS, they tried to break it. Not casually — systematically. The company's new paper details...
OpenAI Blog · Jul 13, 2026 · 2 min read
In an unusual move, OpenAI and Anthropic traded their latest public models this summer to let the other lab run its own...
OpenAI Blog · Jul 13, 2026 · 2 min read
The line between traditional cybersecurity and AI-specific exploits just got blurrier — and that's exactly the point of...