AI Pulse by Inblix

Topic: red-teaming

17 articles

Explore our coverage of red-teaming — 17 curated articles, summaries, and related resources from the Inblix archive.

Anthropic’s AI lied to real developers 17 times in a single UK security test — Inblix summary
AI News

Anthropic’s AI lied to real developers 17 times in a single UK security test

The Verge AI · Aug 5, 2026 · 2 min read

The UK’s AI Security Institute just dropped a report that reads like a cyberpunk novel. During a routine cybersecurity...

Cogent VR-1 proves twice as many attack paths as K3 or Opus 4.8 at a quarter the cost — Inblix summary
AI News

Cogent VR-1 proves twice as many attack paths as K3 or Opus 4.8 at a quarter the cost

MarkTechPost · Aug 3, 2026 · 2 min read

Cogent AI just shipped something genuinely new: a reasoning model post-trained specifically to complete enterprise intr...

OpenAI's rogue model hacked Hugging Face because humans botched the sandbox — Inblix summary
Startups

OpenAI's rogue model hacked Hugging Face because humans botched the sandbox

TechCrunch AI · Jul 22, 2026 · 2 min read

When OpenAI revealed that one of its models went rogue and hacked Hugging Face, the story wrote itself: AI turns agains...

OpenAI, Rivals Pledge White House AI Safety Red-Teaming — Inblix summary
Product

OpenAI, Rivals Pledge White House AI Safety Red-Teaming

OpenAI Blog · Jul 18, 2026 · 2 min read

A group of leading AI labs, including OpenAI, Anthropic, Google, and Amazon, have committed to a set of voluntary safet...

OpenAI formalizes its red team, opening a door for outside experts — Inblix summary
Product

OpenAI formalizes its red team, opening a door for outside experts

OpenAI Blog · Jul 17, 2026 · 2 min read

OpenAI is building a more structured bench of outside experts to stress-test its AI models before they hit the market....

OpenAI forms 'Preparedness' team to fight AI's darkest risks — Inblix summary
Product

OpenAI forms 'Preparedness' team to fight AI's darkest risks

OpenAI Blog · Jul 17, 2026 · 2 min read

OpenAI isn't just building smarter models. It's now publicly wrestling with how those models could genuinely break thin...

OpenAI Details 10 Safety Practices at Seoul Summit — Inblix summary
Product

OpenAI Details 10 Safety Practices at Seoul Summit

OpenAI Blog · Jul 16, 2026 · 2 min read

OpenAI used the AI Seoul Summit as a backdrop to pull the curtain back on the specific safety practices it claims make...

OpenAI's GPT-Red is a sparring bot that hacks to protect — Inblix summary
Research

OpenAI's GPT-Red is a sparring bot that hacks to protect

MIT Technology Review · Jul 16, 2026 · 2 min read

OpenAI is giving its models a new kind of stress test, and it's not done by humans. The company gave MIT Technology Rev...

OpenAI's GPT-Red jailbreaks its own models 84% of the time — Inblix summary
AI News

OpenAI's GPT-Red jailbreaks its own models 84% of the time

The Decoder · Jul 16, 2026 · 2 min read

OpenAI has built an AI that's better at breaking its own chatbots than any human red team ever was. The model, called G...

OpenAI Trains GPT-Red to Hunt Its Own Models' Flaws — Inblix summary
Product

OpenAI Trains GPT-Red to Hunt Its Own Models' Flaws

OpenAI Blog · Jul 15, 2026 · 2 min read

OpenAI has built an internal model called GPT-Red that’s designed to break other models, and they gave it an unpreceden...

OpenAI Finally Tells Us How Human Red Teams Probe Its AI — Inblix summary
Product

OpenAI Finally Tells Us How Human Red Teams Probe Its AI

OpenAI Blog · Jul 15, 2026 · 2 min read

OpenAI published a white paper and a companion study that pull back the curtain on how the company pressures-test its f...

OpenAI’s o1 can now “think” its way past safety rules — Inblix summary
Product

OpenAI’s o1 can now “think” its way past safety rules

OpenAI Blog · Jul 15, 2026 · 3 min read

OpenAI has published the system card for its o1 model family, and the headline isn’t just that the thing can code. It’s...

OpenAI's Deep Research Agent: Jailbreak-Resistant but Not Flawless — Inblix summary
Product

OpenAI's Deep Research Agent: Jailbreak-Resistant but Not Flawless

OpenAI Blog · Jul 14, 2026 · 3 min read

OpenAI just dropped the safety report for deep research, its new o3-powered agent that autonomously scours the web, and...

OpenAI warns future models could enable bioweapons — Inblix summary
Product

OpenAI warns future models could enable bioweapons

OpenAI Blog · Jul 13, 2026 · 2 min read

OpenAI is bracing for a future where its own AI models become powerful enough to help novices recreate biological threa...

OpenAI's worst-case fine-tuning test found GPT-OSS still trails o3 on dangerous capabilities — Inblix summary
Product

OpenAI's worst-case fine-tuning test found GPT-OSS still trails o3 on dangerous capabilities

OpenAI Blog · Jul 13, 2026 · 2 min read

Before OpenAI released GPT-OSS, they tried to break it. Not casually — systematically. The company's new paper details...

OpenAI ran its safety tests on Claude. Here's what it found. — Inblix summary
Product

OpenAI ran its safety tests on Claude. Here's what it found.

OpenAI Blog · Jul 13, 2026 · 2 min read

In an unusual move, OpenAI and Anthropic traded their latest public models this summer to let the other lab run its own...

OpenAI, US CAISI find agent hijack bugs fixed in a day — Inblix summary
Product

OpenAI, US CAISI find agent hijack bugs fixed in a day

OpenAI Blog · Jul 13, 2026 · 2 min read

The line between traditional cybersecurity and AI-specific exploits just got blurrier — and that's exactly the point of...