OpenAI's GPT-Red jailbreaks its own models 84% of the time
Curated by the Inblix editorial team
OpenAI has built an AI that’s better at breaking its own chatbots than any human red team ever was. The model, called GPT-Red, was trained through self-play reinforcement learning to hunt down security flaws in GPT systems. It simulates prompt injections and other attacks where malicious instructions get buried in emails, websites, or files — the kind of thing that keeps security engineers up at night.
The numbers are stark. GPT-Red finds successful attacks in 84 percent of test scenarios. Human red teamers? Just 13 percent. That’s not a marginal improvement. That’s a sixfold gap. In one internal test, the model manipulated an AI-powered vending machine in OpenAI’s office, changed prices, and canceled other customers’ orders. The sort of prank that becomes a lot less funny when scaled to production systems handling real money or sensitive data.
Those results aren’t just sitting in a report somewhere. They feed directly into training. OpenAI says GPT-5.6 Sol now shows six times fewer failures on direct prompt injections than the best model from four months ago, all without degrading general performance. That’s the holy grail — better safety without the usual tradeoffs that make models dumber. But here’s the catch: about 3.8 percent of stronger prompt injections still get through. As one engineer I spoke with put it, that’s great until you’re an enterprise customer running millions of queries a day.
Scale that 3.8 percent across hundreds or thousands of attempts, and a meaningful number of exploits will succeed. Claude Opus 4.5 shows similar residual vulnerability. OpenAI is keeping GPT-Red internal for now, with a technical paper promised soon. What’s clear is that automated red teaming isn’t a research curiosity anymore — it’s becoming the default way to harden models before they ship. The question isn’t whether this approach works. It’s whether we’re comfortable with the fact that even the best defenses leave doors cracked open.
💡 Key Takeaways
- GPT-Red achieved an 84% attack success rate in testing, compared to just 13% for human red teamers — a gap wide enough to make manual security testing look obsolete.
- GPT-5.6 Sol shows a sixfold reduction in direct prompt injection failures versus models from four months ago, but roughly 3.8% of stronger attacks still succeed.
- OpenAI demonstrated the model's real-world potential by having it manipulate an internal AI vending machine, altering prices and canceling orders.
- Automated self-play red teaming is shifting from experimental to operational, with attack results now feeding directly into model training pipelines.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.