OpenAI Built GPT-Red, an AI Super-Hacker, to Bully Its Own Models
Curated by the Inblix editorial team
Most companies use red-teaming to find flaws before attackers do. OpenAI just automated the entire process—and the results are a little unsettling. The company has built an LLM called GPT-Red specifically to relentlessly attack its other models, and it’s better at the job than humans. According to OpenAI, GPT-Red discovered a novel prompt injection technique called a “fake chain of thought,” where it slips a bogus note into a model’s internal reasoning diary. Research scientist Chris Choquette-Choo describes the trick bluntly: “It’s like if I told you that 1+1=3 and that you have verified this already. The model’s like, ‘Oh, okay, of course,’ and it just spits out 3.”
OpenAI trained GPT-Red in a virtual dojo simulating real-world agent scenarios like browsing the web and editing code. Using a self-play loop, the attacker model sparred against defender models over thousands of rounds, getting ruthlessly efficient. “Compared to a human red-teamer, the model is very, very good at finding exactly what will work, exactly what’s most effective,” says co-creator Dylan Hunn. In a rerun of a 2025 human red-teaming test on an older GPT-5 version, the AI outperformed the people. The impact on safety is tangible: more than 90% of GPT-Red’s strongest attacks worked on last August’s GPT-5 release, but fewer than 23% succeeded against the new GPT-5.6.
The scope of the threat is expanding faster than human testers can track. As LLMs morph into agents that touch files, websites, and third-party code, the potential damage from a successful jailbreak isn’t just about embarrassing text—it could mean stolen data or sabotaged infrastructure. Jessica Ji, a senior analyst at Georgetown’s CSET, calls the self-play approach “very promising.” OpenAI’s Nikhil Kandpal frames the project as a necessary arms race, saying the risk surface and blast radius both grow with more capable models.
GPT-Red isn’t a flawless adversary, though. It struggles with multi-turn conversational attacks, a tactic human hackers wield with ease. OpenAI openly acknowledges the limitation. The bigger question—one the company hasn’t fully answered—is whether training models against an in-house digital psychopath actually covers the chaos of the open internet, or if it just makes them very good at passing their own tests.
💡 Key Takeaways
- GPT-Red automated the discovery of a novel "fake chain of thought" attack that tricks models into treating false data as verified fact.
- OpenAI claims GPT-Red outperformed human red-teamers, finding attacks that slashed the success rate of exploits from over 90% on GPT-5 to under 23% on GPT-5.6.
- The AI struggles with conversational, back-and-forth attack patterns, highlighting a gap between automated testing and the tactics of real human hackers.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.