OpenAI's Aardvark AI caught 92% of bugs in benchmark test
Curated by the Inblix editorial team
OpenAI is taking its GPT-5 model and pointing it squarely at a problem that costs the software industry billions: the 40,000-plus CVEs reported every year. The result is Aardvark, an agentic security researcher that doesn’t just scan code—it thinks like a human pentester. Unlike traditional tools that lean on fuzzing or software composition analysis, Aardvark reads repositories, builds threat models, and then validates vulnerabilities by actually trying to trigger them inside a sandbox.
The early numbers are eye-opening. In benchmark testing on so-called “golden” repositories, Aardvark identified 92% of known and synthetically-introduced vulnerabilities. That’s a high-recall figure that suggests the agent isn’t just finding trivial bugs; it’s catching issues that occur under complex conditions. OpenAI has been dogfooding the system internally for months, where it surfaced meaningful vulnerabilities and shored up the company’s own defensive posture.
But the real flex might be what’s happened in the wild. Applied to open-source projects, Aardvark has already discovered vulnerabilities that earned ten separate CVE identifiers. The company is now offering pro-bono scanning to select non-commercial repositories, a move that positions this as a public-good play for the software supply chain. The agent also hunts for logic flaws, incomplete fixes, and privacy issues—scope creep that most security teams would welcome with open arms.
Workflow integration is the quiet killer feature here. Aardvark monitors commits, annotates code for human review, and attaches Codex-generated patches to each finding for one-click fixes. The agent is rolling out now to ChatGPT Enterprise, Business, and Edu customers under the new Codex Security brand, with free usage for the next month. There’s a clear strategy at play: make the tool sticky by embedding it directly into the developer workflow, then let the results speak for themselves.
💡 Key Takeaways
- Aardvark caught 92% of vulnerabilities in benchmark testing, signaling LLM-based reasoning is now competitive with traditional program analysis.
- The agent has already earned ten CVE identifiers for vulnerabilities it found in open-source projects—real-world impact, not just lab results.
- OpenAI is rebranding Aardvark as Codex Security and offering free access to enterprise customers for a month to drive rapid adoption.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.