AI Pulse by Inblix

OpenAI, US CAISI find agent hijack bugs fixed in a day

OpenAI Blog · Jul 13, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI, US CAISI find agent hijack bugs fixed in a day

The line between traditional cybersecurity and AI-specific exploits just got blurrier — and that’s exactly the point of a deepening partnership between OpenAI and the US Center for AI Standards and Innovation (CAISI). The two organizations have been collaborating for over a year on model evaluations in national security domains, but their latest work moved into a more practical, and frankly scarier, arena: tearing apart the security of OpenAI’s agentic products before attackers could.

In July, CAISI received early access to ChatGPT Agent and got to work with a multidisciplinary team spanning cybersecurity and AI agent security. What they found wasn’t a simple bug. They discovered two novel vulnerabilities that, on their own, looked like dead ends due to OpenAI’s existing security measures. But the team then chained a traditional cyber vulnerability with an AI agent hijacking attack, creating a proof-of-concept exploit chain that worked around half the time. The attack could let a sophisticated adversary bypass protections, remotely control the computer systems the agent accessed during a session, and impersonate the user on other logged-in websites.

Here’s the detail that should make security professionals pause: CAISI actually used ChatGPT Agent itself to help discover these vulnerabilities. AI systems hunting for flaws in other AI systems isn’t a sci-fi concept anymore — it’s a working methodology. The vulnerabilities were reported to OpenAI immediately and fixed within one business day. That turnaround speed is notable, but the real story is the type of attack. It required combining classical software exploitation techniques with an understanding of how AI agents interact with the web, something most red teams aren’t structured to do.

OpenAI frames this as proof that voluntary government collaborations produce real-world security improvements, not just paperwork. The CAISI work complements the company’s existing partnerships with the UK AI Security Institute, including previous red-teaming against biological misuse risks. The push into agentic system testing is a preliminary step, both sides acknowledge, but the fact that it immediately surfaced a chainable exploit suggests we’re in the early innings of understanding how to secure autonomous AI. The question nobody’s answering yet: what happens when the agent you’re trying to secure is also the tool attackers use to find the next vulnerability?

💡 Key Takeaways

  1. CAISI chained a traditional cyber exploit with an AI agent hijack, achieving a roughly 50% success rate on a full exploit chain that OpenAI fixed within one business day.
  2. The CAISI team used ChatGPT Agent itself as a tool to discover the vulnerabilities, marking a concrete example of AI assisting in adversarial security testing of AI systems.
  3. The vulnerabilities were initially dismissed as unexploitable, proving that evaluating agentic AI requires multidisciplinary teams comfortable with both software exploitation and machine learning attack surfaces.
  4. This voluntary collaboration expands a year-old partnership beyond model evaluations in biosecurity and cyber domains into live product red-teaming for consumer-facing agent systems.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

← Back to all articles