AI Pulse by Inblix

OpenAI ran its safety tests on Claude. Here's what it found.

OpenAI Blog · Jul 13, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI ran its safety tests on Claude. Here's what it found.

In an unusual move, OpenAI and Anthropic traded their latest public models this summer to let the other lab run its own internal safety evaluations on them. The results, published Tuesday, offer a rare, side-by-side view of how Claude Opus 4, Claude Sonnet 4, GPT‑4o, GPT‑4.1, o3, and o4-mini handle jailbreaking, hallucination, and instruction hierarchy tests—without the usual apples-to-apples framing. OpenAI was careful not to declare winners, noting that differences in API access and deep familiarity with one’s own models make direct comparisons tricky. “It is therefore not appropriate to draw sweeping claims from these results,” the company wrote.

The most stark divergence showed up in hallucination tests. Claude 4 models refused to answer up to 70% of the time, a sign they know what they don’t know. But when they did answer, the accuracy was still low. OpenAI’s o3 and o4-mini, by contrast, refused far less often—and hallucinated more. That’s a design tradeoff, not a bug. On jailbreaking, Claude models generally lagged behind o3 and o4-mini, though one detail caught the evaluators’ attention: disabling reasoning actually made Claude more resistant to certain jailbreaks. OpenAI called that out as a counterintuitive finding worth digging into.

Claude held its own on instruction hierarchy tests, which stress-test whether a model follows developer-set system messages over conflicting user prompts. The Claude 4 models slightly edged out o3 in avoiding system message vs. user message conflicts and performed as well or better on resisting system-prompt extraction attempts. Both labs relaxed some model-external safeguards—standard practice for dangerous-capability testing—so these aren’t real-world behavior estimates. They’re probes of what models might attempt under adversarial pressure.

OpenAI framed the collaboration as a transparency exercise, not a final exam. “Safety testing is never finished,” the post reads, and the company says GPT‑5, launched after this work, already shows improvements on sycophancy, hallucination, and misuse resistance. The bigger story may be the process itself: two rival labs opening their models to each other’s red-teaming infrastructure and publishing the results in parallel. It’s the kind of cooperative pressure-testing that safety researchers have been asking for, and it’s happening between organizations that are otherwise locked in a fierce commercial race.

💡 Key Takeaways

  1. Claude 4 models refused to answer up to 70% of hallucination-triggering prompts, a conservative posture that avoids errors but limits usefulness.
  2. Disabling reasoning unexpectedly improved Claude's resistance to some jailbreaks, a counterintuitive result that warrants further investigation.
  3. OpenAI and Anthropic relaxed model-external safeguards for these tests, meaning the results probe model propensities under adversarial conditions—not real-world behavior.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles