AI Pulse by Inblix

OpenAI arms defenders with GPT-5.6-Cyber as refusal rate drops from 98.5% to 5%

OpenAI Blog · Aug 10, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI arms defenders with GPT-5.6-Cyber as refusal rate drops from 98.5% to 5%

The clock is ticking louder than ever. OpenAI just gave a select group of defenders a serious power-up with GPT-5.6-Cyber, a model so specialized it barely knows how to say “no” to advanced security tasks. The headline number is stark: when asked to handle dicey requests like developing exploit chains or bypassing authentication, standard GPT-5.6 Sol complies a measly 1.5% of the time. The new cyber-focused model? It jumps to a 95% completion rate. That isn’t a minor tweak; it’s a fundamental rewiring of the model’s risk posture for a very specific audience.

This capability is gated behind a new two-tier program called Daybreak. If you’re doing incident response or malware analysis, you get “Daybreak Blue,” which essentially lifts the standard guardrails off GPT-5.6 Sol but keeps the core model intact. The real heavy artillery is in “Daybreak Red.” That tier unlocks GPT-5.6-Cyber, which is purpose-trained not just to stop refusing, but to actually get better at the dark arts of vulnerability discovery and exploit validation. OpenAI reports the model outperforms its predecessors on internal benchmarks like ExploitGym, which tests the ability to turn known vulnerabilities into working code execution exploits.

What’s genuinely new here is the focus on finding novel, or “zero-day,” vulnerabilities. OpenAI created an internal test where models get a fresh open-source repo and are asked to find the maximum-impact flaw and write it up. GPT-5.6-Cyber didn’t just find bugs—it better calibrated their severity and produced higher-quality technical documentation. This addresses a long-standing gripe from security researchers working with GPT-5.5-Cyber, which still refused 42.7% of requests. That older model often got in its own way, forcing red teams to waste time arguing with a safety filter rather than probing a target.

OpenAI’s framing is one of preemptive defense, arguing that fully autonomous AI attacks are coming and defenders need this edge first. That’s a compelling story, but it also puts immense trust in the approval process for Daybreak Red. The line between a “trusted defender” and a malicious actor is only as strong as the vetting. With a model that practically vaults over ethical guardrails, the margin for error in deciding who gets the keys is razor-thin.

💡 Key Takeaways

  1. GPT-5.6-Cyber reduces refusal rates on advanced dual-use tasks from 98.5% to just 5%, a near-complete reversal that makes it a practical offensive security tool.
  2. Daybreak Red access explicitly trains the model to improve on finding zero-day vulnerabilities and calibrating their severity, not just to stop blocking requests.
  3. The specialized model outperforms general-purpose GPT-5.6 Sol on turning known vulnerabilities into working code execution exploits, as measured by the ExploitGym benchmark.
  4. OpenAI is betting on a pre-emptive strike strategy, reasoning that arming vetted defenders now is critical before fully autonomous AI attacks proliferate.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles