AI Pulse by Inblix

Sakana’s Fugu-Cyber nudges past GPT-5.5 on one security benchmark—but it’s still gated

MarkTechPost · Jul 26, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Sakana’s Fugu-Cyber nudges past GPT-5.5 on one security benchmark—but it’s still gated

Sakana AI dropped Fugu-Cyber on Monday, and no, it’s not a new frontier model. It’s a cybersecurity-tuned endpoint that plugs into the company’s existing Fugu orchestrator—the same one that launched a month ago and already had researchers arguing about agentic scaffolds versus raw model capability.

The orchestrator is the interesting bit. Fugu reads a query and builds a multi-agent workflow on the fly, delegating subtasks to specialist models. For security work, that means one agent might surface a vulnerability while another—trained specifically for verification—double-checks the finding before anything gets patched. Sakana’s team points to the verifier role as the secret sauce, though routing stays proprietary so you can’t see which model handled what.

The numbers: 86.9% on UC Berkeley’s CyberGym benchmark and 72.1% on Microsoft’s CTI-REALM. On CyberGym—which tests whether an agent can write a proof-of-concept that crashes unpatched code but not the fixed version—that’s a hair past OpenAI’s GPT-5.5-Cyber at 85.6% and Anthropic’s Claude Mythos Preview at 83.1%. Context matters. When CyberGym first launched, the best any agent-model pair could manage was around 20%. The field has moved fast.

CTI-REALM tells a different story. Microsoft’s own evaluation put the top three configurations—all Claude—in a 0.624 to 0.685 band on what’s actually a trajectory reward score, not a pass/fail rate. Sakana calls 72.1% a “success rate” anyway, which is worth flagging. The benchmark asks agents to map threat reports to MITRE ATT&CK techniques and emit validated Sigma rules across Linux, AKS, and Azure cloud environments.

Don’t expect to kick the tires without paperwork. Access requires an application form, manual review, and a commitment to defensive use only. The API isn’t available in the EU or EEA yet—GDPR compliance work is ongoing—and pricing runs $6 per million input tokens, double that above 272K context. Long codebase runs will cross that threshold routinely.

💡 Key Takeaways

  1. Fugu-Cyber’s 86.9% on CyberGym barely edges GPT-5.5-Cyber’s 85.6%—it’s iteration, not a breakthrough
  2. The orchestrator’s verifier role is where Sakana claims the real value, but proprietary routing means you can’t audit which model made which call
  3. CTI-REALM’s 72.1% isn’t a pass rate; it’s a trajectory reward score Sakana is mislabeling, which should raise eyebrows among detection engineers
  4. Gated access, no EU availability, and a 20% price premium over Fugu-Ultra signal this is for serious security teams with budget and patience

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles