Sakana's Fugu router now beats a model it can't even access, claiming 7.9-point gain
Curated by the Inblix editorial team
Sakana AI just dropped an update to Fugu Ultra, its AI model router, and the claim is genuinely weird: version 1.1 can now beat Fable 5, a model that isn’t even in the router’s selection pool. That’s a neat trick if true—like a sommelier recommending a wine better than anything on the menu. The Tokyo-based startup reports performance jumps of up to 7.9 points over v1.0, with the biggest gains on coding benchmarks ProgramBench and TerminalBench 2.1. All numbers are self-reported, so take them with the appropriate grain of salt until someone independent kicks the tires.
Pricing hasn’t budged—still $5 per million input tokens and $30 per million output. That puts it squarely in premium territory for a routing layer, especially when the first version got dinged for high token usage and glacial speed. Sakana says it needs about two weeks of training and evaluation before a new top-tier model can join the pool, which suggests a fairly heavyweight integration process rather than a quick API swap. The update also ships with a Claude Code-compatible endpoint, letting developers call Fugu directly from the terminal—a nod to the growing developer preference for command-line workflows.
The original Fugu launch landed with a thud. Critics hammered it for slow response times and mediocre output quality. Sakana seems to be addressing the quality side of that equation with v1.1, but the latency question remains open. The company still won’t serve the EU or EEA, citing GDPR compliance hurdles. That’s a non-trivial chunk of the global developer market they’re leaving on the table, and it signals that a small AI lab can still struggle with the regulatory overhead that bigger players have long since absorbed.
What’s actually happening under the hood? The technical report describes how the router distributes queries across publicly available models, but the mechanism that lets it outperform a model it can’t route to is the genuinely interesting bit. It implies the router isn’t just picking the best model for a prompt—it’s synthesizing something from its available pool that collectively surpasses what any individual model, even an external one, can do. If that holds up, it’s a real argument for routing architectures over single-model deployments. If it doesn’t, it’s just another benchmark flex with an asterisk.
💡 Key Takeaways
- Fugu v1.1 claims to beat Fable 5 without including it in the router's pool, implying the collective output of its model ensemble outperforms the unavailable model.
- Performance gains of up to 7.9 points target coding benchmarks specifically, suggesting Sakana is optimizing for developer use cases over general chat.
- The continued EU service gap means a significant portion of the developer market can't even test these claims, limiting independent verification.
- Sakana's two-week integration cycle for new models makes Fugu inherently slower to adapt than the models it routes between.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.