AI Pulse by Inblix

Claude Opus 5 beats Fable 5 on benchmarks but hallucinates 50% of the time

The Decoder · Jul 25, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Claude Opus 5 beats Fable 5 on benchmarks but hallucinates 50% of the time

Anthropic’s Claude Opus 5 just claimed the top spot on the Artificial Analysis Intelligence Index with a score of 61, nudging past Fable 5 and GPT-5.6 Sol. It’s a genuine leap in analytical quality and knowledge work — the AA-Briefcase benchmark for office tasks puts Opus 5 at an Elo of 1720, a 146-point gap over its closest rival. In coding, the model ties for first place when paired with Claude Code. Not bad for a model that costs $2.03 per average benchmark task, about 26% less than Fable 5 with fallback.

But here’s where the story gets complicated. That 50% hallucination rate is hard to ignore. The jump from Opus 4.8’s 36% came because Opus 5 tends to answer confidently even when it’s unsure, which is exactly the behavior you don’t want in high-stakes applications. On factual accuracy benchmarks like AA-Omniscience, it still can’t touch Fable 5 despite a 7-point improvement. Anthropic clearly prioritized capability over caution this time around.

There’s a pricing lesson buried in the performance data too. Vals.ai found that Opus 5 peaks at the “high” reasoning tier — 89.8% on Vibe Code Bench — before actually getting worse at “xhigh” and “max.” The higher tiers produce needlessly complex solutions that introduce more errors. Terminal-Bench 2.1 tells the same story: the “max” tier wastes so much time per attempt that it gets fewer shots within the time limit. Anthropic seems to agree, setting “high” as the default across the API and Claude Code.

Epoch AI’s numbers confirm what the broader pattern suggests: no single model has broken away. Opus 5 scores 159 on the Epoch Capability Index, barely behind Fable 5 at 161. The frontier is a knife fight now. That’s great for pricing pressure and terrible for anyone betting on a single provider’s dominance. Token pricing remains at $5 per million input and $25 per million output, with cache hits at a dirt-cheap $0.50. If you’re doing knowledge work and can tolerate some factual sloppiness, Opus 5 is the best deal going right now.

💡 Key Takeaways

  1. Opus 5's 50% hallucination rate — up from 36% in the previous model — makes it risky for any task where factual accuracy is non-negotiable, despite its benchmark dominance
  2. Paying for the highest reasoning tier is often a waste: the "high" tier delivers better results than "max" on both coding and terminal benchmarks while costing significantly less
  3. The frontier AI race is effectively tied, with Opus 5, Fable 5, and GPT-5.6 Sol separated by razor-thin margins that will accelerate commoditization

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles