AI Pulse by Inblix

Claude Opus 5 quadruples ARC-AGI-3 reasoning record, hitting 30.2%

The Decoder · Jul 26, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Claude Opus 5 quadruples ARC-AGI-3 reasoning record, hitting 30.2%

Anthropic’s Claude Opus 5 didn’t just edge past the competition on the ARC-AGI-3 benchmark—it obliterated the previous record. The model scored 30.2%, nearly four times higher than OpenAI’s GPT-5.6 Sol (Max), which managed 7.8%. For a benchmark designed to measure genuine reasoning in novel situations rather than memorized knowledge, that’s a staggering gap. The ARC Prize team says the lead comes down to genuinely stronger logical reasoning, which lets the model explore and plan more autonomously in unfamiliar environments.

During testing, Opus 5 did something researchers had never seen before. It started translating tasks into algebraic notation and independently formulating reflection equations. It also cracked five environments that had never been solved, four of them at or above human level. That puts it well ahead of Anthropic’s own Fable-class models, which topped out around 20%. Six of the 25 public demo environments are now solved. On older benchmarks, Opus 5 hit 90.4% on ARC-AGI-2 and 97.5% on ARC-AGI-1.

Not everyone is convinced the gains are as broad as they look. Independent testing on Witness, a private benchmark for interactive puzzle games, showed a more modest picture. Opus 5 scored 43.4 there—statistically tied with Kimi K3 and Fable 5, and a much smaller jump over its predecessor Opus 4.8 than what ARC-AGI-3 suggests. The model did identify a conventional puzzle’s hidden rules before taking action, but it trailed Opus 4.8 on a game with less familiar mechanics. Guanghan Ning, who created Witness, says that pattern fits training on genre-specific data, though the benchmark can’t identify what data Anthropic used.

Greg Kamradt, one of the ARC-AGI-3 researchers, pushed back gently. A game built on familiar mechanics doesn’t test adaptation to novelty, he noted, and one weak result shouldn’t overshadow the overall improvement. Ning later clarified that Opus 5 did generalize to Witness—just far less dramatically. He compared the situation to how coding benchmarks evolved. ARC-AGI-3 is now the major target for interactive reasoning, so it’ll attract the most training effort first. Covering more edge cases could eventually help models generalize to a wider range of abstract reasoning tasks, much like coding models moved from saturated benchmarks like HumanEval to today’s coding agents. The open question: how much of Opus 5’s leap is better thinking, and how much is better training for this specific test?

💡 Key Takeaways

  1. Claude Opus 5 scored 30.2% on ARC-AGI-3, nearly 4x the previous record, and solved five environments no model had ever cracked before.
  2. Researchers observed genuinely new behavior: the model translated tasks into algebraic notation and wrote reflection equations on its own.
  3. Independent testing on the Witness benchmark shows far narrower gains, suggesting some of Opus 5's improvement may be specific to ARC-AGI-3's puzzle format rather than pure reasoning.
  4. The ARC Prize team maintains the results reflect real reasoning gains, but the debate over training contamination versus genuine generalization is far from settled.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles