AI Pulse by Inblix

Claude Opus 5 just scored 4x higher than GPT-5.6 Sol on novel problem-solving

The Decoder · Jul 24, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Claude Opus 5 just scored 4x higher than GPT-5.6 Sol on novel problem-solving

Anthropic dropped a new flagship model that reshapes the price-performance conversation. Claude Opus 5 doesn’t just inch past the competition on a few benchmarks—it posts a genuinely weird result on ARC-AGI-3, a test explicitly designed to measure reasoning without memorized patterns. Opus 5 scores 30.2 percent. GPT-5.6 Sol manages 7.8 percent. That’s not a gap, that’s a different weight class.

The model achieves this while costing the same as its predecessor: $5 per million input tokens and $25 per million output tokens. That’s half the price of Anthropic’s own Claude Fable 5, which Opus 5 beats in agentic coding (43.3% vs 33.7% on Frontier-Bench) and knowledge work. It’s a direct shot at pricing pressure from OpenAI and Chinese labs, making Opus 5 the default on Claude Max and the best model available on Pro plans.

But token rates don’t tell the whole story. Anthropic’s own prompting guide warns that higher effort settings don’t always mean better results. On two benchmarks, Opus 5 actually performs slightly worse at max effort than at the second-highest setting, while burning more tokens. It’s a repeat of a pattern we saw with Opus 4.7, where per-task costs ended up 30 to 40 percent higher despite identical base rates. The company is nudging users toward “low” and “medium” settings for most work.

The most revealing detail isn’t a benchmark number. It’s a specific test where Opus 5 had to recreate a 3D model from a drawing it couldn’t directly view. The model wrote its own computer vision pipeline from scratch to extract geometry from raw pixels, then reconstructed the part. No other model solved it in five attempts. An engineer at a trading firm reportedly shipped a market data feed in one session—something previous models couldn’t do with detailed plans. Benchmarks are useful, but that kind of adaptive tool-building hints at something benchmarks miss.

💡 Key Takeaways

  1. Opus 5's 30.2% ARC-AGI-3 score is nearly 4x GPT-5.6 Sol's 7.8%, suggesting a leap in novel reasoning that may not show up in standard coding benchmarks.
  2. Higher effort settings don't reliably deliver better results—Opus 5 performs worse at max effort on two benchmarks while costing more, echoing the token-efficiency trap from Opus 4.7.
  3. The model built its own computer vision pipeline to solve a 3D modeling task after failing to access the drawing directly, a behavior that separates tool-using models from task-completing ones.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles