GPT-5.6 Sol hits 38.3% on ARC-AGI-3, but only after OpenAI tweaks the rules
Curated by the Inblix editorial team
The ARC-AGI-3 leaderboard just got messy. OpenAI claims GPT-5.6 Sol scored 38.3 percent on the tough reasoning benchmark, leapfrogging Anthropic’s Claude Opus 5 at 30.2 percent. But there’s a catch — a big one. That number came from OpenAI’s own API with two special settings turned on, not the official test harness. In the standard environment, the exact same model limped to 7.8 percent.
The two settings doing the heavy lifting are “Retained Reasoning” and “Compaction.” The first preserves the model’s chain of thought between actions instead of wiping it clean each time. The second summarizes old context rather than chopping it off. OpenAI’s argument is straightforward: benchmarks measure the whole pipeline, not just a naked model. And they’re hinting the official setup used an older completions API that disadvantaged GPT-5.6 Sol compared to what Claude had access to.
ARC Prize co-founder François Chollet didn’t exactly shut this down. He drew a line between custom harnesses built to game the benchmark and general-purpose API settings available to all users — putting OpenAI’s approach in the latter, fair-game category. “We had a lot of back and forth with OpenAI about how to best test their models,” Chollet said, adding he’s glad they’re “starting to figure out the answer.” He acknowledged a “potential parity issue” when different providers use different settings but called it acceptable with transparent reporting.
The whole episode exposes a tension that’s been simmering for years. What are we actually measuring here? Pure model intelligence or the cleverness of the plumbing around it? If a feature like retained reasoning genuinely makes a model more capable in the real world, should we penalize it for using it? The 7.8 percent number feels artificially low — nobody would deploy GPT-5.6 Sol that way. But the 38.3 percent figure also feels cherry-picked. The truth is probably somewhere in between, and that’s the problem. Benchmarks are supposed to settle arguments, not create new ones.
💡 Key Takeaways
- OpenAI's 38.3% score relies entirely on API settings that preserve reasoning between steps — without them, GPT-5.6 Sol scores just 7.8% on ARC-AGI-3
- ARC Prize co-founder François Chollet accepts OpenAI's approach as fair game since the settings are general-purpose and available to all API users, not custom-built for the benchmark
- The disagreement exposes a fundamental tension in AI evaluation: benchmarks that strip away real-world tooling may measure a model's raw capability but fail to reflect how anyone actually uses it
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.