AI Pulse by Inblix

Claude 3.5 Sonnet Scores Just 21% on Brutal AI Research Replication Test

OpenAI Blog · Jul 14, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Claude 3.5 Sonnet Scores Just 21% on Brutal AI Research Replication Test

If you want to know how far AI agents really are from replacing PhD-level machine learning researchers, a new benchmark has a sobering number for you: 21%. That’s the top score achieved by Anthropic’s Claude 3.5 Sonnet on PaperBench, an exceptionally rigorous new test that demands agents replicate entire research papers from scratch.

The benchmark, open-sourced by its creators, doesn’t mess around with toy problems. It compiles 20 papers accepted as Spotlight or Oral presentations at the prestigious ICML 2024 conference. To fully replicate one, an agent must digest the paper’s core contributions, build a complete codebase from the ground up, and execute the experiments necessary to reproduce the reported results. This isn’t about answering multiple-choice questions; it’s about doing the job.

Scoring such a complex, long-horizon task required a novel approach. The team worked directly with the original authors of each ICML paper to co-develop hierarchical rubrics. These granular scoring guides break down every replication into thousands of smaller, verifiable sub-tasks. The entire benchmark hinges on a staggering 8,316 individually gradable tasks, all designed for objective evaluation. To manage this scale, the researchers also built an LLM-based automated judge, creating a separate benchmark just to validate the judge’s own accuracy to ensure reliable grading.

When pitted against this gauntlet, frontier models stumbled hard. Even the best-performing system, Claude 3.5 Sonnet equipped with open-source scaffolding to manage its environment and tools, could only nail about a fifth of the required work. To set a human baseline, the authors recruited top-tier machine learning PhDs to tackle a subset of the papers. The models, it turns out, couldn’t beat them. That gap highlights just how much tacit knowledge, iterative debugging, and experimental intuition remains uncaptured by current AI engineering pipelines, leaving a wide chasm between an impressive demo and true scientific labor.

💡 Key Takeaways

  1. The top-scoring agent achieved only a 21.0% average replication score, failing to execute the majority of complex experimental tasks required by the ICML papers.
  2. The benchmark's 8,316 gradable tasks were co-developed with original paper authors, grounding the evaluation in expert-validated criteria rather than generic metrics.
  3. A human baseline established with top ML PhDs confirmed that frontier AI agents cannot yet outperform skilled researchers on authentic, end-to-end scientific replication.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

← Back to all articles