SWE-bench Verified no longer measures real coding ability
Curated by the Inblix editorial team
SWE-bench Verified, once the gold standard for measuring AI models’ autonomous software engineering skills, is now effectively broken for frontier models. A new analysis reveals two critical flaws. First, many of the benchmark’s test cases are flawed: at least 59.4% of audited problems reject functionally correct solutions because of bad tests. Second, and more damningly, all frontier models tested have likely been exposed to the benchmark’s problems during training, since they’re sourced from public open-source repositories. This means they don’t have to reason through problems — they can just recall or reconstruct answers they’ve seen before. As a result, recent performance gains (from 74.9% to 80.9% over six months) reflect data contamination, not real improvement in coding ability. OpenAI has stopped reporting SWE-bench Verified scores and recommends a new benchmark, SWE-bench Pro. Why it matters: This highlights a growing crisis in AI evaluation — as models get trained on more public data, benchmarks that reuse open-source problems become unreliable, forcing the industry to constantly invent new, harder, and more controlled evaluation methods.
💡 Key Takeaways
- A major audit found 59.4% of SWE-bench Verified problems have flawed test cases that reject correct solutions.
- All tested frontier models show signs of having seen benchmark problems during training, allowing them to reproduce known fixes rather than solve problems from scratch.
- OpenAI has stopped reporting SWE-bench Verified scores and recommends the new SWE-bench Pro as a less contaminated alternative.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.