AI Coding Test SWE-Bench Pro Has 30% Flawed Tasks
Curated by the Inblix editorial team
OpenAI just dropped a bombshell on the AI coding world: the popular SWE-Bench Pro test, which measures how well AI models can write software, is riddled with problems. After a deep dive, OpenAI found that roughly 30% of its tasks are broken—too strict, too vague, or just plain misleading. The test pulls its challenges from real software project histories, but those tasks were designed for humans collaborating, not clean AI evaluations. For instance, one task expected a double space but only told the AI to use a single space, so a correct solution would fail. OpenAI used a mix of automated screening, AI agents, and human reviewers to catch these flaws, and even experienced developers flagged more than the bots. This isn’t just an academic quibble—test results influence whether a model gets released or gets safety checks. With top models’ scores skyrocketing from 23% to 80% in eight months, it’s clear something was off. Why it matters: As AI coding tools become more powerful, relying on flawed benchmarks could lead to overhyped claims or unsafe deployments, so the entire industry needs more rigorous, developer-vetted tests that actually reflect real-world programming.
💡 Key Takeaways
- OpenAI found that roughly 30% of tasks in the SWE-Bench Pro coding test have flaws that distort AI performance assessments.
- The flawed tasks fall into categories like being too strict, too vague, too shallow, or having incorrect task descriptions that mislead AI models.
- OpenAI is calling for the industry to build new, more reliable benchmarks using experienced developers rather than repurposing real software project data.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.