OpenAI admits Codex only works 28.8% of the time, opens research push
Curated by the Inblix editorial team
OpenAI is openly acknowledging a hard truth about its code-generating AI: it’s wrong more often than it’s right. The company’s own research agenda document leads with the statistic that Codex, the model powering GitHub Copilot, generates functionally correct code just 28.8% of the time. That number isn’t buried in an appendix—it’s the opening salvo of a call for external researchers to help study the economic ripple effects of a technology that is simultaneously impressive and deeply flawed.
The company, in collaboration with OpenResearch and academics from the University of Toronto and UC Berkeley, is framing this low success rate not as a failure but as the precise reason rigorous study is needed now. The argument is straightforward: if a model that fails seven out of ten times is already reshaping developer workflows, the economic stakes for significantly more capable successors are enormous. Sam Manning and Pamela Mishkin, the paper’s lead authors, are essentially laying down a marker, urging academics and policymakers to treat Codex as a living laboratory. They want methodologies established today that can be applied to the inevitably more powerful models of tomorrow.
Their proposed research agenda is organized around six pressure points: productivity, employment, skill development, competition between firms, consumer prices, and economic inequality. The document doesn’t just list these as abstract concerns. It ties them directly to three specific decision-making realms—deployment policy, AI system design, and public policy. The implication is that waiting for models to become flawless before studying their impact would be a strategic error. By the time an AI codes correctly 90% of the time, the window for proactive governance may have already slammed shut.
The initiative comes with a concrete mechanism for participation. OpenAI is announcing a formal Call for Expressions of Interest, seeking to pair external researchers with the company’s own teams and customers. The goal is to move beyond speculation and generate actual evidence about how these tools alter the economics of software production. The big, unanswered question hanging over the whole endeavor is whether studying a 28.8% success rate tells you anything meaningful about a future where that number is flipped. If the failure modes also fundamentally change, today’s economic data might age poorly.
💡 Key Takeaways
- OpenAI's Codex model produces functionally correct code only 28.8% of the time, a baseline the company is using to justify urgent economic impact research.
- The research agenda targets six specific economic outcomes—including productivity and inequality—to inform deployment, design, and public policy decisions.
- OpenAI is actively recruiting external researchers through a formal call for proposals to measure how code generation models affect firms and labor markets.
- The low success rate raises a critical question about whether studying an early, error-prone model yields insights that will remain valid for more capable future systems.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.