OpenAI quietly buries its own AI text detector
Curated by the Inblix editorial team
OpenAI has pulled the plug on its AI classifier, a tool designed to distinguish human-written text from machine-generated prose, because it simply wasn’t accurate enough. The company made the quiet admission in a July 20 update to a blog post, stating the tool is ‘no longer available due to its low rate of accuracy.’ It’s a stark and somewhat humbling conclusion to a project launched just months ago with the hope of helping educators and platforms fight disinformation.
The numbers were never great. In its own evaluations, the classifier correctly identified only 26% of AI-written text as ‘likely AI-written.’ Even worse, it slapped a false positive on human-written text 9% of the time. For a student sweating a plagiarism accusation or a journalist fact-checking a source, that’s a terrifying margin of error. The company admitted the tool was ‘not fully reliable’ and ‘very unreliable on short texts below 1,000 characters,’ making it practically useless for the kind of quick checks people actually need.
This failure highlights a fundamental problem that’s far from solved. OpenAI’s research team pointed out the inherent cat-and-mouse game, noting that ‘AI-written text can be edited to evade the classifier’ and that ‘it is unclear whether detection has an advantage in the long-term.’ The core issue is a technical one: classifiers based on neural networks are poorly calibrated when they encounter data unlike their training set, leading to moments where the model is ‘extremely confident in a wrong prediction.’ The company is now shifting its focus to provenance techniques for audio and visual content, suggesting a deeper, watermarking-style approach rather than a surface-level text analysis.
OpenAI framed the experiment as a learning experience, a chance to get feedback on ‘whether imperfect tools like this one are useful.’ The consensus from the classroom, apparently, was a resounding ‘no.’ While the company says it remains committed to engaging with U.S. educators and developing new methods, the classifier’s quiet departure leaves the AI-detection market wide open for third-party tools—none of which have cracked the reliability problem either. For now, the most effective detection system remains a teacher with a hunch.
💡 Key Takeaways
- OpenAI's classifier failed miserably on its own terms, catching only 26% of AI text while falsely accusing human writing 9% of the time.
- The tool was practically useless for short-form content, a major blind spot given how much online text falls under 1,000 characters.
- OpenAI admits it's not clear if detection can ever win the long-term arms race against increasingly sophisticated evasion techniques.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.