OpenAI drops GPT-2 774M model, admits synthetic text fools 72%
Curated by the Inblix editorial team
The staged release of GPT-2 is no longer a waiting game for the medium-sized model. OpenAI just put out the 774 million parameter version, the third step in a cautious rollout that started back in February with a 124M model. This is the last checkpoint before the full 1.5 billion parameter behemoth, and the decision to keep that one locked up for now rests on a pile of fresh, slightly unnerving research the company is publishing alongside the code. It’s a rare move: a major lab openly tying a product launch to societally-focused academic findings.
Those findings paint a tricky picture. Research from Cornell’s Sarah Kreps and Miles McCain, published in Foreign Affairs, found that people deemed GPT-2’s synthetic text samples credible 72% of the time. Real New York Times articles fared only slightly better at 83%. In a separate effort, the Allen Institute for AI and the University of Washington showed that text from a system called GROVER could be more plausible than human-written propaganda. “Humans can be convinced by synthetic text,” the post states bluntly, and that realization is driving a more conservative posture. If the gap between fake and real is already this narrow at 774M parameters, the full model’s persuasive power is an open and worrying question.
Detection isn’t saving the day either. The technical report admits that current machine learning detectors scrape the low-to-mid 90% accuracy range, far from the 99.9% to 99.99% reliability a real-world system would need to avoid a flood of false positives. Malicious actors can fine-tune models or tweak sampling methods to evade detection, which makes a purely automated defense a brittle proposition. OpenAI’s stance is that statistical tools will need to be bolted to human judgment and metadata analysis to be effective. That’s not a fix. It’s a workaround.
The path to the 1.5B release now runs through four external research partners. The Middlebury Institute’s CTEC is probing terrorist misuse scenarios, the University of Oregon is building bias probes, and UT Austin is stress-testing statistical detection after domain-specific fine-tuning. OpenAI’s current plan is still to release the big model in a few months, but the post explicitly notes that a partner’s findings, or actual malicious use of this 774M model, could scrap that timeline. They’re also open-sourcing a non-commercial legal agreement to make model-sharing partnerships easier, a quiet signal that they expect this coordination problem to become the norm, not the exception.
💡 Key Takeaways
- People found GPT-2 text credible 72% of the time compared to 83% for genuine New York Times articles, a gap that will likely shrink with larger models.
- OpenAI states that ML-based detection methods are not reliable enough for deployment, requiring human oversight to supplement statistical tools.
- The release of the full 1.5B parameter model is contingent on findings from four external partners studying bias, terrorism risks, and detection evasion.
- Open-source legal agreements for model sharing are being positioned as a foundational tool for responsible AI publication going forward.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.