OpenAI fires back at NYT, calls regurgitation a 'rare bug'
Curated by the Inblix editorial team
OpenAI has broken its silence on the copyright lawsuit filed by The New York Times, framing the legal battle as a misunderstanding of its technology while simultaneously calling the paper’s claims misleading. In a blog post published Monday, the company laid out a four-point rebuttal that mixes a staunch defense of its training practices with an admission that “regurgitation” of paywalled content is a technical flaw it’s working to fix.
More than 92% of Fortune 500 companies now build on OpenAI’s products, a statistic the company leverages to underscore its argument that training on publicly available data is not just legal, but essential. The company leans heavily on the “fair use” doctrine, citing support from a broad coalition of academics, startups, and civil society groups who have submitted comments to the US Copyright Office. OpenAI also points to international momentum, highlighting that jurisdictions like the European Union, Japan, and Singapore have laws permitting the training of models on copyrighted content—a move they frame as critical for staying competitive against global rivals.
But the legal arguments are only half the story. OpenAI is clearly trying to appear conciliatory. The post reveals that negotiations with the Times were apparently progressing well as recently as December 19, just days before the lawsuit dropped. Those talks centered on a “high-value partnership” that would display Times content with real-time attribution in ChatGPT, a deal OpenAI implies the paper walked away from. The company also touts an opt-out mechanism it says the Times began using in August 2023, positioning itself as a “good citizen” that goes beyond what the law requires.
The most technical concession comes in addressing the lawsuit’s core evidence: examples of ChatGPT spitting out near-verbatim copies of Times articles. OpenAI describes this memorization as a “rare failure of the learning process,” one that’s more likely when identical passages appear across multiple public websites. While they’re working to drive the bug to zero, the company also takes a swipe at the Times, suggesting the examples were produced by someone “intentionally manipulating” prompts—a violation of OpenAI’s terms of use. It’s a messy fight that pits a dominant AI lab against a legacy media giant, and neither side seems willing to blink.
💡 Key Takeaways
- OpenAI admits memorization is a real bug that it is actively trying to fix, but claims the examples cited by the Times required manipulative prompting that violates its terms of service.
- The company reveals that partnership talks with the Times were progressing until December 19, suggesting the lawsuit caught them off guard and killed a potential real-time display deal with attribution in ChatGPT.
- OpenAI is wielding its massive enterprise adoption—over 92% of Fortune 500 companies—as a strategic argument that restricting training data would cripple a technology the business world now depends on.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.