NYT Claims OpenAI Hid Evidence of Copyright Infringement
Curated by the Inblix editorial team
The New York Times and The Daily News are accusing OpenAI of lying about its ability to search its training data and chat logs for their copyrighted content. This is the latest twist in a two-year lawsuit where the outlets claim OpenAI trained its AI models on their articles and reproduced their journalism. OpenAI previously argued it couldn’t easily search its massive dataset, but a court-ordered deposition allegedly revealed the company had already built tools to do exactly that. A data privacy engineer reportedly admitted OpenAI had an internal database of 78 million de-identified chat logs and a filter called “Bloom” to detect regurgitation. The plaintiffs say OpenAI made it impossible to get reliable evidence by heavily redacting a sample of 20 million logs and possibly deleting billions of outputs after a preservation order. Why it matters: This case could set a precedent on how AI companies handle copyright disputes and discovery—essentially testing whether tech giants can hide behind technical loopholes when accused of copying creative work.
💡 Key Takeaways
- OpenAI allegedly had internal systems to search its training data for copyrighted works, contradicting earlier claims it couldn't do so.
- The company reportedly built a database of 78 million chat logs and a filter called 'Bloom' to detect content regurgitation.
- Plaintiffs argue OpenAI made it deliberately hard to access evidence, including heavy redactions and possible deletion of logs after a court order.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.