AI Pulse by Inblix

AI speeds up drug discovery but hits a 'data wall' of missing failures

MIT Technology Review · Jul 27, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: AI speeds up drug discovery but hits a 'data wall' of missing failures

AI is letting drug companies design new molecules from scratch, but Paul Belcher, director of protein research strategy at Cytiva, warns the industry has a dirty secret: the models are starving for the data nobody wants to publish. The core promise is real — AI can screen massive virtual libraries and eliminate lousy candidates before a single pipette is lifted. That shifts drug hunting from empirical screening to predictive design, a move Belcher sees as a genuine acceleration. But that acceleration slams straight into a physical bottleneck. Labs are now flooded with more complex, AI-generated compounds that demand richer characterization than the old high-throughput, yes-or-no screening methods were built to handle.

The bigger problem, though, is what feeds the algorithms. AI models are hitting a wall because they’re all trained on the same public datasets, which share a fatal flaw: they’re a highlight reel of positive results. No one publishes the failures, the compounds that didn’t bind, or the dead ends that litter real research. Belcher calls this publication bias a hand tied behind the back of the entire field. Without a broad diet of negative data, models can’t learn what failure looks like, making their predictions less reliable and inherently biased.

Belcher half-jokes about launching a journal of negative data to unlock what’s currently buried in lab notebooks. It’s a pithy line that gets at a hard truth: the industry’s competitive instinct to guard its fumbles is directly at odds with building the robust, authentic datasets AI needs. Adding fabrication to the mix only deepens the trust problem. The technology is clearly capable of compressing timelines, but it’s becoming obvious that the real bottleneck isn’t compute power — it’s the lack of a cultural and technical infrastructure for capturing the full, messy picture of the scientific process.

💡 Key Takeaways

  1. Flooding labs with AI-designed drug candidates is breaking traditional high-throughput screening workflows that weren't built for detailed characterization.
  2. AI models across the industry are converging on the same conclusions because they share the same public training data, leading to diminishing returns.
  3. The absence of published negative data—failed experiments and non-binding compounds—keeps models from learning what failure looks like, undermining prediction reliability.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles