AI Pulse by Inblix

GSK bets $110M that bigger AI models aren't the answer to drug discovery

AI News · Aug 3, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: GSK bets $110M that bigger AI models aren't the answer to drug discovery

GSK is doubling down on a quiet but significant shift in AI-driven drug discovery: the data matters more than the model. The pharma giant expanded its partnership with Relation Therapeutics in a deal worth up to $110 million, tasking the biotech with generating massive, proprietary datasets that measure how human cells react to genetic tweaks and drugs. Forget scaling laws. This is about building a better biological training corpus from scratch.

The collaboration moves beyond the observational studies the two companies previously ran on fibrotic diseases and osteoarthritis. Now, Relation will actively create perturbation data—essentially cause-and-effect cellular readouts—to feed its MORGAN platform. The approach, which Relation calls Lab-in-the-Loop, ties wet-lab experiments directly to computational target identification. It’s a purposeful departure from the industry’s early assumption that hoovering up public data would be enough.

That assumption is crumbling. A June 2025 study in Nature Methods that trained 400 models on 22.2 million cells found something counterintuitive: single-cell AI models hit a performance ceiling fast. Piling on more data didn’t produce the consistent gains you see in large language models. Worse, a 2025 Genome Biology paper showed that foundation models like Geneformer and scGPT often failed to beat simpler methods, while also grappling with batch effects from mismatched public datasets. A separate review flagged the risk of the same cells appearing across multiple public repositories like CZ CELLxGENE, skewing training and leaking into test sets.

What GSK is buying here is a hedge against that messiness. Relation’s proprietary Osteomics bone atlas, built from patient samples, hints at the strategy: tightly controlled, internally generated data tied to specific diseases. It’s an expensive path—and one that acknowledges a hard truth. In biological AI, more data can just mean more noise. The real edge comes from knowing exactly which cellular stories you’re telling your model.

💡 Key Takeaways

  1. GSK's $110 million deal with Relation Therapeutics prioritizes generating proprietary biological data over simply scaling up AI model size.
  2. Recent research shows single-cell foundation models often plateau in performance despite larger training datasets, challenging the 'more data is better' assumption from other AI fields.
  3. Public data repositories introduce batch effects and data leakage risks that can undermine model reliability, making purpose-built datasets a competitive differentiator.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles