AI Pulse by Inblix

SandboxAQ drops 5.2M AI-generated drug complexes, cuts pharma R&D time by 75%

Hugging Face Blog · Sep 2, 2025 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: SandboxAQ drops 5.2M AI-generated drug complexes, cuts pharma R&D time by 75%

The biggest bottleneck in AI-driven drug discovery isn’t algorithms. It’s data — specifically, the lack of high-quality 3D structural data linked to real-world potency measurements. SandboxAQ just torched that bottleneck with SAIR, a massive open-source dataset of 5.24 million computationally co-folded protein-ligand complexes, each paired with a curated IC₅₀ value from ChEMBL or BindingDB. The dataset is free under a CC BY 4.0 license, meaning pharma and biotech teams can plug it into their pipelines starting today.

The scale of the compute job alone is worth noting. Generating SAIR required over 130,000 GPU hours on a cluster of 760 NVIDIA H100s, leveraging NVIDIA DGX Cloud through Google Cloud Platform. Working with NVIDIA’s AI Accelerator team, SandboxAQ engineers squeezed more than 95% GPU utilization out of the system, finishing the job in three weeks instead of the projected three months. That’s a 4x speed-up that slashed what could have been a fiscal quarter of compute time down to a sprint.

But bragging about GPU flops is just vendor theater if the data is garbage. SandboxAQ ran every predicted complex through PoseBusters, the go-to open-source benchmarking tool for structural AI in drug discovery, and 97% of structures passed all chemical sanity and physical plausibility checks. The team also benchmarked affinity prediction methods — including 3D CNNs and graph neural networks — against the dataset’s synthetic structures and experimental IC₅₀ values, with full results available in a bioRxiv preprint.

The real strategic play here is cracking the “dark proteome” — disease-relevant proteins that lack any experimentally validated structure. By providing co-folded structures built with the Boltz1 model, SAIR gives researchers a starting point for targets that were previously invisible to structure-based drug design. Deep-learned affinity models like Boltz-2, trained on similar data, have already demonstrated up to a 1,000x speed-up over traditional first-principles approaches. That doesn’t mean we’re hitting “print drug” from a text prompt anytime soon, but it does mean the most expensive and time-consuming parts of hit-to-lead optimization are shifting decisively from wet lab to in silico.

💡 Key Takeaways

  1. SAIR's 5.24 million co-folded complexes directly link 3D structure to IC₅₀ potency data, filling the gap that has kept AI models from predicting critical drug properties reliably.
  2. The dataset passed PoseBusters validation at a 97% rate, giving computational chemists a trustworthy foundation for affinity prediction and virtual screening without waiting for crystallography.
  3. Training on this class of data has already yielded 1,000x speed-ups in affinity prediction, compressing months of lead optimization into hours of GPU time.
  4. SAIR covers targets in the 'dark proteome' that lack experimental structures, potentially unlocking entirely new target classes for AI-guided drug discovery.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles