AI Pulse by Inblix

Hugging Face and NVIDIA team up to rent GPU clusters by the training run

Hugging Face Blog · Jun 11, 2025 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Hugging Face and NVIDIA team up to rent GPU clusters by the training run

The compute gap between AI’s haves and have-nots has looked like a widening chasm lately, with gigawatt-scale superclusters dominating headlines. Hugging Face and NVIDIA are betting that the fix isn’t building more mammoth datacenters—it’s making the GPUs already out there actually accessible. Their new “Training Cluster as a Service” lets any of the 250,000 organizations on Hugging Face request a GPU cluster sized to their needs, paying only for the duration of a training run.

Here’s the plumbing: NVIDIA Cloud Partners supply the hardware—think Hopper and GB200 systems—housed in regional datacenters and centralized within DGX Cloud. The newly announced DGX Cloud Lepton layer handles provisioning and scheduling. Hugging Face brings the developer tooling and open-source libraries so researchers aren’t wrestling with infrastructure. You request a cluster at hf.co/training-cluster, and if accepted, both companies collaborate to source, price, and configure it to your specs.

A handful of early users are already validating the model. The Telethon Institute of Genomics and Medicine (TIGEM) is using it to train models that predict pathogenic variants for rare genetic diseases. Non-profit Numina, which won the 2024 AIMO progress prize, sees this as their path to building open alternatives to closed systems like DeepMind’s AlphaProof—their bottleneck has been raw compute, not ambition. Startup Mirror Physics is producing high-fidelity chemical models at a scale they couldn’t reach before.

The real play here might be less about the technology and more about the business model. By decoupling GPU access from long-term leases and making it episodic, Hugging Face and NVIDIA are essentially creating a spot market for serious training compute—one where a university lab in Naples gets the same shot at GB200s as a Silicon Valley unicorn. Whether the pricing and availability hold up under demand is the open question, but the direction is clear: making the GPU-rich’s hardware work for the GPU-poor, one training run at a time.

💡 Key Takeaways

  1. Hugging Face and NVIDIA are offering on-demand GPU clusters priced per training run, not per reserved instance, directly targeting the compute access gap for researchers and smaller organizations.
  2. The service combines NVIDIA's DGX Cloud Lepton for orchestration with Hugging Face's developer tools, sourcing hardware from a network of regional cloud partners rather than a single provider.
  3. Early adopters like TIGEM, Numina, and Mirror Physics are already training domain-specific models in genomics, mathematical reasoning, and materials science using the service.
  4. The episodic pricing model could reshape how research compute is funded and accessed, shifting power away from organizations that can afford long-term GPU leases.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

← Back to all articles