AI Pulse by Inblix

Hugging Face and Google Cloud build a CDN to slash open model download times

Hugging Face Blog · Nov 13, 2025 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Hugging Face and Google Cloud build a CDN to slash open model download times

The open model library Hugging Face and Google Cloud are deepening their tie-up with a concrete engineering fix for a massive pain point: slow downloads. The companies announced a new CDN Gateway that caches Hugging Face’s repository of over 2 million models directly on Google Cloud infrastructure. The goal isn’t just speed—it’s about making the supply chain for open AI more robust for the tens of petabytes of models Google Cloud customers download every month.

Jeff Boudier, a Hugging Face exec, framed this as infrastructure for a world where “all companies will build and customize their own AI.” The numbers back up the demand. Usage of Hugging Face on Google Cloud has reportedly grown 10x over the last three years, translating into billions of monthly requests. The new gateway combines Hugging Face’s Xet storage tech with Google’s networking to reduce time-to-first-token, a metric that matters deeply when you’re iterating on model deployments in Vertex AI or GKE.

Google isn’t just the host here; it’s a contributor. Ryan J. Salva from Google Cloud noted the company has put over 1,000 of its own models—like the Gemma family—into the Hugging Face community. The partnership now promises to make deploying those models, or any private model, a two-click affair from a Hugging Face page to Vertex AI’s Model Garden. Hugging Face also plans to make Google’s custom TPU chips a first-class citizen in its libraries, aiming for an experience as seamless as using GPUs.

There’s a security layer baked in, too. The deal will pipe Google’s VirusTotal, Threat Intelligence, and Mandiant capabilities into the Hugging Face Hub to scan models, datasets, and Spaces. It’s a pragmatic move, acknowledging that the speed of open model adoption has outpaced security vetting. For Hugging Face Inference Endpoints customers, the immediate payoff is simpler: Google Cloud’s infrastructure should mean more instances and lower prices. It’s a partnership betting that the future of AI isn’t one massive API, but millions of custom models running on infrastructure companies already control.

💡 Key Takeaways

  1. A new CDN Gateway will cache Hugging Face models on Google Cloud, directly tackling the latency and supply chain fragility of pulling tens of petabytes of open models monthly.
  2. Google is making its custom TPU chips natively supported in Hugging Face libraries, aiming to make them as straightforward to use as GPUs for the 10 million AI builders on the platform.
  3. The partnership integrates Google’s Mandiant and VirusTotal security tools to scan the Hugging Face Hub, addressing the growing attack surface of open model repositories.
  4. Hugging Face Inference Endpoint users should expect price drops and newer Google Cloud instances, as the collaboration moves beyond convenience and into cost performance.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles