AI Pulse by Inblix

Google TPUs Hit Hugging Face at $1.37/Hour, Targeting Llama and Mistral Deploys

Hugging Face Blog · Jul 9, 2024 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Google TPUs Hit Hugging Face at $1.37/Hour, Targeting Llama and Mistral Deploys

Hugging Face developers can now rent Google’s custom TPU v5e chips directly through Inference Endpoints and Spaces, with prices starting at $1.375 per hour for a single-core configuration with 16 GB of memory. The collaboration between the two companies brings Google’s AI-specific silicon to a platform that has largely been associated with NVIDIA GPUs until now. Three instance sizes are available at launch: the entry-level v5litepod-1, a four-core v5litepod-4 with 64 GB of memory at $5.50 per hour, and an eight-core v5litepod-8 with 128 GB of memory at $11 per hour. Hugging Face recommends the single-core option for models up to roughly 2 billion parameters, while suggesting the four-core tier for anything larger to avoid memory constraints. The company also notes that bigger configurations translate directly to lower latency.

The rollout is backed by Optimum TPU, an open-source library built jointly by Hugging Face and Google’s product and engineering teams. That library works alongside Text Generation Inference, the serving framework Hugging Face uses for large language models, to make deployment on TPUs a relatively painless affair. At launch, supported architectures include Gemma, Llama, and Mistral — three of the most widely used open model families in production today. Hugging Face says additional model architectures are in the works.

Pricing here is worth a close look. Google has positioned TPU v5e as a cost-efficient alternative to high-end GPUs for inference workloads, and these hourly rates undercut many comparable GPU offerings on major cloud platforms. Still, the model compatibility list remains narrow compared to NVIDIA’s CUDA ecosystem, where virtually every open model works out of the box. That gap will determine how quickly developers actually adopt TPUs for anything beyond experimentation.

Spaces users get the same three configurations, accessible through the Settings menu. The move signals Google’s growing willingness to push its AI hardware beyond its own cloud walls and into developer communities where mindshare is won and lost. With NVIDIA’s H100 supply constraints still fresh in memory, a credible alternative at these prices could shift how smaller teams think about inference infrastructure.

💡 Key Takeaways

  1. Google TPU v5e instances are now available on Hugging Face Inference Endpoints and Spaces starting at $1.375 per hour for 16 GB of memory.
  2. The integration supports Gemma, Llama, and Mistral model architectures at launch, with more planned through the open-source Optimum TPU library.
  3. Hugging Face recommends the four-core v5litepod-4 tier for models larger than 2 billion parameters to avoid memory budget issues.
  4. TPU pricing undercuts many comparable GPU cloud offerings, but limited model compatibility compared to NVIDIA's CUDA ecosystem remains the key adoption barrier.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles