AI Pulse by Inblix

Hugging Face kills HUGS microservice just 2 years after launch

Hugging Face Blog · Oct 23, 2024 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Hugging Face kills HUGS microservice just 2 years after launch

Hugging Face has quietly pulled the plug on HUGS, its Generative AI Services platform for deploying open models, marking a surprisingly short two-year run for what was once touted as a zero-configuration path to enterprise inference. The shutdown notice, buried in a September 2025 update on the company’s announcement page, comes with a blunt redirect: customers should now look to the Dell Enterprise Hub or Azure AI Foundry’s Hugging Face Collection instead.

When HUGS launched, it was pitched as the answer to a genuine headache. Engineering teams were burning weeks tuning inference workloads for specific GPUs, and HUGS promised to collapse that timeline to minutes on NVIDIA and AMD hardware, with AWS Inferentia and Google TPU support on the roadmap. Early adopters like Polyconseil and Orange were quoted singing its praises — Polyconseil’s CTO said it slashed deployment from a week to under an hour, while an Orange research engineer called it a confidence builder for scaling internal usage. Those testimonials now read like a eulogy for a product that couldn’t find its footing.

What likely went wrong is the same thing that plagues many managed AI services: the economics of running inference containers on cloud marketplaces are brutal. At $1 per hour per container on AWS and GCP, plus the underlying compute costs, HUGS was competing against an ecosystem where open-source tools like vLLM and Ollama are free and increasingly easy to configure. The DigitalOcean partnership, which offered HUGS at no additional cost beyond GPU Droplets, might have been an attempt to find product-market fit in a more price-sensitive channel — but apparently it wasn’t enough to sustain the service.

For teams that built workflows around HUGS, the writing was on the wall the moment Hugging Face started directing users to Dell and Azure. Those aren’t neutral recommendations; they’re off-ramps to partners who can absorb the support burden that Hugging Face clearly decided wasn’t worth carrying. The open model inference space is consolidating fast, and this shutdown suggests that even a company as central to the open-source AI ecosystem as Hugging Face couldn’t make a paid inference layer work against the gravitational pull of free, community-maintained alternatives.

💡 Key Takeaways

  1. HUGS is officially discontinued as of September 2025, with Hugging Face directing users to Dell Enterprise Hub and Azure AI Foundry as replacements.
  2. Early enterprise users reported deployment time dropping from a week to under an hour, but that efficiency gain wasn't enough to sustain the product commercially.
  3. The $1 per container per hour pricing model on top of cloud compute costs likely failed to compete with free, increasingly polished open-source inference tools like vLLM.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles