AI Pulse by Inblix

Scaleway Brings Sub-200ms AI Inference to Hugging Face, Starting at €0.20/M Tokens

Hugging Face Blog · Sep 19, 2025 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Scaleway Brings Sub-200ms AI Inference to Hugging Face, Starting at €0.20/M Tokens

Hugging Face just got a serious European hardware upgrade. Scaleway is now a supported Inference Provider on the Hub, which means developers can tap into powerful open-weight models like DeepSeek R1, Qwen3, and Gemma 3 directly through serverless API calls running out of Paris.

This isn’t just another model listing. The integration lets you fire off requests without leaving Hugging Face’s ecosystem. You can either route through Hugging Face’s billing—using your PRO plan’s $2 monthly inference credits, for instance—or plug in your own Scaleway API key for direct access. The pricing is aggressive: pay-per-token starts at just €0.20 per million tokens, which puts serious pressure on other providers. Scaleway is leaning hard into performance claims too, promising sub-200ms time-to-first-token, a figure that makes agentic workflows and interactive chat apps actually feel snappy.

What sets this apart is the sovereignty play. All inference runs in Scaleway’s French data centers, a concrete answer to the growing demand for EU-based AI infrastructure that doesn’t route data through US clouds. The platform also supports structured outputs, function calling, and multimodal processing, so it’s clearly built for production workloads, not just experimentation. From the Hugging Face side, the integration is clean: model pages now show Scaleway as a provider option, and both the Python and JS client SDKs support it natively with a simple provider="scaleway" parameter.

I’m watching two things closely. First, whether that sub-200ms latency holds up under real load or if it’s a best-case lab number. Second, Hugging Face mentions they’re not adding any markup on routed requests today, but they’re signaling future revenue-sharing deals with providers. That could shift the economics. For now, though, European developers with latency-sensitive apps or data residency requirements just got a very compelling new option that’s deeply woven into the tools they already use.

💡 Key Takeaways

  1. Scaleway's serverless inference promises sub-200ms first-token latency from EU data centers, a critical spec for interactive AI apps.
  2. Pay-per-token pricing begins at €0.20 per million tokens, with no Hugging Face markup on routed requests for the time being.
  3. The integration supports advanced features like function calling and structured outputs, positioning it for production agentic workflows rather than just model demos.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles