AI Pulse by Inblix

DeepInfra joins Hugging Face Hub, slashing serverless AI costs for devs

Hugging Face Blog · Apr 29, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: DeepInfra joins Hugging Face Hub, slashing serverless AI costs for devs

Hugging Face just got a serious cost-performance boost. DeepInfra, the serverless inference platform known for aggressive per-token pricing, is now a supported Inference Provider on the Hugging Face Hub. This isn’t just another logo on a partner page — it means developers can now tap into over 100 models, from heavyweight LLMs like DeepSeek V4, Kimi-K2.6, and GLM-5.1 to upcoming image and video generators, directly from the Hub’s model pages. No more context-switching between dashboards.

The integration is refreshingly practical. You can set your own DeepInfra API key and pay them directly, or let Hugging Face route the request and bill your HF account with no markup. That last part is worth a double-take: “There’s no additional markup from us; we just pass through the provider costs directly.” For PRO subscribers who already get $2 in free inference credits monthly, this is essentially free compute to kick the tires on models that might otherwise sit behind a credit card wall. Free-tier users get a small quota too, but let’s be real — the PRO plan is where this unlocks value.

On the developer experience side, the announcement isn’t just API docs. DeepInfra support is baked into both Python and JS SDKs, and — crucially — into agent harnesses like Pi, OpenCode, and Hermes Agents. That means you can drop a DeepInfra-hosted model into your coding agent toolchain without writing glue code. For anyone who’s wrestled with provider-specific wrappers on a Friday afternoon, this is the kind of integration that actually saves time.

What’s launching now is text generation and conversational tasks. Image, video, and embedding endpoints are on the roadmap. The real question is whether DeepInfra’s infrastructure holds up under a sudden influx of Hub users who are used to Google or Amazon levels of reliability. If they crack under pressure, the cost advantage won’t matter. But if they deliver, this partnership quietly reshapes the default economics of building with open-weight models.

💡 Key Takeaways

  1. DeepInfra brings its low-cost inference pricing directly into the Hugging Face Hub, with no markup when using Hugging Face's routing.
  2. Devs can swap between billing their DeepInfra account or their Hugging Face account, giving flexibility for teams with existing credits.
  3. Integration with agent harnesses like OpenCode and Pi means you can plug cheap, powerful models into your workflow with no boilerplate code.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles