Baseten Joins Hugging Face Hub, Bringing DeepSeek V4 Flash to Serverless Inference
Curated by the Inblix editorial team
Baseten is now a supported Inference Provider on the Hugging Face Hub, a move that plugs its serverless infrastructure directly into the platform’s model pages and SDKs. For developers, this means you can call models like the latest DeepSeek V4 Flash or Kimi K3 without leaving the Hugging Face ecosystem — no separate billing dashboard, no extra integration work.
The integration covers conversational and text-generation tasks at launch, with more modalities promised soon. It surfaces inside the Hub’s web UI where you can set API keys and rank providers by preference, and it’s baked into the Python and JavaScript client libraries. If you use your own Baseten key, requests go straight to Baseten and you’re billed there. If you authenticate through Hugging Face, the call gets routed automatically and you pay standard provider rates with no markup — at least for now. Hugging Face hints at future revenue-sharing deals but isn’t adding a surcharge today.
Code examples show how trivial the switch is. You point your OpenAI client at the Hugging Face router, append :baseten to the model name, and that’s it. Agent harnesses like Pi, OpenCode, and OpenClaw already support this routing, so teams building on those tools can swap in Baseten-hosted models without glue code.
This is less about flashy technology and more about distribution. Baseten gets a direct line to Hugging Face’s massive developer base, while Hugging Face expands its inference provider roster beyond the usual suspects. The real question is how pricing and performance compare when these models are served through Baseten versus alternatives — and whether the convenience of a unified billing experience is enough to make developers stick around. PRO users get $2 in monthly inference credits to experiment across providers, which is a smart, low-friction way to drive adoption.
💡 Key Takeaways
- Baseten's integration lets developers call frontier LLMs like DeepSeek V4 Flash and Kimi K3 directly from Hugging Face model pages and SDKs.
- Routed requests through Hugging Face carry no markup — you pay standard provider rates, while PRO subscribers get $2 in monthly inference credits.
- Agent harnesses including Pi and OpenClaw already support the routing, eliminating custom integration work for teams using those tools.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.