Alibaba's 2.4T param Qwen3.8-Max dethrones DeepSeek on Arena.AI, but at 14x the cost
Curated by the Inblix editorial team
The race to build the biggest—and cheapest—AI model in China just got a new frontrunner, depending on what you’re measuring. Alibaba dropped Qwen3.8-Max, a 2.4 trillion parameter behemoth that instantly shot to the top of Arena.AI’s Chinese text leaderboard. It’s a mixture-of-experts architecture that only activates about 95 billion parameters per query, a design trick that keeps it nimble despite its enormous size. The model handles text, images, and video across a 1-million-token context window, and Alibaba claims it even completed a 16-day software engineering project autonomously.
But raw scale is only half the story. Moonshot AI’s Kimi K3, which launched in July, still technically has more parameters at 2.8 trillion, with roughly 104 billion active. The real knife fight is over pricing. Alibaba set Qwen3.8-Max at $2 per million input tokens and $6 per million output tokens, undercutting Kimi K3’s $3 and $15. Yet both look wildly expensive next to DeepSeek’s latest move. The V4-Flash model clocks in at a lean 284 billion total parameters and just 13 billion active, priced at a staggering $0.14 per million input tokens and $0.28 for output.
DeepSeek isn’t trying to win the parameter-count war; it’s waging a price war. And the numbers from Artificial Analysis make that brutally clear. In benchmark testing, V4-Flash averaged three cents per task. For the same tests, Kimi K3 cost 86 cents, OpenAI’s GPT-5.6 Sol hit $1.86, and Anthropic’s Claude Fable 5 topped out at $3.15. That’s a 28x gap between DeepSeek and Kimi K3 for completing an actual piece of work. The per-token rates you see on a pricing page are a mirage—what matters is the total token burn and the number of back-and-forth calls a model needs to finish a job.
Here’s the nuance nobody’s talking about: DeepSeek’s efficiency comes with a performance ceiling. Its V4-Flash scored a 40 on the Intelligence Index, while Kimi K3 hit a 57. You’re saving pennies, but you might be getting a meaningfully dumber model. The real insight from this latest wave of Chinese releases is that deployment strategy is as important as architecture. All three companies continue to release open-weight versions under permissive licenses like MIT, giving developers the option to bypass API markups entirely. That’s a pressure release valve on pricing that Western labs have been much slower to adopt.
💡 Key Takeaways
- Qwen3.8-Max reached #1 on Arena.AI's Chinese text leaderboard but sits behind Anthropic's models in overall rankings, showing that scaling up parameters doesn't automatically close the capability gap with Western frontier labs.
- DeepSeek V4-Flash costs three cents per benchmark task compared to 86 cents for Kimi K3, proving that per-token pricing is a misleading metric when models use vastly different amounts of compute to finish the same job.
- Alibaba, DeepSeek, and Moonshot AI all release open-weight versions alongside their APIs, giving developers a path to sidestep the escalating token economics and run these models on their own infrastructure.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.