AI Pulse by Inblix

China's 3-way open-weight war: Kimi K3 leads on smarts, DeepSeek V4 Pro wins on price

MarkTechPost · Jul 19, 2026 · 3 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: China's 3-way open-weight war: Kimi K3 leads on smarts, DeepSeek V4 Pro wins on price

The open-weight leaderboard just got a Chinese accent. Moonshot AI’s Kimi K3, DeepSeek V4 Pro, and Zhipu’s GLM-5.2 now sit at the top, and they’re not just lab experiments—these are trillion-parameter-class Mixture-of-Experts (MoE) models shipping with a million-token context windows and a clear focus on long-horizon coding and agent workloads. If you’re an AI team lead, the decision tree has three branches: raw capability, licensing freedom, and serving cost. The one you pick says more about your infrastructure budget than your technical ambition.

On pure capability, Kimi K3 is the heavyweight. This 2.8-trillion-parameter beast scores a 57 on the Artificial Analysis Intelligence Index, landing it in third place globally behind only Anthropic’s Claude and OpenAI’s GPT-5.6. That’s a genuinely elite result. But there’s a catch—a big one. You can’t download the weights until July 27, 2026, and when you do, Moonshot’s Modified MIT license kicks in with an attribution clause triggered at 100 million monthly active users. Until then, you’re renting it through an API at a list price where a dollar buys you about 67,000 output tokens. Moonshot’s own benchmarks show K3 clobbering GLM-5.2 on every shared coding test, but for now it’s more trophy than tool for self-hosters.

DeepSeek V4 Pro is the pragmatist’s choice. At 1.6 trillion total parameters with 49 billion active, it’s MIT-licensed with weights on Hugging Face from day one—no waiting, no asterisks. And the pricing is almost absurd: $0.18 per million tokens on a blended basis, per Artificial Analysis. That means a dollar gets you roughly 1.15 million output tokens. For context, that’s about 17 times more tokens per dollar than K3. It’s not the absolute smartest of the three, but a SWE-bench Verified score of 80.6%—tying Gemini 3.1 Pro—makes it competitive where it counts. If you’re building a product on top of an open model today, this is the default.

GLM-5.2 is the interesting middle child. At 744 billion parameters with roughly 40 billion active, it’s the smallest of the trio and the fastest, clocking 168 tokens per second in third-party testing. It’s also MIT-licensed and fully self-hostable—though you’ll need over 1TB of VRAM in BF16, which means roughly eight H200s at FP8. Its $0.90 per million token price puts it between the two rivals, and its capabilities punch above its weight class, having held the open-weight crown until K3 shipped. For teams that want a downloadable model today with strong performance and can’t stomach DeepSeek’s slower inference speed, GLM-5.2 is the sleeper pick. The real question is whether K3’s 2026 weight release will even matter by then, given how fast this field is moving.

💡 Key Takeaways

  1. Kimi K3 is the most capable open-weight model at a 57 on the AI Analysis Index, but its weights remain locked behind an API until July 2026.
  2. DeepSeek V4 Pro delivers 1.15 million output tokens per dollar—roughly 17 times cheaper than Kimi K3—with a clean MIT license and immediate Hugging Face access.
  3. GLM-5.2 is the fastest of the three at 168 tokens per second and remains the most practical self-hosting option for teams that need strong performance today.
  4. The real differentiator for teams isn't benchmark scores but the trade-off between K3's elite capability, DeepSeek's extreme cost efficiency, and GLM-5.2's deployment speed.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles