AI Pulse by Inblix

Kimi K3 hits near-frontier performance, but the era of dirt-cheap Chinese AI is over

The Decoder · Jul 17, 2026 · 3 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Kimi K3 hits near-frontier performance, but the era of dirt-cheap Chinese AI is over

The era of surprisingly cheap, high-performance Chinese AI models might be coming to an end. Kimi’s new K3 model is a multimodal beast — 2.8 trillion parameters, a one-million-token context window, and a mixture-of-experts architecture with 896 experts — and its pricing tells a new story. At $3 per million input tokens and $15 per million output tokens, K3 is a dramatic price hike from its predecessor. Per-task costs land around $0.94, which puts it squarely in the same ballpark as Western mid-range models like Sonnet 5 and GPT-5.6 Sol, and about half the price of Opus 4.8. The price of admission for cutting-edge performance is clearly rising, regardless of geography.

On the performance front, K3 is the real deal. Kimi’s own benchmarks place it just behind the top proprietary heavyweights — Claude Fable 5 and GPT 5.6 Sol — while it handily beats everything else, including Opus 4.8 and Chinese rival GLM-5.2. Independent testing by Artificial Analysis largely backs up the hype, placing K3’s overall intelligence on par with Opus 4.8 and GPT-5.5. On agentic tasks, it posted a massive jump to an Elo rating of 1,668, leaving GLM-5.2 and GPT-5.5 in the dust. Artificial Analysis even called it “well-rounded,” with analytical quality nipping at Fable 5’s heels.

There is a glaring catch: K3 hallucinates more. Artificial Analysis flagged that the model’s hallucination rate climbed from 39 percent in its predecessor to 51 percent. It’s getting more questions right — accuracy jumped to 46 percent — but it’s also confidently fabricating answers more often. For a model pitched squarely at long-running, minimally supervised software development, a tendency to hallucinate is not a minor bug. Kimi is positioning K3 as a tool that can analyze massive codebases and execute complex projects using a “Vision in the Loop” system, even building 3D browser games. That level of autonomy requires reliability that a 51 percent hallucination rate simply doesn’t support yet.

Kimi plans to release the full weights by the end of July, calling K3 the first open model in the ~3 trillion parameter class. The technical underpinnings are genuinely interesting, with a new Kimi Delta Attention architecture promising 6.3x faster decoding for massive contexts. But the real signal here is economic. The Chinese AI labs that once undercut everyone on price are now charging rates comparable to their American competitors. The race to the bottom on cost is over; the race to the top on capability just got a lot more expensive.

💡 Key Takeaways

  1. Kimi K3's pricing — $3/$15 per million input/output tokens — signals that the era of ultra-cheap access to high-end Chinese AI models is effectively over.
  2. Independent testing confirms K3's near-frontier agentic performance, but its hallucination rate jumped to 51%, creating serious trust issues for its primary use case in autonomous coding.
  3. The model's new 'Vision in the Loop' system points to a future where AI agents don't just write code but visually verify their own output, a shift toward more autonomous development workflows.
  4. With 2.8 trillion parameters and a million-token context window, K3's architecture is a technical flex, but its real-world value will be determined by whether developers can trust its outputs.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles