Qwen3.8 Max catches Claude Opus 4.8 but burns 15x the tokens to do it
Curated by the Inblix editorial team
Alibaba’s Qwen3.8 Max just posted a 56 on the Artificial Analysis Intelligence Index — a 10-point leap over its predecessor and enough to tie Claude Opus 4.8. That’s the headline number. But the way it gets there tells a messier story.
The model’s GDPval-AA score, which measures performance on work-related tasks, jumped 468 Elo points to 1,739. Only Claude Opus 5 scores higher. Yet Qwen3.8 Max required 64 steps per task to hit that mark, compared to just 14 for the previous version. Input token consumption ballooned 15x because the benchmark resends the full conversation history at every step. That’s not efficiency — it’s brute force with a bigger compute bill.
Alibaba did slash token prices: input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits halved to $0.25. But the per-task math still stings. A single Intelligence Index task now costs $1.14, more than double Qwen3.7 Max at $0.53. Meanwhile, Kimi K3 scores one point higher on the same index for just $0.86 per task, and GLM-5.2 manages a 51 at $0.57. Price-to-performance ratios don’t lie — Qwen3.8 Max is working harder, not smarter.
There are regressions worth flagging. The AA-LCR benchmark, which tests whether a model can pull together information from very long texts, dropped 2 points. AA-Omniscience fell 10 points, and while the accuracy rate held around 31 percent, the hallucination rate spiked from 23 to 40 percent. That’s a significant shift toward guessing rather than admitting ignorance. If you’re deploying this in a setting where honest uncertainty matters more than confident wrong answers, that 40 percent number should give you pause. Alibaba clearly prioritized benchmark scores over the kind of reliability that actually matters in production.
💡 Key Takeaways
- Qwen3.8 Max ties Claude Opus 4.8 on the Intelligence Index but uses 4.6x more steps and 15x more input tokens to get there
- Per-task cost more than doubled to $1.14, making it pricier than higher-scoring Kimi K3 at $0.86 per task
- The hallucination rate jumped from 23% to 40%, meaning the model guesses far more often instead of admitting it doesn't know
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.