DeepSeek V4 Pro update doubles agent scores, then doubles some API prices
Curated by the Inblix editorial team
DeepSeek kicked off the week with a three-part announcement: a new build of its V4-Pro flagship, an open-source agent framework, and API pricing changes that penalize heavy cache users. The model update is substantial. Terminal Bench 2.1 scores jumped from 72.1 to 87.9, while DeepSWE climbed from 12.8 to 62.7 — a fivefold improvement that suggests the company has been focused on agentic workloads rather than general chat. On several agent benchmarks, V4-Pro now beats Claude Opus 4.8.
But the broader picture is more complicated. Artificial Analysis puts V4-Pro at 53 on its Intelligence Index, up from 45. That ties GLM-5.2 but still trails Muse Spark, Qwen 3.8 Max, Kimi K3, and Claude Opus 5, which leads the pack at 63. In other words, DeepSeek has closed the gap on agent-specific tasks while remaining mid-tier for general intelligence. The company hasn’t published weights for the new build, and the April preview remains on Hugging Face — a pattern that’s becoming familiar for a lab that talks about openness but holds back its best artifacts.
DeepSeek Harness v0.1, shipping under the MIT license, is pitched as an open alternative to OpenAI’s Codex and Claude. Built on the Cordis plugin system, it makes tools, sandboxes, sessions, and even the UI swappable. A continuous session log records every prompt, tool call, and result, and runs can be resumed, branched, or replayed. Project lead Cui Tianyi joined DeepSeek from Jane Street in March 2026, and when beta testing opened in early August, 712 projects signed up in three days. That’s real traction for a developer preview.
The pricing changes will hurt the most for agent builders. Off-peak input jumps from $0.435 to $0.66 per million tokens, and output nearly doubles to $1.98. Peak hours — 1 a.m. to 4 a.m. and 6 a.m. to 10 a.m. UTC — double those rates again. Cache hits are the real sting: from $0.003625 to $0.022 off-peak, a sixfold increase that shrinks the cache discount dramatically. For agents that repeatedly read the same files, that’s the most expensive part of the change. The move partially reverses May’s price cuts and arrives as DeepSeek raises capital ahead of a planned IPO. The message is clear: the era of DeepSeek as the cheap option is ending.
💡 Key Takeaways
- DeepSeek's V4-Pro update delivers massive gains on agent benchmarks — DeepSWE jumped from 12.8 to 62.7 — but general intelligence still trails Claude Opus 5 by 10 points on the Intelligence Index.
- DeepSeek Harness v0.1 is now open source under MIT, with a modular plugin system and full session logging, and attracted 712 beta testers in its first three days.
- Cache hit pricing is increasing sixfold off-peak, from $0.003625 to $0.022 per million tokens, making repeated file reads dramatically more expensive for agent workloads.
- The pricing shift reverses part of May's rate cuts and comes as DeepSeek raises capital ahead of an IPO, signaling a strategic move toward monetization over market share.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.