AI Pulse by Inblix

DeepSeek's V4 Flash matches GPT-5.6 Luna for 60% less, thanks to a 98% cache trick

The Decoder · Jul 31, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: DeepSeek's V4 Flash matches GPT-5.6 Luna for 60% less, thanks to a 98% cache trick

DeepSeek just dropped V4 Flash “0731” and the numbers make OpenAI’s recent price cuts look timid. The Artificial Analysis Intelligence Index pegs the new model at 50 points — a ten-point jump from the April 2026 version, and just one point shy of GPT-5.6 Luna. But here’s the kicker: it’s doing that work for about 60 percent less cost per task. Even after OpenAI slashed Luna’s prices by 80 percent, DeepSeek still undercuts them dramatically.

The secret weapon isn’t just efficient engineering. It’s a 98 percent cache discount, which blows past the industry standard of 90 percent. Most developers won’t notice the difference until their monthly bill arrives, but for anyone serving repeated prompts or building applications with shared context, the savings stack up fast. The model also uses 12 percent fewer tokens than its predecessor, so you’re paying for less computation from the jump. DeepSeek didn’t touch the architecture — it’s still 284 billion total parameters with 13 billion active and a million-token context window — which suggests this leap came purely from training improvements.

What’s genuinely impressive is where the gains show up. On GDPval, a benchmark that simulates messy, multi-step office work rather than clean academic tasks, the model jumped from 1,189 to 1,559 Elo points. That’s not incremental. That’s the kind of leap that makes a budget model usable for actual agentic workflows — the stuff enterprises are trying to build right now and failing at because the cheap models keep hallucinating or losing the thread. Speaking of which, hallucination rates are down too, though DeepSeek hasn’t published the raw numbers yet.

The model weights are on Hugging Face under an MIT license, which means anyone can grab them and run. No API key, no usage limits, no surprise deprecation notices. That’s the real threat to OpenAI’s pricing strategy: not just a cheaper competitor, but a competent one you can host yourself. If DeepSeek can keep this cadence — matching frontier budget models within a point or two of quality while staying dramatically cheaper — the economic argument for locking into proprietary APIs gets weaker every quarter.

💡 Key Takeaways

  1. DeepSeek's 98 percent cache discount — well above the 90 percent industry norm — is the primary driver of its 60 percent cost advantage over GPT-5.6 Luna, not just lower base pricing.
  2. The model's largest performance leap came in agentic tasks, with GDPval scores jumping from 1,189 to 1,559 Elo, making it viable for complex office automation where budget models typically fail.
  3. Weights are available on Hugging Face under an MIT license, letting enterprises self-host a model that nearly matches GPT-5.6 Luna without API dependency or recurring costs.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles