AI Pulse by Inblix

Microsoft's MAI Code 1.1 Flash gets crushed by DeepSeek on benchmarks and cost

The Decoder · Aug 12, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Microsoft's MAI Code 1.1 Flash gets crushed by DeepSeek on benchmarks and cost

Microsoft has rolled out MAI Code 1.1 Flash, a new code-generation model purpose-built for GitHub Copilot, but the launch raises an awkward question: why bother? The model is undeniably better than its predecessor—25 percent more token-efficient, a quarter of the cost, and developers accepted 4 percent more of its suggestions. Training leaned heavily on Copilot telemetry, with Microsoft citing “hundreds of thousands of reinforcement-learning environments” derived from real-world usage. On paper, it’s a solid iteration. On benchmarks, however, it gets flattened by DeepSeek-V4-Flash-0731, an openly available alternative that beats MAI Code 1.1 Flash on both raw capability and cost per token.

The pricing optics aren’t great. Microsoft touts a cheaper model than what it shipped in June, but it’s still pricier than DeepSeek’s more performant offering. The company’s public messaging is conspicuously vague—the official announcement skips direct comparisons entirely and instead highlights squishy internal metrics like a 9 percent bump in “return visits.” The real benchmark data sits buried in the model card, which tells you everything about who won. It’s a familiar playbook: when the numbers don’t flatter, change the subject to engagement stats.

What makes this genuinely puzzling is Microsoft’s recent rhetorical pivot toward positioning itself as an open-source AI booster. Satya Nadella’s team has been talking a big game about embracing open-weight models, yet here they are pouring resources into a proprietary, closed model that objectively underperforms a freely available competitor. The strategy contradiction is hard to ignore. If you’re going to champion open AI, maybe don’t build a walled garden with an inferior product inside it.

The motivation likely mirrors the quiet model swap Microsoft pulled in Copilot recently—yanking out Anthropic and OpenAI models in favor of homegrown MAI alternatives. It’s a margin play. Cheaper in-house models mean fatter profits, even if the output quality takes a hit. Most Copilot users never manually select which model powers their autocomplete, so Microsoft can make MAI the default and capture a massive, passive user base. The genius of defaults is that nobody changes them. DeepSeek may win every benchmark that matters, but if Microsoft controls the tap, the better model never reaches the glass.

💡 Key Takeaways

  1. MAI Code 1.1 Flash is cheaper and more efficient than its predecessor but loses to DeepSeek-V4-Flash-0731 on both benchmark performance and token pricing.
  2. Microsoft buried direct benchmark comparisons in the model card and led its announcement with vague engagement metrics instead of head-to-head results.
  3. The closed-source MAI push contradicts Microsoft's recent open-AI rhetoric and mirrors a broader Copilot strategy to swap pricier third-party models for higher-margin in-house ones.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles