AI Pulse by Inblix

Google slashes Gemini 3.7 Flash pricing by 50% to $0.75 per 1M tokens

MarkTechPost · Aug 13, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Google slashes Gemini 3.7 Flash pricing by 50% to $0.75 per 1M tokens

Google shipped Gemini 3.7 Flash just three weeks after its predecessor, and the move signals a clear strategy: compete on price, not just benchmarks. The new model isn’t a ground-up retraining — the model card calls it a refinement of 3.6 Flash with algorithmic improvements to the core reasoning engine. Same 1M-token context window, same March 2026 knowledge cutoff, same multimodal inputs across text, images, audio, and video. What changed is the price tag and the performance in targeted areas.

The headline number is cost. Gemini 3.7 Flash lists at $0.75 per 1M input tokens and $3.75 per 1M output tokens — half the original 3.6 Flash rate and roughly a third the blended cost of Claude Sonnet 5 or GPT-5.6 Terra. At an 80/20 input-output mix, you’re looking at $1.35 blended per 1M tokens versus $3.60 for Sonnet 5 and $4.00 for Terra. That’s the kind of gap that makes CFOs pay attention, especially for teams running long-horizon coding agents or document-heavy automation at volume. The rate is introductory and expires December 31, 2026, after which it doubles to $1.50/$7.50. Google is essentially buying evaluation cycles.

Benchmark results are mixed but directionally positive. On FrontierCode 1.1 Main, 3.7 Flash jumps from 34.4% to 43.6%. DeepSWE v1.1 climbs from 48.6% to 65.3%. WebDev Arena Elo ticks up to 1588, the top score in Google’s comparison table. The document work shows the biggest relative gains: GDP.pdf goes from 22.0% to 34.0%, and AutomationBench leaps from 17.0% to 30.4% — ahead of Sonnet 5’s 10.7% and Terra’s 23.6%. Long-context retrieval at 128k hits 97.0%. But GPT-5.6 Terra still wins on terminal and computer-use benchmarks, and CharXiv Reasoning regressed slightly from 85.2% to 84.5%.

The catch for some teams: there are no open weights. Access runs through the Gemini API, Google AI Studio, Android Studio, and the Gemini Enterprise stack. Startups and mid-market teams get the clearest benefit — always-on agents without a Pro-tier budget. Regulated enterprises with data-residency or air-gap requirements are out of luck. This mirrors the broader industry tension we’ve seen since OpenAI and Anthropic started shipping frontier models: hosted convenience versus deployment control. Google’s bet is that for most buyers, the price-performance curve outweighs the lack of self-hosting options.

💡 Key Takeaways

  1. Gemini 3.7 Flash is an algorithmic refinement of 3.6 Flash, not a new base model, shipped just three weeks after its predecessor.
  2. The introductory pricing of $0.75 input and $3.75 output per 1M tokens is roughly a third the blended cost of Claude Sonnet 5 and GPT-5.6 Terra, but doubles on January 1, 2027.
  3. AutomationBench jumps from 17.0% to 30.4%, putting 3.7 Flash ahead of both Sonnet 5 and GPT-5.6 Terra on enterprise workflow automation.
  4. No open weights means regulated enterprises with air-gap or data-residency requirements cannot deploy it on their own infrastructure.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles