AI Pulse by Inblix

Google ships 3 budget Gemini Flash models while its missing Pro cedes ground to rivals

The Decoder · Jul 21, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Google ships 3 budget Gemini Flash models while its missing Pro cedes ground to rivals

Google dropped three new Gemini models on Tuesday, but the launch felt less like a power move and more like a stall tactic. The headliner everyone wants — Gemini 3.5 Pro — remains stuck in partner testing with no ship date, while the company admits pretraining for Gemini 4 is already underway. Logan Kilpatrick, a member of Google’s technical staff, deflected criticism on X by insisting the explicit goal for these Flash releases was efficiency, not raw performance.

What did ship is a trio of workhorses with very specific job descriptions. Gemini 3.6 Flash cuts output token usage by roughly 17 percent compared to its predecessor, with savings hitting 65 percent on coding benchmarks like DeepSWE. At $1.50 per million input tokens, it’s priced to move high-volume workloads. The 3.5 Flash-Lite variant cranks throughput to 350 output tokens per second for just $0.30 per million input tokens — numbers that will matter to anyone running massive inference pipelines. Then there’s the restricted one: Gemini 3.5 Flash Cyber, a security-focused model that scored 83.2 percent on the CyberGym benchmark, just two points shy of OpenAI’s GPT-5.5-Cyber despite being a fraction of the size.

Google’s Big Sleep vulnerability research team put Flash Cyber to work hunting zero-days. Scanning commits in Chrome’s V8 JavaScript engine, it surfaced 55 confirmed unique findings — outpacing the standard Flash model’s 47 and Anthropic’s Claude Opus 4.6 at 36. Ten of those findings were exclusive to the Cyber model. In a separate test, Google’s cloud vulnerability researchers used it to find remote code execution flaws in public APIs within two hours and produce a working exploit that bypassed security protections. That offensive capability is precisely why Google is keeping access locked down to governments and trusted partners.

Honestly, the whole launch reads like damage control. While OpenAI, Anthropic, and even Meta push their frontier models forward, Google is selling efficiency gains and a cybersecurity side project. The benchmark improvements are real — 3.6 Flash jumps from 37 to 49 percent on DeepSWE — but efficiency wins don’t grab headlines when your competitors are shipping models that redefine capability ceilings. Google knows this. The buried acknowledgment that Gemini 4 is already in training feels like a nervous promise: we’re still in the race, just not from the front.

💡 Key Takeaways

  1. Google's flagship Gemini 3.5 Pro remains delayed with no ship date, while pretraining for Gemini 4 has already begun, signaling a widening gap behind OpenAI and Anthropic on frontier capability.
  2. Gemini 3.5 Flash Cyber found 55 confirmed unique vulnerabilities in Chrome's V8 engine, 10 of which no other tested model discovered, but Google restricts access due to its offensive potential.
  3. The new Flash models prioritize cost and speed — 3.6 Flash cuts token usage by up to 65% on coding tasks and Flash-Lite hits 350 tokens per second — but none challenge competitors on raw performance.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles