Google kills Gemini 3.5 Flash after 2 months, replaces it with cheaper 3.6
Curated by the Inblix editorial team
Google’s Gemini 3.5 Flash — the model it hyped at I/O just two months ago — is already dead. In its place, the company has dropped Gemini 3.6 Flash, a model that quietly fixes what its predecessor couldn’t get right. Google admits the update was driven by user feedback that 3.5 Flash wasn’t delivering on its coding promises. The numbers back that up: on the DeepSWE coding benchmark, 3.6 Flash scores 49 percent, a solid jump from the previous model’s 37 percent. That’s not an incremental tweak. That’s Google acknowledging a genuine shortfall.
The efficiency story is even more striking. Google claims Gemini 3.6 Flash uses about 17 percent fewer tokens than the model it replaces. In an industry where every token costs real money, that’s the kind of optimization that makes CFOs pay attention. The new API pricing reflects this — output tokens drop from $9 per million to $7.50 per million. For anyone running agentic workflows at scale, fewer steps, fewer tokens, and higher accuracy could translate into real savings. Not earth-shattering, but genuine.
Google isn’t just burying old models. It also released Gemini 3.5 Flash Lite, a speed demon that cranks out 350 tokens per second, and its first cybersecurity-focused model, Gemini 3.5 Flash Cyber. The Flash Lite is interesting as a budget workhorse. Its benchmarks trail the current frontier, but it roughly matches top-tier models from a year ago. At $0.30 per million input tokens, it’s positioned as the model you use when you need to deploy agents without burning cash. The Cyber variant remains more of a mystery, with no detailed benchmarks available yet.
Conspicuously absent from all this is Gemini 3.5 Pro. That model was supposed to launch in June. It’s now late July with no word. Google’s frantic pace on the Flash line — deprecate, replace, iterate — looks less like a strategy and more like a scramble. The Flash models are getting better and cheaper, but the flagship is MIA. That’s either a very bad sign or a sign that Google is rethinking what “Pro” even means in a world where small models keep getting scarily good.
💡 Key Takeaways
- Google deprecated Gemini 3.5 Flash after just two months, replacing it with a 3.6 version that scores 12 percentage points higher on the DeepSWE coding benchmark.
- The new Gemini 3.6 Flash uses 17 percent fewer tokens and has a lower API price for output, directly addressing enterprise concerns about runaway AI costs.
- Despite the Flash lineup expanding to include a Lite and a cybersecurity variant, the still-missing Gemini 3.5 Pro suggests Google's flagship model strategy is in disarray.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.