Google's Gemini 3.6 Flash slashes output tokens by up to 65%, cutting agent costs
Curated by the Inblix editorial team
Google dropped a trio of new Gemini Flash models today, and the through-line is clear: the company is optimizing aggressively for the economics of agentic work, not just raw benchmark scores. The headliner is Gemini 3.6 Flash, a direct upgrade to the workhorse 3.5 Flash that Google says uses up to 65% fewer output tokens on the DeepSWE coding benchmark. That verbosity cut, combined with a trimmed output price of $7.50 per million tokens (down from $9.00), meaningfully lowers the bill for high-volume developers. Quality didn’t take a backseat—on DeepSWE, 3.6 Flash hits 49% versus the previous 37%, and MLE Bench scores jump from 49.7% to 63.9%. Early customers like Hebbia and Harvey are already citing gains in document parsing and report drafting.
Then there’s Gemini 3.5 Flash-Lite, a speed demon clocking 350 output tokens per second. Priced at a minuscule $0.30 per million input tokens, it’s built for low-latency search and document processing slogs. What’s surprising is it doesn’t just beat the older 3.1 Flash-Lite—it outmuscles the original 3.5 Flash on some evals, scoring 54.2% on SWE-Bench Pro versus 49.6%. Google is exposing configurable thinking levels here, letting developers trade smarts for speed on a per-task basis. It’s a pragmatic nod to the fact that not every subagent needs a PhD.
The most specialized release is Gemini 3.5 Flash Cyber, fine-tuned for automated vulnerability discovery and packaged inside a new tool called CodeMender. The design philosophy is a bet against giant, monolithic reasoning models for security work. Instead of one expensive call, CodeMender launches up to five parallel Flash Cyber agents to explore a massive execution search space, merging their findings into a single report. The internal numbers are genuinely eye-catching: on Google’s Big Sleep evaluation, it trounced both mainline 3.5 and 3.6 Flash, and on the V8 JavaScript engine it found 55 unique confirmed bugs at a fixed invocation count—19 more than Claude Opus 4.6.
Community reaction was split right down the middle. Builders cheered the price-to-performance ratio, but frustration boiled over on Hacker News about Google’s ability to provision the promised capacity, with some users recounting painful hands-on coding sessions. The gated release of Flash Cyber also reignited the dual-use debate: who should have access to automated exploit-finding tools that can surface remote-code-execution flaws in public APIs within two hours? It’s a valid question Google will have to answer as these specialized agents grow more capable. Gemini 3.6 Flash and 3.5 Flash-Lite are available now via Google AI Studio and Android Studio, with 3.6 Flash also rolling out in Google Antigravity and GitHub.
💡 Key Takeaways
- Gemini 3.6 Flash achieves up to a 65% reduction in output tokens on the DeepSWE benchmark, directly lowering per-task costs for agentic workflows.
- The dirt-cheap Gemini 3.5 Flash-Lite ($0.30/M input tokens) can outscore the original 3.5 Flash on certain benchmarks like SWE-Bench Pro, making it a compelling option for high-throughput tasks.
- Gemini 3.5 Flash Cyber, used inside the CodeMender agent, found 19 more confirmed vulnerabilities in the V8 JavaScript engine than Claude Opus 4.6, proving cheap parallelized models can beat massive ones for specific security tasks.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.