Google split its Gemini Flash line into three distinct models on July 21: Gemini 3.6 Flash for general coding and knowledge work, 3.5 Flash-Lite for high-volume low-latency tasks, and a cyber-specialized 3.5 Flash variant built into CodeMender, its automated vulnerability-patching agent. The split matters because Google is no longer selling “cheap and fast” as one product. It now sells three, each tuned for a different constraint agent builders actually hit in production.
3.6 Flash is the workhorse replacement for 3.5 Flash, and Google frames the upgrade around token efficiency rather than raw capability. According to the Artificial Analysis Index, 3.6 Flash uses 17 percent fewer output tokens than its predecessor on equivalent tasks, and Google cites reductions as steep as 65 percent on DeepSWE, a coding benchmark built by Datacurve. Pricing moved down at the same time, to $1.50 per million input tokens and $7.50 per million output tokens, so the cost-per-task drop compounds with the token savings rather than offsetting them.
Google’s own comparisons show 3.6 Flash ahead of 3.5 Flash on four measures:
- Coding precision on DeepSWE: 49 percent versus 37 percent
- ML research tasks on MLE Bench: 63.9 percent versus 49.7 percent
- Computer-use accuracy on OSWorld-Verified: 83.0 percent versus 78.4 percent
- Knowledge work on GDPval-AA v2: 1,421 versus 1,349
Google said customers Hebbia and Harvey are already using 3.6 Flash for document parsing and report drafting. None of the figures above come from an independent evaluator. The announcement, published on Google’s company blog, includes no third-party benchmark results for any of the three models.
3.5 Flash-Lite targets the opposite end of the tradeoff: throughput. Google says it runs at 350 output tokens per second, the fastest model in the 3.5 family, priced at $0.30 per million input tokens and $2.50 per million output tokens. It beats the prior Flash-Lite generation on Terminal-Bench 2.1, 54 percent to 31 percent, and on long-context retrieval on GDM-MRCR v2, 72.2 percent to 60.1 percent. Google also claims Flash-Lite outright beats the larger 3 Flash model on SWE-Bench Pro (54.2 percent versus 49.6 percent) and OSWorld-Verified (74.0 percent versus 65.1 percent).
That last comparison is the one worth sitting with. A Flash-Lite model beating a full-size Flash model from the same company on agentic and coding evals squeezes the market for mid-tier models built purely on cost. Anthropic’s Haiku and OpenAI’s mini-tier models now compete against a $2.50-per-million-output-token product that Google says clears frontier-adjacent scores on some agentic benchmarks. Cheap no longer means slow or shallow, at least by Google’s own measurement.
The third model, 3.5 Flash Cyber, is narrower by design. Google built it on top of 3.5 Flash and fine-tuned it specifically to find and patch code vulnerabilities inside CodeMender. Multiple Flash Cyber instances run in parallel inside CodeMender and merge their findings into one combined report, and Google says the setup reaches competitive frontier-level performance on CyberGym, a benchmark for automated vulnerability discovery, at a lower per-token cost than running a full frontier model for the same job.
Google is restricting 3.5 Flash Cyber to governments and what it calls trusted partners, through a limited-access CodeMender pilot, citing the dual-use risk of a model built to find security holes. The company’s stated goal is giving defenders a head start over attackers already probing code with general-purpose models. Whether a gated, narrow cyber model actually shifts that balance depends on adoption numbers and access terms Google has not published.
Google also confirmed, briefly, that Gemini 3.5 Pro is in partner testing ahead of a wider release and that pre-training has started on Gemini 4, without elaborating on either. For now, the three Flash models are what shipped: 3.6 Flash and 3.5 Flash-Lite are live today in the Gemini API, Google AI Studio, Android Studio, the Gemini Enterprise platform, and the Gemini app, with Flash-Lite also rolling into Google Search. Teams currently pricing cheap-tier model contracts for the second half of 2026 should re-run their cost-per-task math against Flash-Lite’s published numbers before they renew.
Announced by Google on July 21, 2026, in a post by product lead Tulsee Doshi on the company’s official blog.