Google released Gemini 3.8 Flash on September 2, its third Flash-tier model in six weeks, alongside a specialized variant called Gemini 3.8 Flash Cyber that finds and patches vulnerabilities for a restricted list of defenders. Both share the same underlying training, according to Google, which credits gains in the general model partly to lessons pulled from the harder cybersecurity work. The pairing signals that Google now treats offensive-grade capability testing as an input to its mainstream coding models, not a separate track.
The standard 3.8 Flash keeps 3.7 Flash’s introductory pricing: $0.75 per million input tokens and $3.75 per million output tokens. Google says it closes ground on costlier frontier models across software engineering and multi-step reasoning, citing DeepSWE v1.1 for long-horizon coding tasks, Vals Finance Agent V2, and Harvey’s Legal Agent Benchmark, plus a 54.9 percent score on HLE-Verified. Every one of those figures comes from Google’s own evaluation, run on its own model, published in its own launch post. None of it has been checked by an outside lab.
Google attributes the jump to the model simply working harder: taking more reasoning steps, calling tools iteratively, and burning more tokens at higher effort settings. That is a real tradeoff, not a free upgrade. Teams optimizing for token cost can dial effort down or stay on 3.7 Flash, which Google says remains fully supported.
Flash Cyber is the more consequential release. It ships only through a new access program called Fairwind, whose intended audience is government cyber agencies, the people who run critical infrastructure, and those who maintain widely used software, rather than developers at large. Google reports frontier-level results on CyberGym, the standard external benchmark for autonomous vulnerability discovery, and says the model beats its own prior Cyber model along with larger general-purpose systems. On an internal test spanning 20 programming languages, Google puts the discovery success rate above 70 percent.
On patching, the side Google says it prioritized over offensive capability, Flash Cyber posted a 47.2 percent pass rate on CWE-Bench, a third-party benchmark from Collinear, against 47.8 percent for a larger frontier model Google does not name, at what it describes as a significantly lower cost. Google says it is already running the model internally: Chrome’s security team reports 2.6 times more correct patches than larger commercial models, Wiz reports higher recall at lower cost on its own penetration-testing benchmark, and a Google cloud vulnerability team says it found a critical bug in under two hours on work that typically takes months. These are Google’s characterizations of its own deployments, not independent audits.
Google frames the gating as deliberate: 3.8 Flash Cyber carries looser safety mitigations than the general release specifically because it needs broader offensive-adjacent capability to be useful for defense, and that tradeoff is why access runs through an application process rather than an API key. The move lands one day after OpenAI said its Astra model had crossed the company’s own “Critical” threshold for cyber capability. Two labs reaching similar capability claims within a day of each other, each responding with its own gated program rather than a shared standard, suggests the restricted defender-access model is becoming the default release lane for security-capable AI, with each vendor setting its own bar for who gets in.
Security teams evaluating either model should ask for evidence beyond Google’s benchmarks before switching workflows, and any organization eligible for Fairwind should weigh applying now: first-wave access to a gated program tends to shape how its evaluation criteria get set.
Google detailed both releases in a September 2, 2026 post on its official blog.