Just three weeks after Gemini 3.7 Flash, Google has thrown its third Flash card in six weeks: Gemini 3.8 Flash — billed by the company as its "best reasoning & coding model yet" at 3.7's speed and price. The same announcement hides a second name: Gemini 3.8 Flash Cyber, a cybersecurity-specialized model available only to "trusted defenders."
Two models, one brain. One goes to work; one stands guard.

One brain, two jobs: the strongest worker and the fastest gatekeeper
The announcement makes the product logic unapologetically clear: 3.8 Flash is the "most intelligent workhorse," built for software engineering, agentic tasks and multi-step reasoning in specialized domains; 3.8 Flash Cyber is the "most capable defensive model," specialized in vulnerability discovery and automated patching, offered to trusted defenders through the new Fairwind Program.
Notably, the two share the same foundational intelligence, accelerated by long-running agentic loops that recursively evaluate and refine the underlying models. Google states plainly that the coding and reasoning gains rest substantially on rigorous training in the highly demanding domain of cybersecurity: drill on the toughest battlefield, then bring the muscle back to the everyday workstation.
The worker: the upper-right corner of the price-performance curve
DeepSWE v1.1 (long-horizon software engineering) is the headline chart: 3.8 Flash outperforms most larger, more expensive frontier models in autonomously solving complex engineering problems end to end, at a fraction of the cost — which is exactly what the cover chart shows: roughly 72% success at about $1.50 per task, while models eight times pricier are not visibly more accurate.
Specialized domains hold up too: on benchmarks demanding advanced analysis and reporting — Vals Finance Agent V2 and Harvey's Legal Agent Benchmark — 3.8 Flash beats 3.7 Flash and other frontier models; it scores 54.9% on HLE-Verified, covering multi-step reasoning across STEM, humanities and professional fields.
The demos show the playful side of this brain: in Google Antigravity, 3.8 Flash built a wizard-castle 3D game with puzzles and environmental storytelling from a single looping-instruction prompt (textures by Nano Banana); it recreated a fully playable DOS version of Google Maps in one prompt — locations, directions and Street View included — and built an interactive topographic cross-section explorer from real U.S. Geological Survey datasets.
How it earns it: working harder, not growing bigger
The core design choice is disarmingly plain: "3.8 Flash works harder." On complex tasks it executes extra reasoning steps and calls tools iteratively, sometimes spending more tokens to maximize performance, especially at higher effort levels.
This is a candid capability disclosure — and pricing as an art: Google simultaneously tells developers that if compute efficiency is the primary constraint, they can dial effort levels down to cut token overhead, or keep using the fully supported 3.7 Flash. The two generations are explicitly split into "performance-first" and "efficiency-first."
The gatekeeper: a physician specialized in vulnerabilities
The Cyber variant's report card has three parts.
First, autonomous vulnerability discovery: frontier-level performance on the industry-standard CyberGym benchmark, surpassing both 3.5 Flash Cyber and significantly larger frontier models. On a comprehensive internal benchmark spanning 20 programming languages across complex codebases, it exceeds a 70% success rate — an "impressive leap" over predecessors.
Second, automated patching: on CWE-Bench, run by Collinear, 3.8 Flash Cyber posts a pass@1 of 47.2% against a leading frontier model's 47.8% — under one percentage point apart at significantly lower cost, squarely on the Pareto frontier. Google stresses a stance here: investing in fixing from the start, prioritized over offensive capabilities like exploitation.
Third, the most persuasive part — Google is already using it: the Chrome Security team found it produced 2.6 times more correct patches for Chrome vulnerabilities than the best, much larger commercial models; on Wiz's internal penetration testing benchmark it achieves 7.5–9.7% higher recall at 2.3–5.2x lower cost; and Google's Cloud Vulnerability Research team used it to find a critical foundational vulnerability in under 2 hours — research and discovery that usually takes months.
HLE-Verified
Value: 54.9% Context: multi-step reasoning across STEM, humanities and professional fields
Internal vulnerability benchmark (20 languages)
Value: >70% Context: 3.8 Flash Cyber autonomous discovery success rate
CWE-Bench pass@1
Value: 47.2% Context: vs. a leading frontier model's 47.8%, at significantly lower cost
Chrome patch output
Value: 2.6x Context: correct patches vs. larger best commercial models
Wiz pen-test recall
Value: +7.5–9.7% Context: at 2.3–5.2x lower cost
Introductory price (input/output, per million tokens)
Value: $0.75 / $3.75 Context: reverts to $1.50 / $7.50 from Jan 1, 2027
Safety design: the standard model tightens, the Cyber variant loosens — selectively
Per the Frontier Safety Framework, 3.8 Flash ships with safeguards against misuse in Chemical, Biological, Radiological and Nuclear (CBRN) domains and cyber offense; the Cyber variant deliberately ships more permissive cybersecurity mitigations — the price being availability only to trusted defenders who require a fuller set of cyber capabilities. The structure is worth pondering: capability is no longer gated only by price, but by trust. Prompt-injection robustness as measured by Gray Swan also takes a significant leap forward.
Pricing and availability
The introductory price is $0.75 per million input tokens and $3.75 per million output tokens — identical to 3.7 Flash. Mind the footnote: the price expires December 31, 2026; from January 1, 2027 it reverts to $1.50 / $7.50. The half-price window conveniently covers a full procurement evaluation cycle.
Developers can build in Google Antigravity, the Gemini API (AI Studio / Android Studio) and Stitch; enterprises via Gemini Enterprise; consumers via Google AI Pro and Ultra subscriptions across the Gemini app, AI Mode in Search and Gemini in Google Sheets; and cyber defenders via the Fairwind Program — prioritized access for government authorities, critical infrastructure operators and software maintainers.
Google begins its dense Flash release cadence — 3.8 is officially the third Flash release in six weeks
About three weeks before 3.8; Google says 3.8 delivers significantly stronger capability at the same speed and price, while 3.7 remains fully supported for efficiency-first workloads
The "best reasoning & coding model yet" ships at the same price; the Cyber variant debuts alongside the Fairwind Program for trusted defenders
3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively.
Opinion: three things worth remembering
"Working harder" is candor — and a business. Trading more tokens for higher scores rewrites compute cost into an API bill; but effort levels and the continued availability of 3.7 leave the choice (and the bill) with developers. An honest cost model beats vague magic.
The trust threshold is becoming the new pricing dimension. The Cyber variant is cheap, stronger and more "empowered," yet not for sale to everyone. When capability allocation shifts from "who can pay" to "who is trusted," the industry needs to start discussing how trust itself is granted and audited — the Fairwind Program is the first public answer to that question.
Three releases in six weeks makes Flash a disposable-monthly commodity. When a model's lifecycle is measured in weeks, leaderboards stop asking "who is strongest" and start asking "who can iterate." 3.8 Flash welds price-performance onto its own corner of the curve; how rivals respond will shape the model market for the year ahead.