The rapid release cycle continues at Google. On Wednesday, the company unveiled another pair of Gemini Flash and Flash Cyber variants, marking the third Flash launch in just six weeks. Despite Gemini 3.7 Flash being only three weeks old, the newest iteration demonstrates substantial improvements, particularly in agentic coding and computer use scenarios.

Gemini 3.8 Flash carries an introductory pricing structure of $0.75 and $3.75 per million input and output tokens respectively. This promotional rate remains in effect until December 31, 2026, after which pricing will increase to $1.50 and $7.50 per million tokens.

Fairwind gates Flash 3.8 Cyber

While Gemini 3.8 Flash positions itself as a general-purpose workhorse model, the Cyber variant warrants separate attention. Google is adopting a strategy similar to Anthropic's approach with specialized models. The company characterizes Flash 3.8 Cyber as its most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, and has restricted access accordingly.

Distribution of Flash 3.8 Cyber is limited to approximately 650 trusted partners through Google's newly named Fairwind Program. Participating organizations include Accenture, CrowdStrike, the Center for Internet Security, Datadog, Palo Alto Networks, Snowflake, and Wiz.

Four Flynn, Google's VP for Security and Privacy, explained the rationale: The defender's edge comes from shrinking the time between detecting a flaw and patching it. Through Google's Fairwind Program, government and enterprise partners gain autonomous tools to repair systems faster and at scale, keeping them one step ahead of agentic-speed threats.

This represents a formalization of earlier efforts. Flash 3.5 Cyber, released in July, had launched with a limited-access program that lacked formal branding and appeared more ad hoc in its implementation.

Gemini 3.8 Flash: long-horizon coding and agents

Google's benchmarking indicates that Gemini 3.8 Flash frequently surpasses competing models from OpenAI and Anthropic on complex engineering assignments. On DeepSWE, for instance, it matches Opus 5 while outperforming GPT-5.6 Sol and Claude Sonnet 5.

However, a significant caveat applies. Google acknowledges that 3.8 Flash works harder, which translates to increased token consumption. The company elaborates: On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels. Developers can mitigate this through configurable reasoning modes.

Google's evaluation notably excludes Anthropic's Fable 5.1, released just one day prior. While Fable 5.1 represents a substantially stronger model, it commands significantly higher costs. On Terminal-Bench 4.0, a general agentic benchmark, Fable 5.1 achieves 55.8% compared to Flash 3.8's 19.1% and Opus 5's 51.8%. Conversely, on Terminal-Bench 2.1, which emphasizes coding, Flash 3.8 demonstrates superior performance.

Credit: Google.

Performance gaps persist in specific domains. On computer use tasks, Flash 3.8 reaches 59% versus Opus 5's 75.4% on OSWorld-2.0. On GDPVal, which assesses knowledge work capabilities, Flash 3.8 scores 1545, trailing Opus 5's 1824 but approaching Sonnet 5's 1584.

Google attributes its accelerated improvement timeline to a new methodology. The company states: Both of today's releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models.

Chinese models close the gap

Google's comparative analysis focuses exclusively on OpenAI and Anthropic offerings. However, Chinese alternatives merit consideration. On benchmarks like DeepSWE 1.1, models including GLM-5.3, GLM-5.3 Flash, DeepSeek v4 Pro, and Kimi K3 demonstrate competitive positioning, frequently offering superior price-to-performance ratios.

Cyber benchmarks

Credit: Google.

Flash 3.8 Cyber demonstrates what Google characterizes as frontier-level autonomous vulnerability discovery capabilities on CyberGym, outperforming both the July-released Flash 3.5 Cyber and significantly larger frontier models.

Since CyberGym covers only C and C++ code, Google conducted additional testing using an internal benchmark spanning 20 programming languages, reporting success rates exceeding 70%. On Collinear's CWE-Bench, the model achieves a pass@1 score of 47.2%, marginally behind an unnamed leading frontier model at 47.8%.

Real-world validation supplements benchmark results. Chrome's security team reports the model generates 2.6 times more correct patches than larger commercial alternatives. Wiz documents 7.5% to 9.7% higher recall on internal penetration testing benchmarks while operating at 2.3 to 5.2 times lower cost.

Safety

Gemini 3.8 Flash incorporates safeguards targeting Chemical, Biological, Radiological, and Nuclear (CBRN) domains and cyber offense, while preserving beneficial applications consistent with Google's Frontier Safety Framework.

Flash 3.8 Cyber operates under more permissive safety constraints, justified by its restricted user base. The model demonstrates enhanced robustness against prompt injection attacks, achieving near-leading performance on the Gray Swan IPI benchmark.

Gemini Pro?

The timing of Google's next Gemini Pro release remains uncertain. Flash models are advancing at a rapid cadence, and while Google's Pro launch strategy has encountered some missteps, the eventual release may justify the extended development period. By prioritizing Flash models, Google has constructed a compelling price-to-performance narrative that competitors have struggled to match.

At the current pace, a Gemini 3.9 Flash release could precede Gemini 4 Pro's arrival.

Availability

Gemini 3.8 Flash is accessible through Google's standard platforms including Antigravity, Google AI Studio, Android Studio, and Stitch, alongside Gemini Enterprise. Consumers holding AI Pro and Ultra subscriptions can access the model via the Gemini App, AI Mode in Google Search, and Gemini integration in Google Sheets.

Source: The New Stack