Gemini 3.8 Flash

Gemini 3.8 Flash Intro Price Doubles on Jan 1: 54.9% HLE

Gemini 3.8 Flash intro pricing doubles on Jan 1, 2027; Flash Cyber scores 47.2% on CWE-Bench and found a critical vulnerability in under two hours.

Gemini 3.8 Flash Intro Price Doubles on Jan 1: 54.9% HLE — article cover

On September 2, 2026, Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. Coming just three weeks after 3.7 Flash (August 13), Google itself frames this as its “third Flash model release in six weeks.” The pattern is clear: Pro-model updates have slowed while Flash iterates fast — Google is concentrating its energy on the fast, cheap, smart workhorse tier. The accompanying DeepMind model card lists the knowledge cutoff as March 2026.

The two variants share one core intelligence. The official post, credited to Tulsee Doshi and Raluca Ada Popa, says this generation’s gains come from “long-running agentic loops designed to recursively evaluate and refine the underlying models,” with part of the improvement driven by “rigorous training in the highly demanding domain of cybersecurity.”

3.8 Flash Benchmark Results: A Workhorse That Trades Punches With Frontier Models

Google positions 3.8 Flash as its “most intelligent workhorse model,” improving on 3.7 Flash in software engineering, agentic tasks, and multi-step specialized reasoning at the same speed and cost. The numbers:

  • Outperforms most larger frontier models on DeepSWE v1.1, a long-horizon software engineering benchmark
  • Beats 3.7 Flash and other frontier models on Vals Finance Agent V2 and Harvey’s Legal Agent Benchmark
  • Scores 54.9% on HLE-Verified

What matters here is the cost structure: frontier-adjacent scores at Flash-tier pricing. For teams running agentic workloads at volume, that beats a single peak benchmark.

Pricing and the “Works Harder” Tradeoff

3.8 Flash ships with introductory pricing: $0.75 per million input tokens and $3.75 per million output tokens, expiring December 31, 2026. On January 1, 2027 it rises to $1.50 and $7.50. Teams that want to lock in costs should run the math now.

Google is also candid about an engineering tradeoff: 3.8 Flash “works harder” — at high effort it takes more reasoning steps and tool calls, so token usage climbs. For efficiency-first workloads, Google recommends lower effort settings or staying on 3.7 Flash. It is rare for a lab to admit outright that a new model is not universally cheaper to run.

Flash Cyber: Defense as Its Own Model

Flash Cyber is Google’s strongest cybersecurity model, claiming “frontier-level performance in vulnerability detection and automated patching,” designed for defense rather than offense. The key results:

  • Surpasses 3.5 Flash Cyber and larger frontier models on CyberGym
  • Exceeds a 70% success rate on an internal 20-programming-language benchmark
  • Scores 47.2% pass@1 on CWE-Bench (Collinear), nearly matching a leading frontier model’s 47.8% at significantly lower cost
  • Chrome Security measured 2.6x more correct patches than larger commercial models
  • Wiz reported 7.5–9.7 percentage points higher recall at 2.3–5.2x lower cost
  • Google’s own Cloud Vulnerability Research team found a critical foundational vulnerability in under 2 hours, where such finds typically take months

Because Flash Cyber ships with more permissive cyber-safety mitigations, it is not open to general developers. Access is prioritized through the Fairwind Program for trusted government authorities, critical infrastructure operators, and software maintainers.

Availability

3.8 Flash reaches a wide surface: developers get it in Google Antigravity, the Gemini API (AI Studio and Android Studio), and Stitch; enterprises through Gemini Enterprise; consumers via AI Pro/Ultra subscriptions in the Gemini app, AI Mode in Search, and Gemini in Sheets. Safeguards follow Google’s Frontier Safety Framework, with improved prompt-injection robustness per Gray Swan testing.

The signal for product teams is clear: Flash is now Google’s main iteration vehicle, and cyber defense has become a standalone model product line rather than a feature of the general-purpose model.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL