Three Flash releases in six weeks. Google is moving fast, and Gemini 3.8 Flash is the clearest sign yet that the company is serious about owning the mid-tier model segment. According to Google, 3.8 Flash delivers meaningful gains over 3.7 Flash in software engineering, multi-step reasoning, and agentic tasks, all at the same price point. That combination is hard to ignore.
What’s actually new with 3.8 Flash
The headline claim is performance that approaches larger, more expensive frontier models at a fraction of the cost. On DeepSWE v1.1, a long-horizon software engineering benchmark, 3.8 Flash reportedly outperforms most larger frontier models in end-to-end autonomous problem solving. That’s a direct shot at OpenAI’s o3 and Anthropic’s Claude Sonnet class of models, which carry significantly higher token costs.
The model also scores 54.9% on HLE-Verified, a demanding benchmark spanning STEM, humanities, and professional fields. For finance and legal workflows specifically, Google says 3.8 Flash beats both 3.7 Flash and competing frontier models on Vals Finance Agent V2 and Harvey’s Legal Agent Benchmark. Those are niche but meaningful signals for enterprise buyers.
The underlying design choice here is straightforward: 3.8 Flash works harder on complex tasks. It runs extra reasoning steps and calls tools iteratively, which can mean higher token usage at elevated effort levels. Google is giving developers a dial to manage this, with lower effort settings available when compute efficiency matters more than maximum performance. Teams that need to keep costs tight can also stay on 3.7 Flash, which remains supported.
Pricing stays at $0.75 per million input tokens and $3.75 per million output tokens, matching 3.7 Flash. For context, that’s well below GPT-4o and Claude 3.5 Sonnet at their standard rates, and it makes 3.8 Flash genuinely competitive for high-volume agentic deployments.
Gemini 3.8 Flash Cyber: a restricted but significant release
The second model, Gemini 3.8 Flash Cyber, is more unusual. It’s not a general-release product. It’s aimed at what Google is calling the Fairwind Program, a gated access path for government authorities, critical infrastructure operators, and software maintainers who need advanced cybersecurity capabilities.
The performance numbers here are notable. On CyberGym, the standard benchmark for autonomous vulnerability discovery, 3.8 Flash Cyber outperforms both Google’s own 3.5 Flash Cyber and larger frontier competitors. On an internal benchmark covering 20 programming languages across complex codebases, it hits a vulnerability discovery success rate above 70%.
On patching, the picture is similarly strong. CWE-Bench, run by Collinear, shows 3.8 Flash Cyber at 47.2% pass@1 versus a leading frontier model at 47.8%, but at significantly lower cost. That’s competitive parity with better economics.
Real-world results back this up:
- Google’s Chrome Security team found 3.8 Flash Cyber produced 2.6 times more correct vulnerability patches than leading commercial alternatives
- Wiz reported 7.5 to 9.7% higher recall on internal penetration testing benchmarks, at 2.3 to 5.2 times lower cost than other frontier models
- Google’s Cloud Vulnerability Research team used it to find a critical foundational vulnerability in under two hours, work that typically takes months
Safety guardrails and who gets access
Both models ship with safeguards against misuse in Chemical, Biological, Radiological, and Nuclear domains. 3.8 Flash Cyber has looser cybersecurity mitigations by design, which is precisely why access is restricted to vetted defenders. Google is also citing improvements in prompt injection robustness across both models, measured by Gray Swan, which matters for anyone deploying these in agentic pipelines where adversarial inputs are a real risk.
Where you can use it today
3.8 Flash is available now through the Gemini API, Google AI Studio, Android Studio, and Gemini Enterprise. Consumers on Google AI Pro and Ultra plans get access through the Gemini app, AI Mode in Search, and Gemini in Google Sheets. Cyber access requires applying through the Fairwind Program directly.
The bigger picture is this: Google is compressing the performance gap between budget and frontier models faster than most expected. If 3.8 Flash holds up in production, it puts real pressure on Anthropic and OpenAI to justify the cost premium on their comparable offerings.




