Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Gemini 3.6 Flash:
A workhorse model designed for improved coding, knowledge work, and multimodal performance with greater efficiency than 3.5 Flash.
- Efficiency and Pricing: Consumes 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on the DeepSWE benchmark. It is priced lower than 3.5 Flash at $1.50/1M input tokens and $7.50/1M output tokens.
- Features: Includes a built-in client-side computer use tool available via the Gemini API and Gemini Enterprise. It also ships with enhanced Frontier Safety safeguards for Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense misuse, making it more resistant to jailbreaks while minimizing refusals for beneficial uses.
| Benchmark | 3.6 Flash | 3.5 Flash |
|---|---|---|
| DeepSWE | 49% | 37% |
| MLE Bench (ML Research) | 63.9% | 49.7% |
| OSWorld-Verified (Computer Use) | 83.0% | 78.4% |
| GDPval-AA v2 (Knowledge Work) | 1421 | 1349 |
Gemini 3.5 Flash-Lite:
The fastest and most cost-effective model in the 3.5 series, designed for low-latency and high-throughput tasks like agentic search and document processing.
- Speed and Pricing: Achieves 350 output tokens per second according to Artificial Analysis. It is priced at $0.30/1M input tokens and $2.50/1M output tokens.
- Features: Can be configured with different "thinking levels" (minimal, low, higher) to balance latency and cost against task complexity. It also includes a built-in computer use tool.
- Performance: Outperforms 3.1 Flash-Lite significantly and, on some agentic and coding evals, outperforms the larger 3 Flash model.
| Benchmark | 3.5 Flash-Lite | Baseline Model & Score |
|---|---|---|
| Terminal-Bench 2.1 | 54% | 3.1 Flash-Lite (31%) |
| GDM-MRCR v2 (Long Context) | 72.2% | 3.1 Flash-Lite (60.1%) |
| GDPval-AA v2 | 1140 | 3.1 Flash-Lite (642) |
| SWE-Bench Pro | 54.2% | 3 Flash (49.6%) |
| OSWorld-Verified | 74.0% | 3 Flash (65.1%) |
Gemini 3.5 Flash Cyber:
A specialized model for finding and fixing cybersecurity vulnerabilities, deployed within the CodeMender agent system.
- Architecture: Fine-tuned from Gemini 3.5 Flash for cybersecurity tasks.
- Deployment: Used within CodeMender, where multiple 3.5 Flash Cyber agents collaborate to produce a single report. It achieves competitive performance on the CyberGym benchmark.
- Availability: Due to its dual-use nature, the model is available exclusively to governments and trusted partners via CodeMender in a limited-access pilot program.
Availability:
- 3.6 Flash & 3.5 Flash-Lite: Available in the Gemini API (via Google AI Studio and Android Studio) and Gemini Enterprise Agent Platform.
- 3.6 Flash: Also available in Google Antigravity and the Gemini Enterprise app.
- 3.5 Flash-Lite: Also rolling out in the Gemini app and Google Search.
This release signals a clear strategy from Google to compete on the practicalities of production AI, not just on headline benchmark scores. By segmenting their Flash series, they are targeting the primary bottleneck for scaling AI agents: operational cost. The focus on token efficiency, lower pricing, and specialized variants for high-throughput and security tasks is a direct appeal to developers building real-world, cost-sensitive applications. The restricted release of the 'Cyber' model also reflects a growing, necessary caution in the industry regarding the dual-use potential of advanced AI capabilities, pre-emptively limiting access to tools that could be misused for offensive purposes. This is less about a single capability leap and more about maturing the AI toolkit for widespread, economical deployment.
Don't read this site daily. Get it in your inbox.
The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.