AI Intelligence // signal over noise
← back to feed
DeepMind 8/10 signal

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

modelsagentic
Summary
Google DeepMind has released three new models in its efficiency-focused Flash series: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The new models are designed for building AI agents at scale, with 3.6 Flash reducing output token usage by 17% compared to its predecessor while improving performance on coding and knowledge tasks.
Context
As developers and enterprises increasingly build production AI agents, the demand for higher token efficiency, lower latency, and reliable performance at scale has grown. This release builds on the Gemini 3.5 Flash model, which was designed to hit a sweet spot between efficiency and quality. These new models represent a further segmentation of the Flash series to target specific agentic workflows, from general-purpose coding to high-throughput data processing and specialized cybersecurity tasks. This launch comes as Google continues to develop its broader model family, with Gemini 3.5 Pro in partner testing and pre-training for Gemini 4 already underway.
Details

Gemini 3.6 Flash:

A workhorse model designed for improved coding, knowledge work, and multimodal performance with greater efficiency than 3.5 Flash.

  • Efficiency and Pricing: Consumes 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on the DeepSWE benchmark. It is priced lower than 3.5 Flash at $1.50/1M input tokens and $7.50/1M output tokens.
  • Features: Includes a built-in client-side computer use tool available via the Gemini API and Gemini Enterprise. It also ships with enhanced Frontier Safety safeguards for Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense misuse, making it more resistant to jailbreaks while minimizing refusals for beneficial uses.
Benchmark3.6 Flash3.5 Flash
DeepSWE49%37%
MLE Bench (ML Research)63.9%49.7%
OSWorld-Verified (Computer Use)83.0%78.4%
GDPval-AA v2 (Knowledge Work)14211349

Gemini 3.5 Flash-Lite:

The fastest and most cost-effective model in the 3.5 series, designed for low-latency and high-throughput tasks like agentic search and document processing.

  • Speed and Pricing: Achieves 350 output tokens per second according to Artificial Analysis. It is priced at $0.30/1M input tokens and $2.50/1M output tokens.
  • Features: Can be configured with different "thinking levels" (minimal, low, higher) to balance latency and cost against task complexity. It also includes a built-in computer use tool.
  • Performance: Outperforms 3.1 Flash-Lite significantly and, on some agentic and coding evals, outperforms the larger 3 Flash model.
Benchmark3.5 Flash-LiteBaseline Model & Score
Terminal-Bench 2.154%3.1 Flash-Lite (31%)
GDM-MRCR v2 (Long Context)72.2%3.1 Flash-Lite (60.1%)
GDPval-AA v211403.1 Flash-Lite (642)
SWE-Bench Pro54.2%3 Flash (49.6%)
OSWorld-Verified74.0%3 Flash (65.1%)

Gemini 3.5 Flash Cyber:

A specialized model for finding and fixing cybersecurity vulnerabilities, deployed within the CodeMender agent system.

  • Architecture: Fine-tuned from Gemini 3.5 Flash for cybersecurity tasks.
  • Deployment: Used within CodeMender, where multiple 3.5 Flash Cyber agents collaborate to produce a single report. It achieves competitive performance on the CyberGym benchmark.
  • Availability: Due to its dual-use nature, the model is available exclusively to governments and trusted partners via CodeMender in a limited-access pilot program.

Availability:

  • 3.6 Flash & 3.5 Flash-Lite: Available in the Gemini API (via Google AI Studio and Android Studio) and Gemini Enterprise Agent Platform.
  • 3.6 Flash: Also available in Google Antigravity and the Gemini Enterprise app.
  • 3.5 Flash-Lite: Also rolling out in the Gemini app and Google Search.
What's new
The release of three distinct, efficiency-focused models tailored for specific agentic workflows. This represents a strategic segmentation of the Gemini Flash series, moving beyond a single general-purpose efficient model to offer a workhorse (3.6 Flash), a high-speed/low-cost option (3.5 Flash-Lite), and a highly specialized, restricted-access security model (3.5 Flash Cyber). The core novelty is the simultaneous improvement in performance and token efficiency at a lower cost point for production agentic systems.
Limitations
Gemini 3.5 Flash Cyber is not being released publicly. Due to its dual-use nature for finding and fixing vulnerabilities, it will be available only to governments and trusted partners through a limited-access pilot program within the CodeMender agent.
The take

This release signals a clear strategy from Google to compete on the practicalities of production AI, not just on headline benchmark scores. By segmenting their Flash series, they are targeting the primary bottleneck for scaling AI agents: operational cost. The focus on token efficiency, lower pricing, and specialized variants for high-throughput and security tasks is a direct appeal to developers building real-world, cost-sensitive applications. The restricted release of the 'Cyber' model also reflects a growing, necessary caution in the industry regarding the dual-use potential of advanced AI capabilities, pre-emptively limiting access to tools that could be misused for offensive purposes. This is less about a single capability leap and more about maturing the AI toolkit for widespread, economical deployment.

Don't read this site daily. Get it in your inbox.

The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.