Google unveils Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber
Lower token costs and faster, more reliable agentic workflows could meaningfully cut operating costs for AI deployments and enable larger-scale automation.
At a glance
- 3.6 Flash consumes 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index.
- Price: 1.50/1M input tokens and 7.50/1M output tokens.
- 3.5 Flash-Lite runs at 350 output tokens per second.
- 3.6 Flash API rollout starts today and will replace 3.5 Flash in the Gemini app.
The story
Google announced Gemini 3.6 Flash, along with 3.5 Flash-Lite and 3.5 Flash Cyber, in a post describing the new models as delivering improved efficiency, lower latency, and more reliable performance for production AI agents.
The blog notes that 3.6 Flash builds on feedback from 3.5 Flash and delivers a step up in coding and knowledge work while markedly improving token efficiency. It cites a 17% reduction in output tokens on the Artificial Analysis Index and mentions fewer reasoning steps and tool calls for multi-step workflows, in addition to a lower cost per task (1.50/1M input tokens and 7.50/1M output tokens).
Google also introduces 3.5 Flash-Lite, described as the fastest model in the 3.5 series, capable of 350 output tokens per second. Pricing for Flash-Lite is 0.30/1M input tokens and 2.50/1M output tokens, with claimed improvements in throughput and cost efficiency for high-volume developer workloads. The post highlights benchmarks where 3.5 Flash-Lite outperforms earlier Flash iterations on several tasks and notes built-in capabilities like computer use as a tool for agentic tasks.
Other notes in the release include that 3.6 Flash will begin rolling out in the API today and will take over from 3.5 Flash in the Gemini app. The company also released 3.5 Flash-Lite and 3.5 Flash Cyber, with Cyber described as a limited pilot for cybersecurity applications, and a fuller 3.5 Pro in testing with partners and planned for broader availability when ready. Google also mentions that it has started a pre-training run for Gemini 4 and has added enhanced Frontier Safety safeguards for CBRN and cyber offense misuses to 3.6 Flash.
Ars Technica’s coverage adds a performance datapoint from the DeepSWE coding test, noting 3.6 Flash achieves 49% vs 37% for 3.5 Flash in that benchmark, and that 3.6 Flash uses fewer tokens overall while improving coding performance. The report also notes 3.5 Flash Lite’s deployment in the ecosystem and that 3.5 Pro remains in testing with partners, with a stated goal to release “as soon as it’s ready.”