Google launches Gemini 3.6 Flash (17% fewer tokens), 3.5 Flash-Lite, and cyber-security model
Google released three new Gemini models on Tuesday as it prepares for earnings and faces intensifying competition from Chinese AI rivals. Gemini 3.6 Flash improves coding, multimodal, and knowledge-work performance while reducing output token usage by 17% compared to 3.5 Flash and lowering cost per output token to $1.50/M input, $7.50/M output. The model shows 49% accuracy on DeepSWE code editing (vs. 37% for 3.5 Flash) and 83.0% on OSWorld computer-use tasks (vs. 78.4% for prior generation).
Gemini 3.5 Flash-Lite targets high-volume, low-latency workloads at just $0.30/M input tokens and $2.50/M output tokens, running at 350 output tokens per second according to Artificial Analysis—the fastest in the 3.5 family. It significantly outperforms older Flash-Lite generations on agentic workflows and includes built-in computer-use capabilities. Gemini 3.5 Flash Cyber, launched in limited pilot for governments and trusted partners, is a specialized model paired with CodeMender, Google's code security agent, designed to detect and patch software vulnerabilities at lower cost per token than larger models.
Google's broader roadmap signals confidence: 3.5 Pro is in partner testing ahead of broader release, and the company has begun its largest pre-training run yet for Gemini 4. The launch responds to Chinese rivals' momentum—Moonshot AI's Kimi K3 hit capacity constraints from demand, and Alibaba is teasing Qwen 3.8 Max. For practitioners, 3.6 Flash's 17% token reduction and improved agentic performance at lower cost per task make it the clear choice for multi-turn agent systems and document analysis at scale.