Google ships Gemini 3.6 Flash at 17% fewer tokens, $7.50 output price; 3.5 Pro still delayed
Google released Gemini 3.6 Flash on July 21, 2026 as its workhorse model for agentic coding, knowledge work, and multimodal tasks. Output token pricing fell from $9.00 to $7.50 per million (input unchanged at $1.50), and the model consumes 17% fewer output tokens per task than Gemini 3.5 Flash on the Artificial Analysis Index, reaching 65% efficiency gains on specific agentic benchmarks like DeepSWE. The combined sticker-price cut plus token efficiency compounds to an effective 31% cost reduction per completed task, with up to 71% savings on agentic coding workloads.
Google also shipped Gemini 3.5 Flash-Lite at $0.30/$2.50 per million tokens for high-throughput tasks and announced Gemini 3.5 Flash Cyber, a security-tuned model restricted to government and partner access. The March 2026 knowledge cutoff represents a 14-month advance over 3.5 Flash's January 2025 date. On applied agentic benchmarks, 3.6 Flash gained ground: DeepSWE improved from 37% to 49%, MLE Bench from 49.7% to 63.9%, and OSWorld computer-use from 78.4% to 83.0%.
However, on the independent Artificial Analysis Intelligence Index, Gemini 3.6 Flash scores exactly 50—unchanged from 3.5 Flash. This is an efficiency release, not a capability jump. The model trades raw reasoning power for token efficiency and lower latency, optimizing for the specific economics of agentic workflows rather than frontier reasoning depth. Artificial Analysis independently measured average task completion time dropping from 2.7 minutes to 1.3 minutes and average cost per task from $0.59 to $0.50.
Notably absent: Gemini 3.5 Pro remains in partner testing with no public availability date, having missed its May 2026 I/O promise and subsequent June target. While Google is shipping three models, the flagship of the 3.5 generation is stalled. Gemini 4 pre-training has begun and will ship later. For teams running high-volume agentic inference, 3.6 Flash's pricing and efficiency gains matter; for those evaluating flagship capability, Google's delivery timeline continues to lag Anthropic and OpenAI.
Sources
- Primary source
- 9to5google.com
“Gemini 3.6 Flash consumes 17% fewer output tokens compared to 3.5 Flash, while taking fewer reasoning steps and tool calls to accomplish multi-step workflows, and is priced lower at $1.50/1M input tokens and $7.50/1M output tokens”
- trilogyai.substack.com
“Gemini 3.6 Flash pricing is $1.50/$7.50 per million tokens — but token efficiency gains mean the effective cost per completed task drops ~31%, and up to ~71% on agentic coding workloads”
- felloai.com
“On the Artificial Analysis Intelligence Index it scores around 50 which is above the field average, but roughly flat against 3.5 Flash”
- 9to5google.com
“Gemini 3.6 Flash scores 49 percent on DeepSWE benchmark, a notable increase from the 37 percent achieved by version 3.5, and pushes machine learning engineering performance higher, scoring 63.9 percent on MLE-Bench compared to 49.7 percent previously”
- techcrunch.com
“Google teased the release of Pro as part of the 3.5 Flash release in May, saying the Pro version was already being used internally, and we look forward to rolling it out next month. Last week, Bloomberg reported that Google was facing internal delays in launching the 3.5 Pro”