Skip to main content
株式会社オブライト
AI2026-07-226 min read

Gemini 3.6 Flash: Price, Speed & Benchmarks vs 3.5 (2026)

Released July 21, 2026, Gemini 3.6 Flash is Google's new default workhorse: $7.50 per 1M output tokens, 17% fewer output tokens than 3.5, and stronger agentic coding scores. A fast rundown of pricing, benchmarks, and how it differs from Flash-Lite.


Gemini 3.6 Flash is a high-efficiency model Google released on July 21, 2026. It is priced at $1.50 per 1M input tokens and $7.50 per 1M output tokens — cheaper than the previous 3.5 Flash (output $9.00). Because it also completes the same work with roughly 17% fewer output tokens, the effective cost drops further. Positioned as a "new default workhorse" for coding and agentic workloads (where the model autonomously runs multiple steps), it launched alongside a lightweight Flash-Lite and a security-focused Flash Cyber the same day.

What Gemini 3.6 Flash is

Gemini 3.6 Flash is the latest in Google's "Flash" tier — the mid-range models built for speed and cost. It doesn't carry the reasoning power of the top frontier tier, but it is designed to handle coding, knowledge work, and multimodal processing (working with images as well as text) cheaply and quickly. Its hallmark is generating "ready-to-use" output with fewer unnecessary edits and less hedging, while cutting the number of model calls and tokens needed to finish a task.

Pricing

ItemGemini 3.6 FlashFor reference: 3.5 Flash
Input (1M tokens)$1.50Same level
Output (1M tokens)$7.50$9.00
Context caching$0.15 / 1M tokens
Cache storage$1.00 / 1M tokens·hour
Output token usage~17% less than 3.5Baseline

On top of the lower output rate ($9.00 → $7.50), the same task finishes in fewer tokens, so the effective savings grow the more output-heavy the workload is — as in agentic use. Google states it can cut agent token costs by up to 65% on long-horizon engineering tasks.

What it can do (key features)

- 1M-token context window: feed in long documents and codebases at once (max output 64k tokens)
- Multimodal input: handles images as well as text
- Computer Use: run agentic tasks involving screen operation from the API
- Function calling and structured output: supports JSON-schema output and tool integration
- Search-as-a-tool: the model can use web search as a tool
- Thinking controls: adjust the depth of reasoning
- Fast output: generation speed on the order of ~300 tokens per second

Benchmarks (vs. 3.5 Flash)

In Google's published figures, agentic and coding benchmarks improve clearly over 3.5 Flash. On OSWorld-Verified (validating tasks that involve screen operation) in particular, it is cited as the top level within Google's comparison table.

BenchmarkGemini 3.6 Flash3.5 Flash
DeepSWE (software engineering)49%37%
MLE Bench (machine-learning tasks)63.9%49.7%
OSWorld-Verified (screen operation)83.0%78.4%
GDPval-AA v214211349

It is cited at 50 on the composite Artificial Analysis Intelligence Index, rated as high-efficiency for the Flash tier. Scores vary with benchmark and measurement conditions, so it is safest to read them as a trend rather than absolutes.

How to use it and where it's available

Gemini 3.6 Flash and 3.5 Flash-Lite are available from launch day through several channels: the Gemini API, Google AI Studio, Android Studio, GitHub Copilot, Gemini Enterprise, and the Gemini app. Developers typically call the model by name via the Gemini API or Google AI Studio. Flash-Lite is also being rolled into Google Search in stages.

How the three launch models differ

Three models were announced together, including 3.6 Flash. Because their purposes differ, here is a rough guide to choosing between them.

ModelPositioningInput / output (1M tokens)Main use
Gemini 3.6 FlashNew default workhorse$1.50 / $7.50Coding, knowledge work, agentic tasks
Gemini 3.5 Flash-LiteLow-latency, high-throughput$0.30 / $2.50High-volume jobs, agentic search, document processing
Gemini 3.5 Flash CyberSecurity-focusedLimited accessFinding, validating, and patching vulnerabilities (gated pilot for governments and trusted partners)

Flash-Lite is characterized by 350 tokens per second and low $0.30 / $2.50 pricing, suited to high-volume lightweight jobs and low-latency needs. Flash Cyber is tuned to find, validate, and patch software vulnerabilities; it uses a design that calls a cheap model up to five times in parallel rather than one massive call, and for now is gated to governments and trusted partners.

How to choose

- If coding or agents are the main use, pick 3.6 Flash: better performance and cost-efficiency than 3.5
- For high-volume, low-latency lightweight jobs, pick Flash-Lite: an order-of-magnitude cheaper unit price and high throughput
- If you need top accuracy, consider a higher Pro tier: Flash prioritizes efficiency, so compare against higher models for hard reasoning
- The more output-heavy the task, the bigger the benefit: the lower output rate and token reduction compound

Frequently asked questions

How much does Gemini 3.6 Flash cost?

It is $1.50 per 1M input tokens and $7.50 per 1M output tokens. That is cheaper than the previous 3.5 Flash (output $9.00), and because it completes the same work with roughly 17% fewer output tokens, the effective savings are larger still.

What changed from 3.5 Flash?

The output rate dropped from $9.00 to $7.50, and output token usage fell by about 17%. Agentic and coding benchmark scores such as DeepSWE and OSWorld-Verified also improved. It is positioned as the new default workhorse for coding and agentic tasks.

How should I choose between it and Flash-Lite?

3.6 Flash is the default model for coding, knowledge work, and agentic tasks, while 3.5 Flash-Lite is cheaper ($0.30 input / $2.50 output) and fast at 350 tokens per second, suited to high-volume or low-latency jobs. Favor 3.6 Flash for accuracy and Flash-Lite for cost and throughput.

Where can I use it?

From launch day it is available via the Gemini API, Google AI Studio, Android Studio, GitHub Copilot, Gemini Enterprise, and the Gemini app. Developers typically call the model by name through the Gemini API or Google AI Studio.

Summary

Gemini 3.6 Flash is a new default workhorse that meaningfully lowers effective cost for agentic use, through a lower output rate ($9.00 → $7.50) and roughly 17% fewer output tokens. Coding and screen-operation benchmarks like DeepSWE and OSWorld-Verified improve over 3.5 Flash, and it carries the features needed to build agents — a 1M-token context, Computer Use, and function calling. For high-volume, low-latency work the same-day Flash-Lite is cheaper, and where accuracy is paramount, comparing against a higher Pro tier is realistic. Choosing within the Flash tier by workload is the smart way to use it.

Feel free to contact us

Contact Us