Gemini 3.6 Flash

Googlebalanced New

Google's 21 Jul 2026 mid-tier release, replacing Gemini 3.5 Flash. $1.50 input / $7.50 output per 1M tokens with cached input at $0.15 and batch mode at half rate — a 17% cut on output against 3.5 Flash for a materially better model. 1M context, native audio and video understanding, and the fastest response times in the frontier-adjacent band. Note that Google shipped no Gemini 3.2–3.5 Pro: the Pro line still tops out at Gemini 3.1 Pro, and the 3.5/3.6 releases are all Flash-tier.

Context Window

1M

tokens

Max Output

64K

tokens

Input Price

$1.5

per 1M tokens

Output Price

$7.5

per 1M tokens

Speed

148

tokens/sec

Released

Jul 2026

2026-07-21

Blended Cost

$4.50

per 1M tokens

Value Score

20.0

quality per $

Capabilities

ChatVisionFunction CallingCode GenerationReasoningAudio

Benchmarks

Quality Index
90
MMLU Pro
88.1
HumanEval (Coding)
91
MATH
89.4
Arena ELO
1481

About Gemini 3.6 Flash

Gemini 3.6 Flash is a balanced AI model by Google, released on July 21, 2026. It supports a context window of 1000K tokens and can generate up to 64K output tokens.

At $1.5 per million input tokens and $7.5 per million output tokens, its blended cost of $4.50/1M tokens puts it in the mid-range pricing tier. Its value score of 20.0 reflects the balance of quality and cost.

Using Gemini 3.6 Flash with Swfte

Access Gemini 3.6 Flash through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try Gemini 3.6 Flash in our AI Playground before integrating.