Gemini 3.5 Flash-Lite

Googlefast New

Google's cheapest current-generation model at $0.30 input / $2.50 output per 1M tokens ($0.03 cached), released 21 Jul 2026 alongside 3.6 Flash. 1M context. Competitive on high-volume classification and extraction work, though DeepSeek V4 Flash still undercuts it roughly 2x on input and 9x on output.

Context Window

1M

tokens

Max Output

33K

tokens

Input Price

$0.3

per 1M tokens

Output Price

$2.5

per 1M tokens

Speed

192

tokens/sec

Released

Jul 2026

2026-07-21

Blended Cost

$1.40

per 1M tokens

Value Score

56.4

quality per $

Capabilities

ChatVisionFunction CallingCode Generation

Benchmarks

Quality Index
79
MMLU Pro
81.2
HumanEval (Coding)
83.6
MATH
78.9
Arena ELO
1418

About Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite is a fast AI model by Google, released on July 21, 2026. It supports a context window of 1000K tokens and can generate up to 33K output tokens.

At $0.3 per million input tokens and $2.5 per million output tokens, its blended cost of $1.40/1M tokens puts it in the mid-range pricing tier. Its value score of 56.4 reflects the balance of quality and cost.

Using Gemini 3.5 Flash-Lite with Swfte

Access Gemini 3.5 Flash-Lite through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try Gemini 3.5 Flash-Lite in our AI Playground before integrating.