Gemini 3.8 Flash
Released 2 Sep 2026 alongside a restricted Gemini 3.8 Flash Cyber variant — by Google's own count the third Flash release in six weeks, and further evidence that the Pro line has stalled: Gemini 3.1 Pro (19 Feb 2026) is still the current Pro model, and every release since has been Flash-tier. The headline number is speed. At 345.1 tokens per second it is the fastest model of the 199 Artificial Analysis measures, roughly five times the 68 tok/s median, while scoring 41 on the v4.3 Intelligence Index. The offset is a 30.49s time to first token, so the throughput only pays off on long generations. Pricing holds at the 3.7 Flash introductory rate of $0.75/$3.75 with a 90% cache discount — but note that this is introductory: from 1 Jan 2027 the rates double to $1.50 and $7.50, which is worth modelling now if you are sizing a 2027 budget. Google's own benchmark table has 3.8 Flash ahead of 3.7 Flash on every published row: DeepSWE v1.1 73.7% vs 65.3%, Terminal-bench 2.1 89.4% vs 85.8%, OSWorld-2.0 59.0% vs 50.6%, Vals Finance Agent v2 61.4% vs 59.0%, HLE-Verified 54.9% vs 53.6% — all first-party runs. Google attributes the gains to the model executing extra reasoning steps and calling tools iteratively, which also explains why Artificial Analysis flags it 'very verbose' at 170M output tokens against a 90M median. 1M context. The Cyber variant is gated behind the new Fairwind Programme for government authorities, critical infrastructure operators and software maintainers.
1M
tokens
66K
tokens
$0.75
per 1M tokens
$3.75
per 1M tokens
345
tokens/sec
Sep 2026
2026-09-02
$2.25
per 1M tokens
41.3
quality per $
Capabilities
Benchmarks
Price vs Quality
Context Window in Context
- ▸ Gemini 3.8 Flash1M
- GPT-4.11M
- Gemini 2.5 Pro1M
- Gemini 2.0 Flash1M
- Llama 4 Maverick1M
- GPT-5.51M
- GPT-5.5 Pro1M
- Claude Opus 4.71M
Pas encore sur notre gateway
Swfte Connect ne dispose pas de clé pour Gemini 3.8 Flash : il n'y a donc rien à exécuter ici. Ajoutez votre propre clé fournisseur et appelez-le avec la même API que tous les autres modèles de cet annuaire.
curl https://api.swfte.com/agents/v2/gateway/chat/completions \
-H "Authorization: Bearer $SWFTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google:gemini-3.8-flash",
"messages": [{"role": "user", "content": "Hello"}]
}'About Gemini 3.8 Flash
Gemini 3.8 Flash is a fast AI model by Google, released on September 2, 2026. It supports a context window of 1000K tokens and can generate up to 66K output tokens.
At $0.75 per million input tokens and $3.75 per million output tokens, its blended cost of $2.25/1M tokens puts it in the mid-range pricing tier. Its value score of 41.3 reflects the balance of quality and cost.
Using Gemini 3.8 Flash with Swfte
Access Gemini 3.8 Flash through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try Gemini 3.8 Flash in our AI Playground before integrating.