DeepSeek V4.1 Flash
Released 10 Sep 2026 with weights on Hugging Face under a genuine MIT licence — the permissive one, with no revenue-share clause, no territorial exclusion and no custom terms to read. That matters because it posts Terminal-Bench 2.1 of 90.6, ahead of Claude Opus 5 (89.1) and GPT-5.6 Sol (88.8), at $0.30/$1.20 hosted. The smallest model in DeepSeek's new architecture family: a multimodal MoE with a 552B backbone but only 8B parameters active while processing a prompt and 16B while generating, using a Causal Encoder-Decoder design — 40 layers arranged as a 20-layer causal encoder followed by a 20-layer decoder, where the decoder's global KV cache is projected from the final encoder hidden states rather than from each decoder layer's own. KV entries are stored in four-bit floating point at roughly 890 bytes per token, about a quarter of what V4 Flash needed, and the persistent SSD cache drops to around an eighth of the previous generation. 1M context, native image and text in. Two caveats worth pricing in: Artificial Analysis rates it 40 on the v4.3 Intelligence Index — below GLM-5.3 (45) and Kimi K3 (44) on the composite even as it leads them on Terminal-Bench — and it is flagged 'very verbose' at roughly 250M output tokens across the index against a 140M median, which quietly multiplies that cheap output rate. Call it as `deepseek-flash`; V4 Flash and V4 Flash Vision Exp are retired and their old names route here temporarily. DeepSeek had planned to route all `deepseek-v4-pro` traffic here from 14 Sep but reversed that after user pushback, so V4 Pro continues with billing unchanged. Available through 8 API providers.
1M
tokens
66K
tokens
$0.3
per 1M tokens
$1.2
per 1M tokens
217
tokens/sec
Sep 2026
2026-09-10
$0.75
per 1M tokens
122.7
quality per $
Capabilities
Benchmarks
Price vs Quality
Context Window in Context
- ▸ DeepSeek V4.1 Flash1M
- GPT-4.11M
- Gemini 2.5 Pro1M
- Gemini 2.0 Flash1M
- Llama 4 Maverick1M
- GPT-5.51M
- GPT-5.5 Pro1M
- Claude Opus 4.71M
Compare With
Open Source: Licensed under MIT
Пока нет в нашем шлюзе
DeepSeek V4.1 Flash — модель с открытыми весами, её можно запустить самостоятельно. Swfte Connect превращает ваш деплой в управляемый эндпоинт с тем же API, что и у остальных моделей каталога.
curl https://api.swfte.com/agents/v2/gateway/chat/completions \
-H "Authorization: Bearer $SWFTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek:deepseek-flash",
"messages": [{"role": "user", "content": "Hello"}]
}'About DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a open-source AI model by DeepSeek, released on September 10, 2026. It supports a context window of 1000K tokens and can generate up to 66K output tokens.
At $0.3 per million input tokens and $1.2 per million output tokens, its blended cost of $0.75/1M tokens makes it one of the most affordable models available. Its value score of 122.7 reflects the balance of quality and cost.
DeepSeek V4.1 Flash is available as an open-source model under the MIT license, meaning you can self-host it for predictable costs or use it through API providers like Swfte Connect.
Using DeepSeek V4.1 Flash with Swfte
Access DeepSeek V4.1 Flash through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try DeepSeek V4.1 Flash in our AI Playground before integrating.