DeepSeek V4.1 Flash

DeepSeekopen-source NewOpen Source

Released 10 Sep 2026 with weights on Hugging Face under a genuine MIT licence — the permissive one, with no revenue-share clause, no territorial exclusion and no custom terms to read. That matters because it posts Terminal-Bench 2.1 of 90.6, ahead of Claude Opus 5 (89.1) and GPT-5.6 Sol (88.8), at $0.30/$1.20 hosted. The smallest model in DeepSeek's new architecture family: a multimodal MoE with a 552B backbone but only 8B parameters active while processing a prompt and 16B while generating, using a Causal Encoder-Decoder design — 40 layers arranged as a 20-layer causal encoder followed by a 20-layer decoder, where the decoder's global KV cache is projected from the final encoder hidden states rather than from each decoder layer's own. KV entries are stored in four-bit floating point at roughly 890 bytes per token, about a quarter of what V4 Flash needed, and the persistent SSD cache drops to around an eighth of the previous generation. 1M context, native image and text in. Two caveats worth pricing in: Artificial Analysis rates it 40 on the v4.3 Intelligence Index — below GLM-5.3 (45) and Kimi K3 (44) on the composite even as it leads them on Terminal-Bench — and it is flagged 'very verbose' at roughly 250M output tokens across the index against a 140M median, which quietly multiplies that cheap output rate. Call it as `deepseek-flash`; V4 Flash and V4 Flash Vision Exp are retired and their old names route here temporarily. DeepSeek had planned to route all `deepseek-v4-pro` traffic here from 14 Sep but reversed that after user pushback, so V4 Pro continues with billing unchanged. Available through 8 API providers.

Context Window

1M

tokens

Max Output

66K

tokens

Input Price

$0.3

per 1M tokens

Output Price

$1.2

per 1M tokens

Speed

217

tokens/sec

Released

Sep 2026

2026-09-10

Blended Cost

$0.75

per 1M tokens

Value Score

122.7

quality per $

Capabilities

ChatVisionFunction CallingCode GenerationReasoning

Benchmarks

Quality Index
92

Price vs Quality

6080100$1.00$10.00$100.00Blended cost per 1M tokens (log scale) →GPT-4o — $6.25/1M, quality 85GPT-4o Mini — $0.375/1M, quality 72o3 Mini — $2.75/1M, quality 88o3 — $5.00/1M, quality 94GPT-4.1 — $5.00/1M, quality 89Claude Opus 4 — $45.00/1M, quality 91Claude Sonnet 4 — $9.00/1M, quality 88Claude 3.5 Haiku — $2.40/1M, quality 75Gemini 2.5 Pro — $5.63/1M, quality 92Gemini 2.0 Flash — $0.250/1M, quality 74Llama 4 Maverick — $0.400/1M, quality 80Llama 4 Scout — $0.275/1M, quality 71Mistral Large 2 — $4.00/1M, quality 79Codestral — $0.600/1M, quality 76DeepSeek V3 — $0.685/1M, quality 86DeepSeek R1 — $1.37/1M, quality 91Grok 3 — $9.00/1M, quality 87Grok 3 Mini — $0.400/1M, quality 78Command R+ — $6.25/1M, quality 68Amazon Nova Pro — $2.00/1M, quality 70Qwen 2.5 72B — $0.600/1M, quality 80Qwen 2.5 Coder 32B — $0.300/1M, quality 74Sonar Pro — $9.00/1M, quality 78GPT-5.5 — $17.50/1M, quality 97GPT-5.5 Pro — $105.00/1M, quality 96Claude Opus 4.7 — $15.00/1M, quality 96Gemini 3.1 Pro — $7.00/1M, quality 96DeepSeek V4 Pro — $1.32/1M, quality 89DeepSeek V4 Flash — $0.210/1M, quality 84Qwen 3.6 Plus — $3.50/1M, quality 86Grok 4.3 — $1.88/1M, quality 93Kimi K2.6 — $2.11/1M, quality 92Kimi K3 — $9.00/1M, quality 97Claude Sonnet 4.6 — $9.00/1M, quality 90GLM-5.1 — $2.03/1M, quality 88Claude Opus 4.8 — $15.00/1M, quality 98Qwen 3.7 Max — $5.00/1M, quality 94Claude Fable 5 — $30.00/1M, quality 99MiniMax M3 — $1.50/1M, quality 89GPT-5.6 Sol — $17.50/1M, quality 98GPT-5.6 Terra — $7.00/1M, quality 93GPT-5.6 Luna — $0.700/1M, quality 83Gemini 3.6 Flash — $2.25/1M, quality 90Gemini 3.5 Flash-Lite — $1.40/1M, quality 79Claude Sonnet 5 — $9.00/1M, quality 93GLM-5.2 — $2.90/1M, quality 91Claude Opus 5 — $15.00/1M, quality 99Claude Mythos 5 — $30.00/1M, quality 99Grok 4.5 — $4.00/1M, quality 94Qwen3.8 Max — $4.00/1M, quality 96GPT-6 Astra — $30.00/1M, quality 99Claude Fable 5.1 — $30.00/1M, quality 99GLM-5.3 — $2.90/1M, quality 96GLM-5.3 Flash — $0.325/1M, quality 93Muse Spark 1.3 — $2.75/1M, quality 97Grok 4.6 — $4.00/1M, quality 95Gemini 3.8 Flash — $2.25/1M, quality 93Claude Opus 5.5 — $12.00/1M, quality 100Claude Sonnet 5.5 — $6.00/1M, quality 99GPT-6.1 Sol — $6.00/1M, quality 98GPT-6 Sol — $6.00/1M, quality 97GPT-6 Luna — $0.300/1M, quality 89DeepSeek V4.1 Flash — $0.750/1M, quality 92DeepSeek V4.1 Flash
Cheaper is left, more capable is up. 1 of 62 priced models are both cheaper and score at least as high. Self-host models are left off the price axis — their cost depends on the hardware you run them on, not on tokens.

Context Window in Context

  • ▸ DeepSeek V4.1 Flash1M
  • GPT-4.11M
  • Gemini 2.5 Pro1M
  • Gemini 2.0 Flash1M
  • Llama 4 Maverick1M
  • GPT-5.51M
  • GPT-5.5 Pro1M
  • Claude Opus 4.71M
Bar length is logarithmic — each full step is 10× the tokens. Shown against the models with the nearest context windows, not the directory extremes.

Open Source: Licensed under MIT

Non è ancora sul nostro gateway

DeepSeek V4.1 Flash è a pesi aperti, quindi puoi eseguirlo tu stesso. Swfte Connect trasforma il tuo deployment in un endpoint gestito che parla la stessa API di ogni altro modello di questa directory.

La richiesta
curl https://api.swfte.com/agents/v2/gateway/chat/completions \
  -H "Authorization: Bearer $SWFTE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek:deepseek-flash",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
Distribuisci questo modello con Swfte Connect

About DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a open-source AI model by DeepSeek, released on September 10, 2026. It supports a context window of 1000K tokens and can generate up to 66K output tokens.

At $0.3 per million input tokens and $1.2 per million output tokens, its blended cost of $0.75/1M tokens makes it one of the most affordable models available. Its value score of 122.7 reflects the balance of quality and cost.

DeepSeek V4.1 Flash is available as an open-source model under the MIT license, meaning you can self-host it for predictable costs or use it through API providers like Swfte Connect.

Using DeepSeek V4.1 Flash with Swfte

Access DeepSeek V4.1 Flash through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try DeepSeek V4.1 Flash in our AI Playground before integrating.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.