DeepSeek V4 Flash

DeepSeekfastOpen Source

The cheapest usable model in the directory at $0.14 input / $0.28 output per 1M tokens, with cache hits at $0.0028 — DeepSeek cut the cache-hit rate to a tenth of its launch price on 26 Apr 2026. 284B MoE / 13B active, MIT weights, 1M context, and an OpenAI-compatible endpoint. On 31 Jul 2026 the `deepseek-v4-flash` API ID began serving the V4-Flash-0731 checkpoint, which posts 79% on SWE-bench Verified and 85.9 on BrowseComp — numbers that would have been frontier a year earlier, at roughly 1/100th of GPT-5.6 Sol's output price. The hosted service is in public beta. Rate limit is 2,500 concurrent requests.

Context Window

1M

tokens

Max Output

16K

tokens

Input Price

$0.14

per 1M tokens

Output Price

$0.28

per 1M tokens

Speed

105

tokens/sec

Released

Apr 2026

2026-04-24

Blended Cost

$0.21

per 1M tokens

Value Score

400.0

quality per $

Capabilities

ChatFunction CallingCode GenerationReasoning

Benchmarks

Quality Index
84
MMLU Pro
84.6
HumanEval (Coding)
87.2
MATH
82.4
Arena ELO
1441

Open Source: Licensed under MIT

About DeepSeek V4 Flash

DeepSeek V4 Flash is a fast AI model by DeepSeek, released on April 24, 2026. It supports a context window of 1000K tokens and can generate up to 16K output tokens.

At $0.14 per million input tokens and $0.28 per million output tokens, its blended cost of $0.21/1M tokens makes it one of the most affordable models available. Its value score of 400.0 reflects the balance of quality and cost.

DeepSeek V4 Flash is available as an open-source model under the MIT license, meaning you can self-host it for predictable costs or use it through API providers like Swfte Connect.

Using DeepSeek V4 Flash with Swfte

Access DeepSeek V4 Flash through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try DeepSeek V4 Flash in our AI Playground before integrating.