DeepSeek-V4.1-Flash: MIT Weights, 552B Parameters, and the VRAM Maths
An MIT-licensed model just topped Terminal-Bench 2.1. The architecture, the serving arithmetic, and when to self-host.
Four facts about DeepSeek-V4.1-Flash, each of which would carry a post on its own. It is MIT licensed. It is a 552-billion-parameter mixture-of-experts backbone that activates 16 billion parameters while generating. It scores 90.6 on Terminal-Bench 2.1, narrowly ahead of Claude Opus 5 at 89.1 and GPT-5.6 Sol at 88.8. And the hosted API costs $0.30 in, $1.20 out per million tokens.
Put them together and you have the most interesting open-weight release of 2026, and one of the few where the licence is the headline rather than the footnote.
The weights went up on Hugging Face at 02:17 UTC on 10 September, at deepseek-ai/DeepSeek-V4.1-Flash, with the X announcement following just under four hours later. That ordering is itself a small statement — the file existed before the press did.
The licence, first, because it is the part that has been going wrong
The recurring frustration with the 2026 open-weight wave has not been capability. It has been terms.
Kimi K3 shipped 2.8 trillion parameters under the custom Kimi K3 License, so anybody embedding it in a product had to read the actual file rather than recognise a familiar name. Qwen3.8 arrived with an open-weight commitment and no published licence, and has since got murkier: Artificial Analysis now lists the 0902 checkpoint's weights as not publicly available at all, which contradicts the August framing badly enough that the hosted API and any open checkpoint are best treated as two separate products. GLM-5.3 is open-weight under the GLM-5.3 License — custom, commercial use permitted subject to restrictions, a sentence that costs legal review time every time it appears in a procurement document.
V4.1-Flash is MIT. Not MIT-with-an-appendix, not a community licence that borrows the shape of MIT. MIT. Commercial use, modification, redistribution, embedding in a closed product you sell — all of it, with attribution and a warranty disclaimer, and nothing else.
Worth being precise rather than triumphal: it is not the only permissive release of the season. GLM-5.3-Flash, which shipped on 26 August at 320B total and 18B active, is also MIT. So the accurate claim is narrower and still significant — V4.1-Flash is the largest and strongest genuinely MIT-licensed model of the wave, and the only one of the top-tier open releases where "can we ship this" is a five-minute question rather than a two-week one.
If your deployment plan has ever died in legal review, that difference is worth more than a few points of benchmark.
The CED architecture, and why it is the interesting part
DeepSeek did not just make a smaller model. It changed the shape.
Confirmed by DeepSeek: V4.1-Flash is a 40-layer Transformer arranged as a 20-layer causal encoder followed by a 20-layer decoder — a Causal Encoder-Decoder, or CED. The decoder maintains a single global KV cache, and that cache is projected from the final encoder hidden states rather than from each decoder layer's own hidden states. KV entries are stored in four-bit floating point, and the resulting global footprint is roughly 890 bytes per token, about a quarter of what V4-Flash required. The persistent SSD cache drops to about an eighth of the previous generation.
Unpack the middle sentence, because it is doing all the work.
In a conventional decoder-only Transformer, every layer keeps its own keys and values for every token in the context. Forty layers means forty caches. Memory scales with layers × tokens × heads × head dimension, and that product is why long context is expensive — not the attention compute, which modern kernels handle, but the volume of tensors you keep resident for the whole session.
CED breaks that link. The encoder runs causally over the prompt and produces a final hidden state per token. The decoder's global cache is projected once, from those states, and every decoder layer reads the same cache. You are no longer storing twenty decoder-layer caches; you are storing one shared representation.
Reasonable inference, labelled as such: DeepSeek publishes the architecture and it publishes the 890 bytes per token, but it does not publish a line-by-line derivation connecting the two. Connecting them is fair — a single projected cache instead of twenty per-layer caches is the obvious mechanism for a roughly-4x reduction, and fp4 storage supplies the rest of it — but treat the causal story as mine rather than theirs. The number is confirmed. The explanation is inferred.
Either way, the consequence is concrete: 890 bytes per token means a full one-million-token context costs about 933 MB of cache for a single sequence. Under a gigabyte. For a megatoken window. That is the figure to hold on to when the rest of this post does arithmetic.
The VRAM arithmetic
This section is derived. Every number below is my own calculation from DeepSeek's published parameter count and cache footprint, not a vendor-published sizing figure. Card capacities are standard published specs. Check my working; that is the point of showing it.
Weights cost roughly two bytes per parameter at bf16, one byte at fp8, and a little over half a byte at 4-bit. At 552 billion parameters:
- bf16: 552 × 10⁹ × 2 bytes = 1,104 GB
- fp8: 552 × 10⁹ × 1 byte = 552 GB
- 4-bit: 552 × 10⁹ × 0.5 bytes = 276 GB
KV cache at 890 bytes per token, per sequence:
| Context | Tokens | KV cache |
|---|---|---|
| 32K | 32,768 | ~29 MB |
| 128K | 131,072 | ~117 MB |
| 1M | 1,048,576 | ~933 MB |
Now the part that surprises people. Take a production-shaped configuration — 128K context, batch 32 — and the cache comes to 32 × 117 MB ≈ 3.7 GB. Against 276 GB of 4-bit weights, the cache is rounding error. Even at the full 1M window with a batch of 8, you are at about 7.5 GB.
| Precision | Weights | KV (128K, batch 32) | Total | Lands on |
|---|---|---|---|---|
| bf16 | ~1,104 GB | ~4 GB | ~1,108 GB | 8x H200 141GB, two nodes in practice |
| fp8 | ~552 GB | ~4 GB | ~556 GB | 4x H200, or 3x B200 192GB with no room to breathe |
| 4-bit | ~276 GB | ~4 GB | ~280 GB | 2x H200 on paper, 4x in reality |
The "in reality" column is not padding. Two H200s give you 282 GB against a 280 GB requirement: nothing for activations, no headroom for a larger batch, and no margin for the higher-precision residue — normalisation parameters, the router, embeddings, the LM head — that does not quantise down with the expert weights. Plan four cards and you have a system rather than a demonstration.
The MoE sizing mistake, stated plainly
V4.1-Flash activates 8 billion parameters while processing a prompt and 16 billion while generating output. Neither of those numbers has anything to do with how much memory you need.
You must hold all 552 billion in VRAM, because the router chooses experts per token and you cannot know in advance which ones it will want. Sparsity cuts compute per forward pass. It does not cut footprint. A 16B-active model is a 552B model as far as your GPU budget is concerned, and it is a 16B model as far as your throughput is concerned — which is precisely why it runs at 214 to 220 tokens per second with a time to first token of 1.13 to 1.21 seconds, against a field median TTFT around 3.7 seconds.
This is the most common planning error at this tier, and it is worth saying bluntly: "16B active" is a speed claim, not a memory claim. Teams who size a cluster against the active count end up with a quarter of the hardware they need and a very confusing first week.
The SSD cache point
The persistent cache drops to about an eighth of the previous generation. For a short session that is invisible. For anything long-lived — an agent holding working context across hours, a coding assistant with a resident repository, a support system keeping conversation state — storage is one of the quieter line items that grows without anybody noticing until it does not fit. An eighth is the kind of change that turns a provisioning conversation into no conversation at all.
Self-hosting versus the hosted API
Here is where honesty is more useful than enthusiasm.
At $0.30 input and $1.20 output per million, with a 98% cache discount and a blended rate of $0.18 per million, V4.1-Flash is served through eight API providers at a price that makes the self-hosting case difficult on cost alone.
Reasonable inference, labelled as such: DeepSeek quotes the blended figure against a 7:2:1 ratio, and the arithmetic only reproduces $0.18 if that ratio reads as seven parts cache-hit input, two parts fresh input, one part output — (7 × $0.006 + 2 × $0.30 + 1 × $1.20) / 10 = $0.184. Which means the headline blended rate assumes a 70% cache hit rate. If your prompts do not share a stable prefix, your real blended cost is meaningfully higher than $0.18 and closer to the raw card.
Now the break-even, also derived. Four H200-class cards on demand run somewhere in the region of $2 to $3 per GPU-hour depending on provider and commitment — a market observation, not a DeepSeek figure, and yours will differ. Call it $10 an hour for the node, kept up continuously: roughly $7,300 a month. At a blended $0.18 per million, that same $7,300 buys about 40 billion tokens on the hosted API.
Forty billion tokens a month is an enormous workload, and most teams reading this do not run a tenth of it. That is before you cost the engineer who keeps the cluster alive, the expert-parallel interconnect that decides your real throughput, and the on-call rotation for a multi-card inference system.
So the case for self-hosting V4.1-Flash is not cost. It is:
- Data residency. Some prompts cannot leave your network, and no price makes that acceptable.
- Latency control. A local endpoint has a floor you own rather than one your provider owns.
- Vendor independence. Weights on disk cannot be repriced, rate-limited or deprecated.
- Modification. An API cannot be fine-tuned. MIT weights can, and can then ship inside a product.
Those are excellent reasons. Cost joins them somewhere past several tens of billions of tokens a month, and not before. The fuller version of this comparison is in our DeepSeek pricing analysis.
The V4 Pro routing saga, and the lesson that is not about DeepSeek
DeepSeek originally announced that from 04:00 UTC on 14 September, all deepseek-v4-pro requests would route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro launched. Then, in response to user demand, it reversed the decision and continued serving V4 Pro past the 14th with billing unchanged.
The reversal was the right call and the episode is still worth reading carefully, because the lesson is structural rather than about any one lab. A model ID is not a stable contract unless you treat it as one. If your production traffic goes to a name, and the name can be pointed somewhere else on a fortnight's notice, a vendor's product decision becomes a change to your system that you did not review, test or approve.
Two operational notes follow from the same release. The new API name is deepseek-flash. And V4 Flash and V4 Flash Vision Exp are retired, with their old names routing to V4.1-Flash temporarily — temporarily being the operative word, since if you are still calling them you are relying on a compatibility shim with no published end date.
Pin your model IDs, pin snapshots where the provider offers them, and put a canary on the endpoint that tells you when the thing behind the name starts behaving differently. Cheap insurance, and almost nobody has it.
The verbosity caveat, which partly eats the cheap headline
Artificial Analysis describes V4.1-Flash as "very verbose" — roughly 250 million output tokens across its Intelligence Index run against a 140 million median. That is about 1.8 times the typical output volume for the same set of tasks.
Output is the expensive side at $1.20 per million, four times the input rate. A 1.8x multiplier on the dear half of the bill is not a rounding error, and it means the effective cost of V4.1-Flash on output-heavy work is materially worse than the card rate implies. Summarisation, long-form generation and anything with a chatty agent loop are where this bites.
Self-hosting does not escape it; it just changes the currency. More tokens per task means fewer tasks per GPU-hour, so verbosity shows up as throughput rather than invoice. Either way, normalise on cost per completed task rather than cost per million tokens when you compare. On that basis V4.1-Flash still looks strong — $0.27 per AA task, against $0.25 for GLM-5.3-Flash, $3.26 for GPT-6 Astra at max effort and $7.63 for Claude Fable 5.1 — but the gap is narrower than the sticker price suggests, and that is the correct way to read it.
One more piece of context the headline hides. V4.1-Flash scores 40 on the AA Intelligence Index, which puts it below GLM-5.3 at 45, Kimi K3 at 44 and GLM-5.3-Flash at 42. A model can lead the field on Terminal-Bench 2.1 and sit near the bottom of a general-intelligence composite at the same time, and both measurements are telling you something true. What they jointly say is that V4.1-Flash is sharply specialised toward agentic and terminal work rather than broadly strong. Pick it for the former and you will be delighted; pick it for the latter and you will not.
Worth noting for fairness: Google reports Terminal-bench 2.1 of 89.4 for Gemini 3.8 Flash, but that is Google's own run rather than an independent board, so it does not belong in the same column as the 90.6. We compare the two leading open-weight options directly in the V4.1-Flash versus GLM-5.3 roundup.
Running it
vLLM remains the sensible default, with SGLang worth benchmarking if your traffic shares long prefixes. Four settings decide whether a deployment is good or merely functional, and three of them are the same ones people always miss.
Set max model length to what you actually use rather than the model's 1M maximum, because vLLM pre-allocates against it and the 890-bytes-per-token figure is only cheap if you are not reserving the whole window for every slot. Turn prefix caching on if your requests share a preamble — given the 98% cache discount on the hosted side, DeepSeek clearly expects that they will. Leave continuous batching on, which it is by default. And with a cache this small, the usual advice to quantise the KV cache is largely moot: it is already fp4 by design, and your headroom problem is the 276 GB of weights, not the 4 GB of cache.
DeepSeek names WorkBuddy (including CodeBuddy) and OpenCode as official partners with full V4.1-Flash support, which is a reasonable shortcut if you want to evaluate the model on real agentic work before committing to infrastructure.
Who should deploy this
Deploy it if you run agentic or terminal-heavy workloads at volume, you need MIT terms for a product you ship, your prompts cannot leave your network, or you want a strong base to fine-tune and own. The licence removes the objection that usually kills these projects, and the architecture means the long-context story is genuinely cheap rather than theoretically cheap.
Stay hosted if your volume is under a few billion tokens a month, which is most people. At $0.18 blended across eight providers, the API is cheaper than the cluster and considerably cheaper than the cluster plus the person who runs it.
Stay on a frontier model if your work is broad reasoning rather than tool use. A 40 on the index is a real ceiling, and Claude Opus 5 or GPT-6 Astra will do better on the hard general slice even though V4.1-Flash beats them at the terminal. The sane architecture is both: route the agentic volume here and keep a frontier model for the difficult remainder.
The genuinely new thing this release delivers is not the 90.6. It is that a model at the top of a frontier benchmark now comes with a licence you do not have to read twice. That has not been true all year, and it should not go unremarked.
Related: Kimi K3's self-hosting requirements, the Qwen3.8 deep dive, and current rankings on the AI model leaderboard.