← The journal
Open Weight Buyer'S Guide

The Best Open-Weight LLM in 2026: DeepSeek-V4.1-Flash vs GLM-5.3

The open tier closed the capability gap in four weeks. Licence terms and cost of ownership did not. A buyer's guide.

Swfte Journal / Open Weight Buyer'S Guide

DeepSeek-V4.1-Flash scores 90.6 on Terminal-Bench 2.1. Claude Opus 5 scores 89.1. GPT-5.6 Sol scores 88.8. The model at the top of that list published its weights to Hugging Face under an MIT licence at 02:17 UTC on 10 September, and its hosted API runs at $0.30 per million input tokens against $1.20 output. In January, every clause of that paragraph would have read like a press release nobody fact-checked.

So the capability argument is, on that benchmark at least, finished. What replaced it is harder to answer because it is not a number. Are you allowed to ship this? Under what terms, on whose hardware, and at what cost once the real bill arrives rather than the headline rate? Three serious open-weight releases landed inside four weeks — GLM-5.3 on 18 August, GLM-5.3-Flash on 26 August, DeepSeek-V4.1-Flash on 10 September — and they sit within five points of each other on Artificial Analysis's Intelligence Index while differing by more than an order of magnitude on price and diverging completely on what their licences permit.

That divergence is the story. A leaderboard row tells you one thing about a model and quietly hides four others. This is a guide to the four.

Start with the licence, not the benchmark

If you take one thing from this post, take this table. It is the part that decides whether the rest of the decision is even yours to make.

ModelReleasedLicenceWhat that actually means
DeepSeek-V4.1-Flash10 Sep 2026MITGenuinely permissive. Commercial use allowed. The strongest licence position in the tier, by a distance.
GLM-5.3-Flash26 Aug 2026MITSame terms. Weights on Hugging Face, commercial use allowed.
GLM-5.318 Aug 2026GLM-5.3 License (custom)Commercial use permitted subject to restrictions. Not MIT, not Apache. You have to read the document.
Kimi K316 Jul 2026Kimi K3 License (custom)Commercial use permitted subject to restrictions. Same posture, different bespoke document.
Qwen3.8 MaxGA 3 Aug 2026; 0902 refresh 2 SepUnresolvedDo not plan against a licence nobody can currently point at.

Two of those rows are clean. Two require a lawyer and an afternoon. One is a genuine mess.

Confirmed about Qwen3.8 Max: Alibaba open-weighted the model in August, the hosted API is live at $2 input against $6 output, and Artificial Analysis records an Intelligence Index of 45 — the joint-highest score in this comparison.

Not confirmed: that the model you would download is the model you just benchmarked. Artificial Analysis lists the 0902 checkpoint's weights as not publicly available, which contradicts the August open-weighting. The only honest reading is that the hosted API and whatever open checkpoint exists are two different products with two different release histories, and we are not going to assert a licence for either one. If your plan depends on self-hosting Qwen, the first step is not benchmarking. It is finding the licence file and reading it. We took the same position when the weights were first promised and nothing since has resolved it.

For GLM-5.3 and Kimi K3, the public position is that commercial use is permitted subject to restrictions. We are not going to characterise those restrictions further, because we have not read a clause that lets us. That is not a knock on either lab. It is the difference between a licence you can clear in five seconds because you already know what MIT says, and one that needs a procurement review before an engineer touches it. On a two-week evaluation cycle, that difference is most of the cycle.

What "open weights" does not promise

Three assumptions travel with the phrase, and none of them follow from it.

It does not mean the licence is OSI-approved. MIT is. "The GLM-5.3 License" and "the Kimi K3 License" are bespoke documents written by the lab that shipped the model, and bespoke means the terms are whatever the document says rather than whatever you remember from the last model you deployed. Kimi K3's terms were the first mainstream example of this pattern and they will not be the last.

It does not mean hosted and self-hosted behave identically. The hosted endpoint has a serving configuration, a quantisation, a system prompt and a set of tool integrations that you do not inherit when you download a checkpoint. Benchmarks are almost always run against the hosted API. Your self-hosted deployment is a different artefact, and the honest expectation is that it performs somewhat worse until you have done the tuning work yourself.

It does not mean the advertised context is usable context. Every model in this comparison advertises a million tokens. What that figure means in practice is a function of your KV cache budget, your batch size and your latency tolerance, not of the model card. Treat the advertised maximum as a ceiling you have to pay to approach.

The capability picture, honestly ranked

On the Artificial Analysis Intelligence Index, the order is: GLM-5.3 at 45, Kimi K3 at 44, GLM-5.3-Flash at 42, DeepSeek-V4.1-Flash at 40.

Now notice that this ordering inverts almost perfectly on price, and inverts again on Terminal-Bench 2.1, where the model ranked last on the composite index posts 90.6 and beats two closed frontier models to do it. That is not a contradiction. It is what happens when you compress a model into one scalar.

Dashes below mean the figure is not published in the data we are working from. We would rather leave a gap than fill it.

GLM-5.3Kimi K3GLM-5.3-FlashDeepSeek-V4.1-Flash
AA Intelligence Index45444240
Total parameters753B2.8T320B552B
Active per token40B104B18B8B prefill / 16B decode
Input modalitiesText—Text + imageText + image
Context1M1M1M1M
Output speed67.0 tok/s36.8 tok/s117.2 tok/s214.4–219.8 tok/s
Time to first token3.11s——1.13–1.21s
API providers22—188
Price per 1M in / out$1.40 / $4.40$3.00 / $15.00$0.15 / $0.50$0.30 / $1.20

A composite index is the right tool when you are choosing between two models of roughly equal cost and you want a general-purpose tiebreak. It is the wrong tool here, because the spread on price across this table is over an order of magnitude and a five-point index spread cannot possibly justify that. Five points is the distance between GLM-5.3 and DeepSeek-V4.1-Flash. Twelve times is the distance between their output rates.

Worth noting that the larger GLM model takes text only, while the cheaper Flash variant takes images as well as text. If your workload involves screenshots, diagrams or scanned documents, the index ordering is irrelevant — the 45 cannot do the job and the 42 can.

The architectural detail underneath DeepSeek's numbers is the reason it can be both fast and cheap. It is a 552B-parameter mixture of experts with only 8B active while it reads your prompt and 16B while it writes, built as a 40-layer causal encoder-decoder: twenty encoder layers feeding twenty decoder layers, with the decoder's global KV cache projected from the final encoder hidden states rather than from each decoder layer's own. We walk through the consequences in the deep dive. For this post, the relevant consequence is that an 8B active footprint on prefill is why the time to first token is roughly 1.2 seconds rather than roughly 3.

Price per unit of intelligence

This is our derivation, not a vendor figure. We blend input and output at 3:1, which is a reasonable shape for retrieval-augmented and agentic workloads and is stated here so you can redo it with your own ratio. Blended cost is (3 × input + output) / 4. Cost per index point is that blended figure divided by the Intelligence Index.

ModelIn / out per 1MBlended at 3:1AA indexCents per index point
GLM-5.3-Flash$0.15 / $0.50$0.238420.57c
DeepSeek-V4.1-Flash$0.30 / $1.20$0.525401.31c
GLM-5.3$1.40 / $4.40$2.150454.78c
Kimi K3$3.00 / $15.00$6.0004413.64c

GLM-5.3-Flash is not marginally better value. It is roughly eight times better value than GLM-5.3 and roughly twenty-four times better value than Kimi K3, for a model three points behind the first and two behind the second. Artificial Analysis said the quiet part about its own bigger sibling: GLM-5.3 is "amongst the leading models in intelligence, but particularly expensive when comparing to other open weight models of similar size."

One cross-check, because our arithmetic deserves scrutiny too. Artificial Analysis's own measured cost to run its full evaluation suite came to $280.28 for GLM-5.3-Flash against $476.89 for DeepSeek-V4.1-Flash — a 1.70x gap, where our blended-rate maths predicts 2.21x. Per task, AA records $0.25 and $0.27 respectively. The three figures do not scale identically, which is precisely why you should run your own prompts before committing. Our table is a first-pass filter, not a procurement document.

Also worth flagging: DeepSeek publishes a blended rate of $0.18 per million on a 7:2:1 input, cached-input and output mix, helped by a 98% cache discount. That is a real number and it is also a flattering blend. Vendors choose the ratio that suits them. Choose yours from your own logs.

And a floor to keep in view: DeepSeek V4 Pro still lists at $0.435 input against $0.87 output and continued to be served past 14 September, after DeepSeek reversed a plan to route its traffic onto V4.1-Flash. If your workload is well served by it, nothing above obliges you to move.

The verbosity tax

Artificial Analysis flags both GLM-5.3-Flash and DeepSeek-V4.1-Flash as "very verbose". The index median is about 140M output tokens. GLM-5.3-Flash spends 180M. DeepSeek-V4.1-Flash spends roughly 250M.

Our derivation from those figures: that is a 1.29x output multiplier for GLM-5.3-Flash and 1.79x for DeepSeek-V4.1-Flash against the median model. Apply each to its own output rate and the effective cost of a unit of useful output becomes $0.64 per million for GLM-5.3-Flash and $2.14 for DeepSeek-V4.1-Flash.

The headline output rates differ by 2.4x. Adjusted for verbosity, they differ by 3.3x. On an output-billed workload — summarisation, code generation, anything agentic that writes as it works — that gap compounds every single request, and it is invisible on a pricing page. This is the single most reliable way a cheap-looking model turns out not to be, and it is worth measuring on your own traffic before you believe any per-token comparison, including ours.

Running it yourself

Weights are only free if you own the machines, so here is the rough shape. The arithmetic below is ours, derived from the published parameter counts at two common quantisation levels, at half a byte per parameter for 4-bit and one byte for 8-bit. It covers weights only and excludes activation memory, framework overhead and KV cache.

  • GLM-5.3-Flash, 320B: about 160 GB at 4-bit, about 320 GB at 8-bit.
  • DeepSeek-V4.1-Flash, 552B: about 276 GB at 4-bit, about 552 GB at 8-bit.
  • GLM-5.3, 753B: about 377 GB at 4-bit, about 753 GB at 8-bit.
  • Kimi K3, 2.8T: about 1.4 TB at 4-bit, about 2.8 TB at 8-bit.

Kimi K3 is rack-scale and no active-parameter count changes that. A 104B active footprint tells you about compute per token; it tells you nothing about memory, because a mixture of experts holds every expert resident whether or not the router picks it. You are provisioning for 2.8 trillion parameters.

GLM-5.3-Flash at 320B is the one most teams can realistically run on hardware they already have or could plausibly buy. That is a different category of decision from the other three, and combined with its position on the value table it is the reason this post has a favourite.

DeepSeek-V4.1-Flash changes the long-session storage budget in a way that is easy to miss. It stores KV entries in four-bit floating point at a global footprint of roughly 890 bytes per token — about a quarter of what V4-Flash required, which puts the previous generation near 3,560 bytes per token by the same arithmetic. Our derivation: a full million-token session therefore costs around 0.89 GB of KV cache, so a hundred concurrent maximum-length sessions fit in roughly 89 GB. On the older footprint the same hundred sessions would have wanted around 356 GB. Its persistent SSD cache also drops to roughly an eighth of the previous generation's. If you serve long conversations, that is a material line item rather than a footnote.

No equivalent KV figures have been published for the GLM models, so we are not going to estimate them. Measure on your own hardware.

Throughput belongs in this section too, because it sets serving cost directly: a machine that emits tokens twice as fast serves twice the traffic for the same capital. DeepSeek-V4.1-Flash runs at roughly 214 to 220 tokens per second with 1.13 to 1.21 seconds to first token. GLM-5.3-Flash manages 117.2. GLM-5.3 manages 67.0 with 3.11 seconds to first token. Kimi K3 manages 36.8. DeepSeek is running about six times Kimi's output speed at roughly a tenth of the blended price, and that is before anyone opens a licence document.

One last resilience number, which is not about the model at all. GLM-5.3 is reachable through 22 API providers, GLM-5.3-Flash through 18, DeepSeek-V4.1-Flash through 8. More providers means more price competition, easier failover, and less exposure when a single vendor changes terms or has a bad week. Eight is not few. It is fewer, and it is a reason to keep your integration layer provider-agnostic.

What is not in this category

Two things keep drifting into open-weight comparisons where they do not belong.

Meta is out. Muse Spark 1.3 shipped on 2 September with closed weights, and Llama 5 has slipped to roughly 2027. Meta has exited open frontier releases for now. If your roadmap contains a line item that assumes an open Llama flagship this year, delete it. The closed tier is where Meta competes now, and the frontier comparison is the right frame for it.

MiniMax H3 is a video model. It is omni-modal with 768p open checkpoints, and it is not a text LLM. It does not belong on a text leaderboard or in this comparison. Its community licence also excludes the United States, the EU, the United Kingdom and Korea, which makes it an odd thing to hold up as an example of open access regardless of modality.

The call

Four workloads, four answers, no hedging.

Cheapest credible production tier: GLM-5.3-Flash. An Intelligence Index of 42 at $0.15 and $0.50 is the outlier on this board and nothing else is close on value. MIT licence, image input, 117.2 tokens per second, 18 providers. Start here and make something else prove it needs to be replaced.

Best licence certainty: DeepSeek-V4.1-Flash, with GLM-5.3-Flash level alongside it. Both are MIT. Both are done in five seconds. If your procurement process is the bottleneck rather than your GPU budget, this pair is the shortlist and the other three are not on it.

Best raw capability: GLM-5.3, and be honest with yourself about paying for it. It leads the open tier at 45 with 22 providers behind it, and it costs roughly eight times more per index point than its own Flash sibling, runs at half the speed, and ships under a custom licence. That is a defensible trade for a narrow slice of hard reasoning work. It is not a default. If the slice you care about is terminal and agentic work specifically, DeepSeek-V4.1-Flash's 90.6 on Terminal-Bench 2.1 says the capability crown depends entirely on which benchmark you asked.

Best for on-premise: GLM-5.3-Flash, with DeepSeek-V4.1-Flash where long sessions dominate. 320B at 4-bit is roughly 160 GB of weights, which is a real deployment rather than a datacentre project. Choose DeepSeek instead when your sessions are long enough that KV cache, not weights, is what you are actually provisioning — 890 bytes per token changes that calculation more than any benchmark does.

The uncomfortable summary is that the two models with the best licences are the two with the lowest composite index scores, and the model with the highest score is the one you would have to get cleared. That is a genuinely new problem. A year ago the open tier was not good enough to have it.


Related: the DeepSeek-V4.1-Flash deep dive, the Kimi K3 self-host guide, Qwen3.8 Max and its unresolved weights, and the live model leaderboard.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.