GLM-5.3 Flash

Z.ai (Zhipu AI)open-sourceOpen Source

Released 26 Aug 2026 under a straight MIT licence and, on price per unit of measured intelligence, the strongest value on the board. A 320B-parameter MoE with 18B active, scoring 42 on the Artificial Analysis Intelligence Index v4.3 — three points behind its own 753B flagship GLM-5.3, at $0.15/$0.50 against that model's $1.40/$4.40. That is roughly a tenth of the input cost for 93% of the score. It is also the open-weight model most teams can realistically self-host: 320B total is demanding but not rack-scale in the way Kimi K3's 2.8T is, and the 18B active count keeps inference cheap. 1M context, text and image in, text out, 117.2 tokens/second, 18 API providers. One caveat: Artificial Analysis flags it 'very verbose' at 180M output tokens across the index against a 140M median, so budget output tokens at roughly 1.3x what the headline rate implies.

Context Window

1M

tokens

Max Output

66K

tokens

Input Price

$0.15

per 1M tokens

Output Price

$0.5

per 1M tokens

Speed

117

tokens/sec

Released

Aug 2026

2026-08-26

Blended Cost

$0.33

per 1M tokens

Value Score

286.2

quality per $

Capabilities

ChatVisionFunction CallingCode GenerationReasoning

Benchmarks

Quality Index
93

Price vs Quality

6080100$1.00$10.00$100.00Blended cost per 1M tokens (log scale) →GPT-4o — $6.25/1M, quality 85GPT-4o Mini — $0.375/1M, quality 72o3 Mini — $2.75/1M, quality 88o3 — $5.00/1M, quality 94GPT-4.1 — $5.00/1M, quality 89Claude Opus 4 — $45.00/1M, quality 91Claude Sonnet 4 — $9.00/1M, quality 88Claude 3.5 Haiku — $2.40/1M, quality 75Gemini 2.5 Pro — $5.63/1M, quality 92Gemini 2.0 Flash — $0.250/1M, quality 74Llama 4 Maverick — $0.400/1M, quality 80Llama 4 Scout — $0.275/1M, quality 71Mistral Large 2 — $4.00/1M, quality 79Codestral — $0.600/1M, quality 76DeepSeek V3 — $0.685/1M, quality 86DeepSeek R1 — $1.37/1M, quality 91Grok 3 — $9.00/1M, quality 87Grok 3 Mini — $0.400/1M, quality 78Command R+ — $6.25/1M, quality 68Amazon Nova Pro — $2.00/1M, quality 70Qwen 2.5 72B — $0.600/1M, quality 80Qwen 2.5 Coder 32B — $0.300/1M, quality 74Sonar Pro — $9.00/1M, quality 78GPT-5.5 — $17.50/1M, quality 97GPT-5.5 Pro — $105.00/1M, quality 96Claude Opus 4.7 — $15.00/1M, quality 96Gemini 3.1 Pro — $7.00/1M, quality 96DeepSeek V4 Pro — $1.32/1M, quality 89DeepSeek V4 Flash — $0.210/1M, quality 84Qwen 3.6 Plus — $3.50/1M, quality 86Grok 4.3 — $1.88/1M, quality 93Kimi K2.6 — $2.11/1M, quality 92Kimi K3 — $9.00/1M, quality 97Claude Sonnet 4.6 — $9.00/1M, quality 90GLM-5.1 — $2.03/1M, quality 88Claude Opus 4.8 — $15.00/1M, quality 98Qwen 3.7 Max — $5.00/1M, quality 94Claude Fable 5 — $30.00/1M, quality 99MiniMax M3 — $1.50/1M, quality 89GPT-5.6 Sol — $17.50/1M, quality 98GPT-5.6 Terra — $7.00/1M, quality 93GPT-5.6 Luna — $0.700/1M, quality 83Gemini 3.6 Flash — $2.25/1M, quality 90Gemini 3.5 Flash-Lite — $1.40/1M, quality 79Claude Sonnet 5 — $9.00/1M, quality 93GLM-5.2 — $2.90/1M, quality 91Claude Opus 5 — $15.00/1M, quality 99Claude Mythos 5 — $30.00/1M, quality 99Grok 4.5 — $4.00/1M, quality 94Qwen3.8 Max — $4.00/1M, quality 96GPT-6 Astra — $30.00/1M, quality 99Claude Fable 5.1 — $30.00/1M, quality 99DeepSeek V4.1 Flash — $0.750/1M, quality 92GLM-5.3 — $2.90/1M, quality 96Muse Spark 1.3 — $2.75/1M, quality 97Grok 4.6 — $4.00/1M, quality 95Gemini 3.8 Flash — $2.25/1M, quality 93Claude Opus 5.5 — $12.00/1M, quality 100Claude Sonnet 5.5 — $6.00/1M, quality 99GPT-6.1 Sol — $6.00/1M, quality 98GPT-6 Sol — $6.00/1M, quality 97GPT-6 Luna — $0.300/1M, quality 89GLM-5.3 Flash — $0.325/1M, quality 93GLM-5.3 Flash
Cheaper is left, more capable is up. Nothing in the directory is both cheaper and at least as capable. Self-host models are left off the price axis — their cost depends on the hardware you run them on, not on tokens.

Context Window in Context

  • ▸ GLM-5.3 Flash1M
  • GPT-4.11M
  • Gemini 2.5 Pro1M
  • Gemini 2.0 Flash1M
  • Llama 4 Maverick1M
  • GPT-5.51M
  • GPT-5.5 Pro1M
  • Claude Opus 4.71M
Bar length is logarithmic — each full step is 10× the tokens. Shown against the models with the nearest context windows, not the directory extremes.

Open Source: Licensed under MIT

Not on our gateway yet

GLM-5.3 Flash is open-weight, so you can run it yourself. Swfte Connect turns your deployment into a managed endpoint that speaks the same API as every other model in this directory.

The request
curl https://api.swfte.com/agents/v2/gateway/chat/completions \
  -H "Authorization: Bearer $SWFTE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zhipu:glm-5.3-flash",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
Deploy this model with Swfte Connect

About GLM-5.3 Flash

GLM-5.3 Flash is a open-source AI model by Z.ai (Zhipu AI), released on August 26, 2026. It supports a context window of 1000K tokens and can generate up to 66K output tokens.

At $0.15 per million input tokens and $0.5 per million output tokens, its blended cost of $0.33/1M tokens makes it one of the most affordable models available. Its value score of 286.2 reflects the balance of quality and cost.

GLM-5.3 Flash is available as an open-source model under the MIT license, meaning you can self-host it for predictable costs or use it through API providers like Swfte Connect.

Using GLM-5.3 Flash with Swfte

Access GLM-5.3 Flash through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try GLM-5.3 Flash in our AI Playground before integrating.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.