GLM-5.2

Z.ai (Zhipu AI)open-sourceOpen Source

Z.ai's 13 Jun 2026 open-weight flagship, and the model that held the top open-weight slot until Kimi K3 landed a month later. 744B MoE / ~40B active, MIT-licensed weights on Hugging Face, 1M context, 131K max output. $1.40 input / $4.40 output per 1M tokens ($0.26 cached) — a blended ~$0.90/1M. The headline result is SWE-bench Pro 62.1%, which beats GPT-5.5's 58.6%: an MIT-licensed model you can self-host outscoring a $5/$30 US flagship on agentic coding. Also the fastest of the big open MoEs at ~168 tokens/sec, roughly 2.7x Kimi K3 and DeepSeek V4 Pro. Self-hosting needs about 1TB VRAM in BF16, or ~8x H200 at FP8.

Context Window

1M

tokens

Max Output

131K

tokens

Input Price

$1.4

per 1M tokens

Output Price

$4.4

per 1M tokens

Speed

168

tokens/sec

Released

Jun 2026

2026-06-13

Blended Cost

$2.90

per 1M tokens

Value Score

31.4

quality per $

Capabilities

ChatVisionFunction CallingCode GenerationReasoning

Benchmarks

Quality Index
91
MMLU Pro
88.9
HumanEval (Coding)
92.2
MATH
90.6
Arena ELO
1483

Open Source: Licensed under MIT

About GLM-5.2

GLM-5.2 is a open-source AI model by Z.ai (Zhipu AI), released on June 13, 2026. It supports a context window of 1000K tokens and can generate up to 131K output tokens.

At $1.4 per million input tokens and $4.4 per million output tokens, its blended cost of $2.90/1M tokens puts it in the mid-range pricing tier. Its value score of 31.4 reflects the balance of quality and cost.

GLM-5.2 is available as an open-source model under the MIT license, meaning you can self-host it for predictable costs or use it through API providers like Swfte Connect.

Using GLM-5.2 with Swfte

Access GLM-5.2 through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try GLM-5.2 in our AI Playground before integrating.