Z.ai GLM model, released June 16, 2026
Released June 16, 2026 under the MIT license: a 744B-parameter (40B active) Mixture-of-Experts model built for long-horizon coding and agent work, with a 1M-token context. The specs, Z.ai's benchmarks, today's API prices, and how to run it.
GLM-5.2 is the third flagship of Z.ai's GLM-5 generation, after GLM-5 and GLM-5.1, and it led the line until GLM-5.3 arrived in August 2026. It keeps GLM-5's Mixture-of-Experts layout (about 744B parameters, 40B active per token) and makes two big moves: the context window grows from 200K to a full 1,048,576 tokens, and coding and agent scores jump well past GLM-5.1. In Z.ai's evaluations Terminal Bench 2.1 rises from 63.5 to 81.0 and SWE-bench Pro from 58.4 to 62.1.
Z.ai also reworked attention for that window. A technique it calls IndexShare reuses one sparse-attention indexer across every four layers, which Z.ai says cuts per-token compute 2.9 times at 1M tokens, and an improved multi-token-prediction layer accepts up to 20% more tokens per speculative-decoding step. The weights are on Hugging Face under a plain MIT license, with no regional limits.
An independent reference, not affiliated with Z.ai. Checked against Z.ai's model card, docs and OpenRouter on October 4, 2026.
1M
Context window (tokens)
744B
Total parameters, 40B active
81.0
Terminal Bench 2.1 (Z.ai)
MIT
License, open weights
Specifications
From Z.ai's model card and the published config, with each row linked to its source.
What changed
The headline change. Z.ai trained it for months on long-horizon coding-agent work so the full window stays usable, not just addressable.
Terminal Bench 2.1 (Terminus-2) 63.5 to 81.0, FrontierSWE 30.5 to 74.4, DeepSWE 18 to 46.2, SWE-bench Pro 58.4 to 62.1, ProgramBench 50.9 to 63.7.
Humanity's Last Exam 31.0 to 40.5, CritPt 4.6 to 20.9, AIME 2026 95.3 to 99.2, GPQA-Diamond 86.2 to 91.2.
IndexShare cuts per-token compute 2.9 times at 1M tokens, and the improved multi-token-prediction layer speeds up speculative decoding.
The parameter count and the MIT license carry over, so a GLM-5.1 deployment can move to GLM-5.2 on the same hardware.
Benchmarks
Every score and every comparison model comes from Z.ai's own model card. These are vendor-reported results: useful for direction, not independent proof.
| Benchmark | GLM-5.2 | GLM-5.1 | Qwen3.7-Max | MiniMax M3 | DeepSeek-V4-Pro | Claude Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|---|---|---|
| Reasoning | ||||||||
| HLE | 40.5 | 31 | 41.4 | 37 | 37.7 | |||
Where to run it
Call it through Z.ai's API or a third-party host with any OpenAI-compatible SDK, use it inside a coding agent, or download the weights and serve them yourself.
Pricing
On Z.ai's own API it costs $1.40 per 1M input tokens and $4.40 per 1M output tokens (cached input $0.26). 26 hosts list it on OpenRouter, and 19 of them charge less than Z.ai. Sorted cheapest first by blended price (three parts input to one part output).
| Host | Input /1M | Output /1M | Cached input /1M | Precision | Context | Max output |
|---|---|---|---|---|---|---|
Baidu | $0.189 | $0.594 | $0.0351 | fp8 | 1,048,576 | 131,072 |
Decart | $0.2795 | $1.72 | $0.1075 | mxfp4 | 1,048,576 | 943,718 |
StreamLake | $0.56 | $1.76 | $0.104 | fp8 | 1,024,000 | 128,000 |
DeepInfra | $0.5625 | |||||
The GLM family
Every GLM model with its own page here, newest first, with Z.ai's current price per 1M input / output tokens.
| Model | Released | Parameters | Context | License | Z.ai price (in / out) |
|---|---|---|---|---|---|
GLM-5.3 Flash Cheap, fast, 1M-context multimodal (the model behind Ox Alpha) | Aug 2026 | 320B total / 18B active | 1M | MIT | $0.15 / $0.50 |
GLM-5.3 Z.ai's current flagship: the strongest GLM for coding and agents, 1M context | Aug 2026 | 753B total (GLM-5.2 base) | 1M | GLM-5.3 License | $1.40 / $4.40 |
GLM-5.2 The newest GLM flagship under plain MIT: 1M context, long-horizon coding | Jun 2026 | 744B / 40B active | 1M | MIT | $1.40 / $4.40 |
FAQ
You have seen the specs, the scores and the price. Turning a model into a product people use is the harder part, and it is the part you can skip wiring up yourself.
GLM-5.3 is built on the same base model with more post-training, and it scores higher on every benchmark in Z.ai's GLM-5.3 comparison. GLM-5.2 still matters for one reason: it is the newest GLM flagship under the plain MIT license, while GLM-5.3 ships under Z.ai's GLM-5.3 License, which adds a security-review condition for the very largest model-as-a-service companies. If license terms decide what you can deploy, GLM-5.2 is the strongest open GLM you can use without reading further. Otherwise GLM-5.3 is the better model.
| 49.8* |
| 41.4* |
| 45 |
| HLE (w/ Tools) | 54.7 | 52.3 | 53.5 | - | 48.2 | 57.9* | 52.2* | 51.4* |
|---|
| CritPt | 20.9 | 4.6 | 13.4 | 3.7 | 12.9 | 20.9 | 27.1 | 17.7 |
|---|
| AIME 2026 | 99.2 | 95.3 | 97 | - | 94.6 | 95.7 | 98.3 | 98.2 |
|---|
| HMMT Nov. 2025 | 94.4 | 94 | 95 | 84.4 | 94.4 | 96.5 | 96.5 | 94.8 |
|---|
| HMMT Feb. 2026 | 92.5 | 82.6 | 97.1 | 84.4 | 95.2 | 96.7 | 96.7 | 87.3 |
|---|
| IMOAnswerBench | 91.0 | 83.8 | 90 | - | 89.8 | 83.5 | - | 81 |
|---|
| GPQA-Diamond | 91.2 | 86.2 | 90 | 93 | 90.1 | 93.6 | 93.6 | 94.3 |
|---|
| Coding | ||||||||
|---|---|---|---|---|---|---|---|---|
| SWE-bench Pro | 62.1 | 58.4 | 60.6 | 59 | 55.4 | 69.2 | 58.6 | 54.2 |
| NL2Repo | 48.9 | 42.7 | 47.2 | 42.1 | 35.5 | 69.7 | 50.7 | 33.4 |
|---|
| DeepSWE | 46.2 | 18 | 18 | 20 | 8 | 58 | 70 | 10 |
|---|
| ProgramBench | 63.7 | 50.9 | - | - | 47.8 | 71.9 | 70.8 | 39.5 |
|---|
| Terminal Bench 2.1 (Terminus-2) | 81.0 | 63.5 | 75 | 65 | 64 | 85 | 84 | 74 |
|---|
| Terminal Bench 2.1 (best reported harness) | 82.7 | 69 | - | - | - | 78.9 | 83.4 | 70.7 |
|---|
| FrontierSWE (Dominance) | 74.4 | 30.5 | - | - | 29.0 | 75.1 | 72.6 | 39.6 |
|---|
| PostTrainBench | 34.3 | 20.1 | - | - | - | 37.2 | 28.4 | 21.6 |
|---|
| SWE-Marathon | 13.0 | 1.0 | - | - | - | 26.0 | 12.0 | 4.0 |
|---|
| Agentic | ||||||||
|---|---|---|---|---|---|---|---|---|
| MCP-Atlas (Public Set) | 76.8 | 71.8 | 76.4 | 74.2 | 73.6 | 77.8 | 75.3 | 69.2 |
| Tool-Decathlon | 48.2 | 40.7 | - | - | 52.8 | 59.9 | 55.6 | 48.8 |
|---|
Z.ai's own evaluations, except FrontierSWE (Proximal), PostTrainBench (PostTrainBench) and SWE-Marathon (Abundant AI), which outside evaluators ran at 1M context. HLE is the text-only subset; * marks a full-set score. Model card
Prices in USD per 1M tokens from OpenRouter's public API, fetched October 8, 2026 at 23:13 UTC and refreshed every 6 hours.
© 2026 Ox Alpha. Independent of Z.ai.
Facts sourced from model cards, public reporting and community forensics. Every benchmark is cited to its source.