Z.ai GLM model, released February 11, 2026
Released February 11, 2026 under the MIT license: the model that doubled GLM's scale to 744B parameters (40B active) and opened the GLM-5 generation. The specs, Z.ai's benchmarks, today's prices, and how to run it.
GLM-5 opened the GLM-5 generation. Compared with GLM-4.5 it doubles the model, from 355B parameters (32B active) to 744B (40B active), and raises pre-training data from 23 trillion to 28.5 trillion tokens. It adopts DeepSeek Sparse Attention (DSA), which Z.ai says largely cuts deployment cost while keeping long-context capacity, and it was post-trained with slime, Z.ai's asynchronous reinforcement-learning system.
Z.ai aimed it at complex systems engineering and long-horizon agent tasks, and its own evaluations show big gains over GLM-4.7: SWE-bench Verified 77.8 (from 73.8), Terminal-Bench 2.0 56.2 (from 41.0), MCP-Atlas 67.8 (from 52.0). Before launch, Z.ai ran it anonymously on OpenRouter as "Pony Alpha", the same playbook it later used for Ox Alpha.
An independent reference, not affiliated with Z.ai. Checked against Z.ai's model card, docs and OpenRouter on October 4, 2026.
744B
Total parameters, 40B active
28.5T
Pre-training tokens
77.8
SWE-bench Verified (Z.ai)
MIT
License, open weights
GLM-5 is the base of everything after it: GLM-5.1, GLM-5.2 and GLM-5.3 share its architecture and size, and each scores higher in Z.ai's evaluations. Its context window stops at 200K tokens, where GLM-5.2 and GLM-5.3 reach 1M. For a new self-hosted deployment, GLM-5.2 runs on the same hardware with more capability under the same MIT license; GLM-5 mainly remains useful where it is already deployed and tested.
Specifications
From Z.ai's model card and the published config, with each row linked to its source.
What changed
355B to 744B parameters, 32B to 40B active, and 23T to 28.5T pre-training tokens.
DeepSeek Sparse Attention replaces dense attention, which Z.ai says largely reduces deployment cost at long context.
SWE-bench Verified 73.8 to 77.8, SWE-bench Multilingual 66.7 to 73.3, Terminal-Bench 2.0 41.0 to 56.2.
MCP-Atlas 52.0 to 67.8, BrowseComp 52.0 to 62.0, Tool-Decathlon 23.8 to 38.0, and Vending Bench 2 nearly doubles ($2,376.82 to $4,432.12).
Benchmarks
Every score and every comparison model comes from Z.ai's own model card. These are vendor-reported results: useful for direction, not independent proof.
| Benchmark | GLM-5 | GLM-4.7 | DeepSeek-V3.2 | Kimi K2.5 | Claude Opus 4.5 | Gemini 3 Pro | GPT-5.2 (xhigh) |
|---|---|---|---|---|---|---|---|
| Reasoning | |||||||
| HLE | 30.5 | 24.8 | 25.1 | 31.5 | 28.4 | 37.2 | |
Where to run it
Call it through Z.ai's API or a third-party host with any OpenAI-compatible SDK, use it inside a coding agent, or download the weights and serve them yourself.
Pricing
On Z.ai's own API it costs $1.00 per 1M input tokens and $3.20 per 1M output tokens (cached input $0.20). 8 hosts list it on OpenRouter, and 4 of them charge less than Z.ai. Sorted cheapest first by blended price (three parts input to one part output).
| Host | Input /1M | Output /1M | Cached input /1M | Precision | Context | Max output |
|---|---|---|---|---|---|---|
GMICloud | $0.60 | $1.92 | $0.12 | fp8 | 202,752 | 182,476 |
StreamLake | $0.60 | $1.92 | $0.12 | fp8 | 198,000 | 128,000 |
Baidu | $0.70 | $2.24 | $0.14 | fp8 | 202,752 | 131,072 |
SiliconFlow | $0.95 | $2.55 | ||||
The GLM family
Every GLM model with its own page here, newest first, with Z.ai's current price per 1M input / output tokens.
| Model | Released | Parameters | Context | License | Z.ai price (in / out) |
|---|---|---|---|---|---|
GLM-5.3 Flash Cheap, fast, 1M-context multimodal (the model behind Ox Alpha) | Aug 2026 | 320B total / 18B active | 1M | MIT | $0.15 / $0.50 |
GLM-5.3 Z.ai's current flagship: the strongest GLM for coding and agents, 1M context | Aug 2026 | 753B total (GLM-5.2 base) | 1M | GLM-5.3 License | $1.40 / $4.40 |
GLM-5.2 The newest GLM flagship under plain MIT: 1M context, long-horizon coding | Jun 2026 | 744B / 40B active | 1M | MIT | $1.40 / $4.40 |
FAQ
You have seen the specs, the scores and the price. Turning a model into a product people use is the harder part, and it is the part you can skip wiring up yourself.
| 35.4 |
| HLE (w/ Tools) | 50.4 | 42.8 | 40.8 | 51.8 | 43.4* | 45.8* | 45.5* |
|---|
| AIME 2026 I | 92.7 | 92.9 | 92.7 | 92.5 | 93.3 | 90.6 | - |
|---|
| HMMT Nov. 2025 | 96.9 | 93.5 | 90.2 | 91.1 | 91.7 | 93.0 | 97.1 |
|---|
| IMOAnswerBench | 82.5 | 82.0 | 78.3 | 81.8 | 78.5 | 83.3 | 86.3 |
|---|
| GPQA-Diamond | 86.0 | 85.7 | 82.4 | 87.6 | 87.0 | 91.9 | 92.4 |
|---|
| Coding | |||||||
|---|---|---|---|---|---|---|---|
| SWE-bench Verified | 77.8 | 73.8 | 73.1 | 76.8 | 80.9 | 76.2 | 80.0 |
| SWE-bench Multilingual | 73.3 | 66.7 | 70.2 | 73.0 | 77.5 | 65.0 | 72.0 |
|---|
| Terminal-Bench 2.0 (Terminus 2) | 56.2 / 60.7 † | 41.0 | 39.3 | 50.8 | 59.3 | 54.2 | 54.0 |
|---|
| Terminal-Bench 2.0 (Claude Code) | 56.2 / 61.1 † | 32.8 | 46.4 | - | 57.9 | - | - |
|---|
| CyberGym | 43.2 | 23.5 | 17.3 | 41.3 | 50.6 | 39.9 | - |
|---|
| Agentic | |||||||
|---|---|---|---|---|---|---|---|
| BrowseComp | 62.0 | 52.0 | 51.4 | 60.6 | 37.0 | 37.8 | - |
| BrowseComp (w/ Context Manage) | 75.9 | 67.5 | 67.6 | 74.9 | 67.8 | 59.2 | 65.8 |
|---|
| BrowseComp-Zh | 72.7 | 66.6 | 65.0 | 62.3 | 62.4 | 66.8 | 76.1 |
|---|
| τ²-Bench | 89.7 | 87.4 | 85.3 | 80.2 | 91.6 | 90.7 | 85.5 |
|---|
| MCP-Atlas (Public Set) | 67.8 | 52.0 | 62.2 | 63.8 | 65.2 | 66.6 | 68.0 |
|---|
| Tool-Decathlon | 38.0 | 23.8 | 35.2 | 27.8 | 43.5 | 36.4 | 46.3 |
|---|
| Vending Bench 2 | $4,432.12 | $2,376.82 | $1,034.00 | $1,198.46 | $4,967.06 | $5,478.16 | $3,591.33 |
|---|
Z.ai's own evaluations as published on the GLM-5 model card. * marks a full-set score; † the second figure is on Z.ai's verified Terminal-Bench 2.0, which fixes ambiguous instructions. Andon Labs ran Vending Bench 2 independently. Model card
Prices in USD per 1M tokens from OpenRouter's public API, fetched October 8, 2026 at 23:13 UTC and refreshed every 6 hours.
© 2026 Ox Alpha. Independent of Z.ai.
Facts sourced from model cards, public reporting and community forensics. Every benchmark is cited to its source.