Z.ai GLM model, released April 2026
Released in April 2026 under the MIT license: GLM-5's 744B architecture retrained for long-running agent work, with a 200K context. The specs, Z.ai's benchmarks, today's prices, and how it compares with GLM-5.2.
GLM-5.1 is the first update to GLM-5. It keeps the same architecture (744B parameters, 40B active, 200K context) and focuses on what Z.ai calls agentic engineering: staying productive over hundreds of rounds and thousands of tool calls instead of plateauing after the first pass. In Z.ai's evaluations it topped the comparison on SWE-Bench Pro (58.4) and pulled well ahead of GLM-5 on repository generation (NL2Repo 42.7 against 35.9) and terminal tasks (Terminal-Bench 2.0 63.5 against 56.2).
Its standout result is security work: 68.7 on CyberGym, up from GLM-5's 48.3 and ahead of every other model in Z.ai's table. Like GLM-5, its weights are on Hugging Face under the MIT license, in BF16 and FP8.
An independent reference, not affiliated with Z.ai. Checked against Z.ai's model card, docs and OpenRouter on October 4, 2026.
58.4
SWE-Bench Pro (Z.ai)
68.7
CyberGym (Z.ai)
200K
Context window (tokens)
MIT
License, open weights
GLM-5.1 has been overtaken twice: GLM-5.2 (June 2026) adds the 1M-token context and much stronger coding on the same architecture and license, and GLM-5.3 (August 2026) improves again. There is little reason to start a new project on GLM-5.1; the main reason to keep it is an existing deployment you have tested and do not want to move. If you self-host, GLM-5.2 needs the same hardware and scores higher on every coding benchmark Z.ai published for both.
Specifications
From Z.ai's model card and the published config, with each row linked to its source.
What changed
Z.ai says earlier models, GLM-5 included, apply familiar techniques early and then plateau; GLM-5.1 keeps revising its strategy and gets better the longer it runs.
SWE-Bench Pro 55.1 to 58.4, NL2Repo 35.9 to 42.7, Terminal-Bench 2.0 (Terminus-2) 56.2 to 63.5, and 69.0 in Claude Code.
CyberGym 48.3 to 68.7, BrowseComp 62.0 to 68.0, MCP-Atlas 69.2 to 71.8.
No change to the 744B architecture or the 200K window; the 1M context arrived with GLM-5.2.
Benchmarks
Every score and every comparison model comes from Z.ai's own model card. These are vendor-reported results: useful for direction, not independent proof.
| Benchmark | GLM-5.1 | GLM-5 | Qwen3.6-Plus | MiniMax M2.7 | DeepSeek-V3.2 | Kimi K2.5 | Claude Opus 4.6 | Gemini 3.1 Pro | GPT-5.4 |
|---|---|---|---|---|---|---|---|---|---|
| Reasoning | |||||||||
| HLE | 31.0 | 30.5 | 28.8 | 28.0 | |||||
Where to run it
Call it through Z.ai's API or a third-party host with any OpenAI-compatible SDK, use it inside a coding agent, or download the weights and serve them yourself.
Pricing
On Z.ai's own API it costs $1.40 per 1M input tokens and $4.40 per 1M output tokens (cached input $0.26). 13 hosts list it on OpenRouter, and 8 of them charge less than Z.ai. Sorted cheapest first by blended price (three parts input to one part output).
| Host | Input /1M | Output /1M | Cached input /1M | Precision | Context | Max output |
|---|---|---|---|---|---|---|
Baidu | $0.9646 | $3.0316 | $0.1791 | fp8 | 202,752 | 131,072 |
StreamLake | $0.966 | $3.036 | $0.1794 | fp8 | 200,000 | 128,000 |
Chutes | $0.98 | $3.08 | $0.098 | fp8 | 202,752 | 65,535 |
SiliconFlow | $1.19 | |||||
The GLM family
Every GLM model with its own page here, newest first, with Z.ai's current price per 1M input / output tokens.
| Model | Released | Parameters | Context | License | Z.ai price (in / out) |
|---|---|---|---|---|---|
GLM-5.3 Flash Cheap, fast, 1M-context multimodal (the model behind Ox Alpha) | Aug 2026 | 320B total / 18B active | 1M | MIT | $0.15 / $0.50 |
GLM-5.3 Z.ai's current flagship: the strongest GLM for coding and agents, 1M context | Aug 2026 | 753B total (GLM-5.2 base) | 1M | GLM-5.3 License | $1.40 / $4.40 |
GLM-5.2 The newest GLM flagship under plain MIT: 1M context, long-horizon coding | Jun 2026 | 744B / 40B active | 1M | MIT | $1.40 / $4.40 |
GLM-5.1 | |||||
FAQ
You have seen the specs, the scores and the price. Turning a model into a product people use is the harder part, and it is the part you can skip wiring up yourself.
| 25.1 |
| 31.5 |
| 36.7 |
| 45.0 |
| 39.8 |
| HLE (w/ Tools) | 52.3 | 50.4 | 50.6 | - | 40.8 | 51.8 | 53.1* | 51.4* | 52.1* |
|---|
| AIME 2026 | 95.3 | 95.4 | 95.1 | 89.8 | 95.1 | 94.5 | 95.6 | 98.2 | 98.7 |
|---|
| HMMT Nov. 2025 | 94.0 | 96.9 | 94.6 | 81.0 | 90.2 | 91.1 | 96.3 | 94.8 | 95.8 |
|---|
| HMMT Feb. 2026 | 82.6 | 82.8 | 87.8 | 72.7 | 79.9 | 81.3 | 84.3 | 87.3 | 91.8 |
|---|
| IMOAnswerBench | 83.8 | 82.5 | 83.8 | 66.3 | 78.3 | 81.8 | 75.3 | 81.0 | 91.4 |
|---|
| GPQA-Diamond | 86.2 | 86.0 | 90.4 | 87.0 | 82.4 | 87.6 | 91.3 | 94.3 | 92.0 |
|---|
| Coding | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| SWE-Bench Pro | 58.4 | 55.1 | 56.6 | 56.2 | - | 53.8 | 57.3 | 54.2 | 57.7 |
| NL2Repo | 42.7 | 35.9 | 37.9 | 39.8 | - | 32.0 | 49.8 | 33.4 | 41.3 |
|---|
| Terminal-Bench 2.0 (Terminus-2) | 63.5 | 56.2 | 61.6 | - | 39.3 | 50.8 | 65.4 | 68.5 | - |
|---|
| Terminal-Bench 2.0 (best self-reported) | 69.0 (Claude Code) | 56.2 (Claude Code) | - | 57.0 (Claude Code) | 46.4 (Claude Code) | - | - | - | 75.1 (Codex) |
|---|
| CyberGym | 68.7 | 48.3 | - | - | 17.3 | 41.3 | 66.6 | 38.8 | 66.3 |
|---|
| Agentic | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| BrowseComp | 68.0 | 62.0 | - | - | 51.4 | 60.6 | - | - | - |
| BrowseComp (w/ Context Manage) | 79.3 | 75.9 | - | - | 67.6 | 74.9 | 84.0 | 85.9 | 82.7 |
|---|
| τ³-Bench | 70.6 | 69.2 | 70.7 | 67.6 | 69.2 | 66.0 | 72.4 | 67.1 | 72.9 |
|---|
| MCP-Atlas (Public Set) | 71.8 | 69.2 | 74.1 | 48.8 | 62.2 | 63.8 | 73.8 | 69.2 | 67.2 |
|---|
| Tool-Decathlon | 40.7 | 38.0 | 39.8 | 46.3 | 35.2 | 27.8 | 47.2 | 48.8 | 54.6 |
|---|
| Vending Bench 2 | $5,634.41 | $4,432.12 | $5,114.87 | - | $1,034.00 | $1,198.46 | $8,017.59 | $911.21 | $6,144.18 |
|---|
Z.ai's own evaluations as published on the GLM-5.1 model card. * marks a full-set score. Vending Bench 2 is the simulated business's final balance. Z.ai re-ran GLM-5 for this card, so a few GLM-5 scores here differ from the GLM-5 card (CyberGym, for one). Model card
Prices in USD per 1M tokens from OpenRouter's public API, fetched October 8, 2026 at 23:13 UTC and refreshed every 6 hours.
© 2026 Ox Alpha. Independent of Z.ai.
Facts sourced from model cards, public reporting and community forensics. Every benchmark is cited to its source.