Z.ai GLM model, released August 14, 2026
Released August 14, 2026, with weights on Hugging Face from August 28: the GLM-5.2 base model with far more post-training, a 1M-token context, and a new license. The specs, Z.ai's benchmarks, today's prices, and what the license means.
GLM-5.3 is Z.ai's current flagship. It uses the same base model as GLM-5.2, and Z.ai says every gain comes from post-training: GLM-5.3 beats GLM-5.2 on every benchmark in Z.ai's comparison, from Terminal Bench 2.1 (88.2 against 81.0) and DeepSWE (66.9 against 46.2) to CyberGym, where its 84.5 is the top score in the table. Z.ai calls it the most capable open-weights model for coding.
It keeps the 1M-token context and adds a reasoning_effort control with three levels: low, high, and max, the default. Z.ai also reports a cyber capability that grew faster than it expected: on exploitation benchmarks GLM-5.3 more than doubles GLM-5.2. The name covers a small family: GLM-5.3 Prime is a faster-serving version of the same model, and GLM-5.3 Flash is a separate, smaller multimodal model under MIT, the one that ran anonymously as Ox Alpha.
An independent reference, not affiliated with Z.ai. Checked against Z.ai's model card, docs and OpenRouter on October 4, 2026.
1M
Context window (tokens)
88.2
Terminal Bench 2.1 (Z.ai)
84.5
CyberGym, top of Z.ai's table
Open
Weights, GLM-5.3 License
Specifications
From Z.ai's model card and the published config, with each row linked to its source.
What changed
Terminal Bench 3.0 rises from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, SWE-Marathon from 19.4 to 42.5, and Z.ai reports a 50% gain on its in-house Z.ai Code Bench.
CyberGym 77.2 to 84.5, and exploitation benchmarks more than double: ExploitGym 29 to 105 solved in two hours, ExploitBench 24.4 to 54.4.
AutomationBench 26.2 to 48.2, Toolathlon Verified 59.9 to 73.0, HLE with tools 54.7 to 62.5.
No new pre-training: the 1M-token context and the parameter count carry over from GLM-5.2.
GLM-5.2 is plain MIT; GLM-5.3 adds the $10B model-as-a-service security-review condition described above.
Benchmarks
Every score and every comparison model comes from Z.ai's own model card. These are vendor-reported results: useful for direction, not independent proof.
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | DeepSeek-V4 Pro-0813 | Qwen3.8-Max | Opus 4.8 | Fable 5 (w/ fallback) | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|---|
| Coding | ||||||||
| Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 87.9 | ||||
Where to run it
Call it through Z.ai's API or a third-party host with any OpenAI-compatible SDK, use it inside a coding agent, or download the weights and serve them yourself.
Pricing
On Z.ai's own API it costs $1.40 per 1M input tokens and $4.40 per 1M output tokens (cached input $0.26). 34 hosts list it on OpenRouter, and 22 of them charge less than Z.ai. Sorted cheapest first by blended price (three parts input to one part output).
| Host | Input /1M | Output /1M | Cached input /1M | Precision | Context | Max output |
|---|---|---|---|---|---|---|
Wafer | $0.039 | $3.39 | $0.038 | not stated | 1,048,576 | 943,718 |
Wafer | $0.17 | $3.39 | $0.16 | not stated | 1,048,576 | 943,718 |
Sail Research | $0.20 | $3.40 | $0.15 | fp8 | 1,048,576 | 943,718 |
DeepInfra | $0.5625 | |||||
The GLM family
Every GLM model with its own page here, newest first, with Z.ai's current price per 1M input / output tokens.
| Model | Released | Parameters | Context | License | Z.ai price (in / out) |
|---|---|---|---|---|---|
GLM-5.3 Flash Cheap, fast, 1M-context multimodal (the model behind Ox Alpha) | Aug 2026 | 320B total / 18B active | 1M | MIT | $0.15 / $0.50 |
GLM-5.3 Z.ai's current flagship: the strongest GLM for coding and agents, 1M context | Aug 2026 | 753B total (GLM-5.2 base) | 1M | GLM-5.3 License | $1.40 / $4.40 |
GLM-5.2 The newest GLM flagship under plain MIT: 1M context, long-horizon coding | Jun 2026 | 744B / 40B active | 1M | MIT | $1.40 / $4.40 |
FAQ
You have seen the specs, the scores and the price. Turning a model into a product people use is the harder part, and it is the part you can skip wiring up yourself.
For most teams GLM-5.3 is the GLM to use today: the strongest results Z.ai has published and the same 1M context as GLM-5.2. The one thing to check is the license. The GLM-5.3 License grants MIT-style rights with one added condition: a company that runs a model-as-a-service business and has more than $10 billion in revenue over any 12 months must pass Z.ai's security review before commercial use. Nearly everyone falls outside that. If you need plain MIT terms, GLM-5.2 is the newest flagship that has them.
| 86.6 |
| 85.0 |
| 88.0 |
| 88.8 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | - | - | 21.1 | 33.7 | 34.6 |
|---|
| DeepSWE (v1.1) | 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | 72.7 |
|---|
| NL2Repo | 58.0 | 48.9 | 58.0 | 61.1 | 55.9 | 69.7 | - | - |
|---|
| ProgramBench (Almost Solved) | 19.0 | 9.5 | 17.5 | - | 10.5 | 15.5 | 33.0 | 23.0 |
|---|
| FrontierSWE | 78.1 | 67.5 | - | - | - | 66.5 | 88.2 | - |
|---|
| SWE-Marathon (v1.1) | 42.5 | 19.4 | 48.1 | - | - | 48.8 | 33.1 | 42.5 |
|---|
| PostTrainBench | 39.8 | 31.7 | 32.0 | - | - | 32.9 | 41.8 | 36.2 |
|---|
| Security | ||||||||
|---|---|---|---|---|---|---|---|---|
| CyberGym | 84.5 | 77.2 | 80.0 | 83.3 | 78.5 | 78.1 | 83.8 | 83.6 |
| ExploitGym (2h / 6h) | 105 / 130 | 29 / 39 | 36 / 70 | - | 14 / 26 | 80 / 120 | 181 / 247 | 216 / 293 |
|---|
| ExploitBench | 54.4 | 24.4 | 32.2 | - | 28.8 | 40.0 | 78.0 | 76.5 |
|---|
| Agents and knowledge work | ||||||||
|---|---|---|---|---|---|---|---|---|
| Toolathlon Verified | 73.0 | 59.9 | 76.5 | 74.1 | 72.5 | 76.2 | 74.7 | 74.9 |
| AutomationBench (v1.0.6) | 48.2 | 26.2 | 46.7 | 43.2 | 39.8 | 41.0 | 46.2 | 45.8 |
|---|
| Agents' Last Exam (ALE-CLI) | 28.5 | 23.8 | 27.6 | 25.7 | 27.0 | 25.7 | 23.8 | 28.6 |
|---|
| HLE w/ Tools | 62.5 | 54.7 | 59.8 | 60.0 | 56.2 | 57.9 | 63.9 | 64.5 |
|---|
| GDPval-AA v2 | 1769 | 1508 | 1682 | 1590 | 1739 | 1588 | 1743 | 1730 |
|---|
Z.ai's own evaluations, mostly in Claude Code at max reasoning effort. FrontierSWE was run by Proximal and GDPval-AA by Artificial Analysis. ExploitGym counts tasks solved within 2 and 6 hours of rescaled inference time. Some GLM-5.2 scores here differ from GLM-5.2's own card because the benchmark versions and dates changed (FrontierSWE, for one). Model card
Prices in USD per 1M tokens from OpenRouter's public API, fetched October 8, 2026 at 23:13 UTC and refreshed every 6 hours.
© 2026 Ox Alpha. Independent of Z.ai.
Facts sourced from model cards, public reporting and community forensics. Every benchmark is cited to its source.