Z.ai GLM model, released December 22, 2025
Released December 22, 2025 under the MIT license: a 355B-parameter (32B active) model with a 200K context, built as a coding partner. The specs, Z.ai's benchmarks, today's prices, and how it compares with GLM-4.6 and GLM-5.
GLM-4.7 is the last flagship built on the GLM-4.5 architecture (about 355B parameters, 32B active, 200K context). Z.ai pitched it as "your new coding partner": against GLM-4.6 it gains 5.8 points on SWE-bench Verified (73.8), 12.9 on SWE-bench Multilingual (66.7) and 16.5 on Terminal Bench 2.0 (41.0), and it scores 42.8 on Humanity's Last Exam with tools, up 12.4.
It also changed how the model thinks inside agents. Interleaved Thinking (thinking before every reply and tool call) was joined by Preserved Thinking, which keeps earlier reasoning blocks across turns in coding agents instead of re-deriving them, and Turn-level Thinking, which switches reasoning on or off per turn to trade accuracy for latency and cost. Z.ai also highlighted cleaner generated web pages and slides.
An independent reference, not affiliated with Z.ai. Checked against Z.ai's model card, docs and OpenRouter on October 4, 2026.
73.8
SWE-bench Verified (Z.ai)
355B
Total parameters, 32B active
200K
Context window (tokens)
MIT
License, open weights
GLM-5 replaced GLM-4.7 as Z.ai's flagship seven weeks later, with twice the parameters and higher scores on nearly every benchmark Z.ai compared. GLM-4.7 is still worth knowing for two reasons: at 355B it needs about half the hardware of a GLM-5 model to self-host (362 GB in FP8 against 756 GB), and its small sibling GLM-4.7-Flash is free on Z.ai's own API. For the best GLM results today, use GLM-5.3 or GLM-5.2.
Specifications
From Z.ai's model card and the published config, with each row linked to its source.
What changed
SWE-bench Verified 68.0 to 73.8, SWE-bench Multilingual 53.8 to 66.7, Terminal Bench 2.0 24.5 to 41.0, LiveCodeBench v6 82.8 to 84.9.
Humanity's Last Exam 17.2 to 24.8 (30.4 to 42.8 with tools), HMMT February 2025 89.2 to 97.1, IMOAnswerBench 73.5 to 82.0.
τ²-Bench 75.2 to 87.4, BrowseComp 45.1 to 52.0.
Preserved Thinking keeps reasoning across turns in coding agents, and Turn-level Thinking lets you switch it off for quick requests.
Z.ai calls it vibe coding: cleaner, more modern web pages and slides with more accurate layout and sizing.
Benchmarks
Every score and every comparison model comes from Z.ai's own model card. These are vendor-reported results: useful for direction, not independent proof.
| Benchmark | GLM-4.7 | GLM-4.6 | Kimi K2 Thinking | DeepSeek-V3.2 | Gemini 3.0 Pro | Claude Sonnet 4.5 | GPT-5-High | GPT-5.1-High |
|---|---|---|---|---|---|---|---|---|
| Reasoning | ||||||||
| MMLU-Pro | 84.3 | 83.2 | 84.6 | 85.0 | ||||
Where to run it
Call it through Z.ai's API or a third-party host with any OpenAI-compatible SDK, use it inside a coding agent, or download the weights and serve them yourself.
Pricing
On Z.ai's own API it costs $0.60 per 1M input tokens and $2.20 per 1M output tokens (cached input $0.11). 6 hosts list it on OpenRouter, and 3 of them charge less than Z.ai. Sorted cheapest first by blended price (three parts input to one part output).
| Host | Input /1M | Output /1M | Cached input /1M | Precision | Context | Max output |
|---|---|---|---|---|---|---|
DeepInfra | $0.40 | $1.75 | $0.08 | fp4 | 202,752 | 131,072 |
Venice | $0.4004 | $1.9292 | $0.0801 | fp4 | 198,000 | 16,384 |
Novita | $0.54 | $1.98 | $0.099 | fp8 | 204,800 | 131,072 |
Google | $0.60 | $2.20 | ||||
The GLM family
Every GLM model with its own page here, newest first, with Z.ai's current price per 1M input / output tokens.
| Model | Released | Parameters | Context | License | Z.ai price (in / out) |
|---|---|---|---|---|---|
GLM-5.3 Flash Cheap, fast, 1M-context multimodal (the model behind Ox Alpha) | Aug 2026 | 320B total / 18B active | 1M | MIT | $0.15 / $0.50 |
GLM-5.3 Z.ai's current flagship: the strongest GLM for coding and agents, 1M context | Aug 2026 | 753B total (GLM-5.2 base) | 1M | GLM-5.3 License | $1.40 / $4.40 |
GLM-5.2 The newest GLM flagship under plain MIT: 1M context, long-horizon coding | Jun 2026 | 744B / 40B active | 1M | MIT | $1.40 / $4.40 |
FAQ
You have seen the specs, the scores and the price. Turning a model into a product people use is the harder part, and it is the part you can skip wiring up yourself.
| 90.1 |
| 88.2 |
| 87.5 |
| 87.0 |
| GPQA-Diamond | 85.7 | 81.0 | 84.5 | 82.4 | 91.9 | 83.4 | 85.7 | 88.1 |
|---|
| HLE | 24.8 | 17.2 | 23.9 | 25.1 | 37.5 | 13.7 | 26.3 | 25.7 |
|---|
| HLE (w/ Tools) | 42.8 | 30.4 | 44.9 | 40.8 | 45.8 | 32.0 | 35.2 | 42.7 |
|---|
| AIME 2025 | 95.7 | 93.9 | 94.5 | 93.1 | 95.0 | 87.0 | 94.6 | 94.0 |
|---|
| HMMT Feb. 2025 | 97.1 | 89.2 | 89.4 | 92.5 | 97.5 | 79.2 | 88.3 | 96.3 |
|---|
| HMMT Nov. 2025 | 93.5 | 87.7 | 89.2 | 90.2 | 93.3 | 81.7 | 89.2 | - |
|---|
| IMOAnswerBench | 82.0 | 73.5 | 78.6 | 78.3 | 83.3 | 65.8 | 76.0 | - |
|---|
| Coding | ||||||||
|---|---|---|---|---|---|---|---|---|
| LiveCodeBench-v6 | 84.9 | 82.8 | 83.1 | 83.3 | 90.7 | 64.0 | 87.0 | 87.0 |
| SWE-bench Verified | 73.8 | 68.0 | 71.3 | 73.1 | 76.2 | 77.2 | 74.9 | 76.3 |
|---|
| SWE-bench Multilingual | 66.7 | 53.8 | 61.1 | 70.2 | - | 68.0 | 55.3 | - |
|---|
| Terminal Bench Hard | 33.3 | 23.6 | 30.6 | 35.4 | 39.0 | 33.3 | 30.5 | 43.0 |
|---|
| Terminal Bench 2.0 | 41.0 | 24.5 | 35.7 | 46.4 | 54.2 | 42.8 | 35.2 | 47.6 |
|---|
| Agentic | ||||||||
|---|---|---|---|---|---|---|---|---|
| BrowseComp | 52.0 | 45.1 | - | 51.4 | - | 24.1 | 54.9 | 50.8 |
| BrowseComp (w/ Context Manage) | 67.5 | 57.5 | 60.2 | 67.6 | 59.2 | - | - | - |
|---|
| BrowseComp-Zh | 66.6 | 49.5 | 62.3 | 65.0 | - | 42.4 | 63.0 | - |
|---|
| τ²-Bench | 87.4 | 75.2 | 74.3 | 85.3 | 90.7 | 87.2 | 82.4 | 82.7 |
|---|
Z.ai's own evaluations as published on the GLM-4.7 model card: temperature 1.0 and up to 131,072 new tokens by default, Terminal Bench and SWE-bench Verified at temperature 0.7 with 16,384 tokens, and τ²-Bench at temperature 0. Z.ai recommends Preserved Thinking for τ²-Bench and Terminal Bench 2.0. Model card
Prices in USD per 1M tokens from OpenRouter's public API, fetched October 8, 2026 at 23:13 UTC and refreshed every 6 hours.
© 2026 Ox Alpha. Independent of Z.ai.
Facts sourced from model cards, public reporting and community forensics. Every benchmark is cited to its source.