Z.ai GLM model, released January 19, 2026
Released January 19, 2026 under the MIT license: a 30B-parameter Mixture-of-Experts model with about 3B active and a 200K context, free on Z.ai's own API. The specs, Z.ai's benchmarks, and how to run it on one machine.
GLM-4.7-Flash is the lightweight member of the GLM-4.7 series: a 30B-A3B Mixture-of-Experts model (about 30B parameters, roughly 3B active per token) that Z.ai calls the strongest model in the 30B class. It keeps GLM-4.7's 200K context and 128K maximum output, and it is tuned for agentic coding, long-horizon planning and tool use.
In Z.ai's comparison it far outscores models of its size on agent and coding work: 59.2 on SWE-bench Verified against 22.0 for Qwen3-30B-A3B-Thinking-2507 and 34.0 for GPT-OSS-20B, and 79.5 on τ²-Bench against 49.0 and 47.7. On contest maths and code the gap is small: GPT-OSS-20B edges it on AIME 2025 and Qwen3-30B on LiveCodeBench.
An independent reference, not affiliated with Z.ai. Checked against Z.ai's model card, docs and OpenRouter on October 4, 2026.
30B
Total parameters, ~3B active
59.2
SWE-bench Verified (Z.ai)
62 GB
BF16 weights
Free
On Z.ai's own API
GLM-4.7-Flash is the cheapest way into the GLM family: free through Z.ai's own API, and small enough to self-host on a single machine (62 GB of BF16 weights, less with community GGUF quantizations). It is not a substitute for GLM-5: for hard coding and agent work its larger siblings are far ahead. Z.ai also sells a faster paid variant, GLM-4.7-FlashX.
Specifications
From Z.ai's model card and the published config, with each row linked to its source.
What changed
SWE-bench Verified 59.2, against 22.0 for Qwen3-30B-A3B-Thinking-2507 and 34.0 for GPT-OSS-20B.
τ²-Bench 79.5 (49.0 and 47.7) and BrowseComp 42.8 (2.29 and 28.3).
Humanity's Last Exam 14.4 against 9.8 and 10.9, and GPQA 75.2 against 73.4 and 71.5.
AIME 2025 91.6 against GPT-OSS-20B's 91.7, and LiveCodeBench v6 64.0 against Qwen3-30B's 66.0.
Benchmarks
Every score and every comparison model comes from Z.ai's own model card. These are vendor-reported results: useful for direction, not independent proof.
| Benchmark | GLM-4.7-Flash | Qwen3-30B-A3B-Thinking-2507 | GPT-OSS-20B |
|---|---|---|---|
| AIME 25 | 91.6 | 85.0 | 91.7 |
| GPQA | 75.2 | 73.4 | 71.5 |
| LCB v6 |
Where to run it
Call it through Z.ai's API or a third-party host with any OpenAI-compatible SDK, use it inside a coding agent, or download the weights and serve them yourself.
Pricing
Z.ai's own API serves it free; Z.ai does not sell it through OpenRouter. 3 hosts list it on OpenRouter. Sorted cheapest first by blended price (three parts input to one part output).
| Host | Input /1M | Output /1M | Cached input /1M | Precision | Context | Max output |
|---|---|---|---|---|---|---|
Venice | $0.06 | $0.40 | $0.01 | fp8 | 128,000 | 16,384 |
Cloudflare | $0.0605 | $0.40 | n/a | not stated | 131,072 | 117,964 |
Novita | $0.07 | $0.40 | $0.01 | bf16 | 200,000 | 128,000 |
Prices in USD per 1M tokens from OpenRouter's public API, fetched October 8, 2026 at 23:13 UTC and refreshed every 6 hours.
The GLM family
Every GLM model with its own page here, newest first, with Z.ai's current price per 1M input / output tokens.
| Model | Released | Parameters | Context | License | Z.ai price (in / out) |
|---|---|---|---|---|---|
GLM-5.3 Flash Cheap, fast, 1M-context multimodal (the model behind Ox Alpha) | Aug 2026 | 320B total / 18B active | 1M | MIT | $0.15 / $0.50 |
GLM-5.3 Z.ai's current flagship: the strongest GLM for coding and agents, 1M context | Aug 2026 | 753B total (GLM-5.2 base) | 1M | GLM-5.3 License | $1.40 / $4.40 |
GLM-5.2 The newest GLM flagship under plain MIT: 1M context, long-horizon coding | Jun 2026 | 744B / 40B active | 1M | MIT | $1.40 / $4.40 |
FAQ
You have seen the specs, the scores and the price. Turning a model into a product people use is the harder part, and it is the part you can skip wiring up yourself.
| 64.0 |
| 66.0 |
| 61.0 |
| HLE | 14.4 | 9.8 | 10.9 |
|---|
| SWE-bench Verified | 59.2 | 22.0 | 34.0 |
|---|
| τ²-Bench | 79.5 | 49.0 | 47.7 |
|---|
| BrowseComp | 42.8 | 2.29 | 28.3 |
|---|
Z.ai's own evaluations as published on the GLM-4.7-Flash model card: temperature 1.0 by default, with SWE-bench Verified at 0.7 and τ²-Bench at 0. Model card
© 2026 Ox Alpha. Independent of Z.ai.
Facts sourced from model cards, public reporting and community forensics. Every benchmark is cited to its source.