Z.ai GLM-5.3 Flash
Z.ai's open-weight (MIT), 320B-A18B multimodal Mixture-of-Experts with a 1M-token context. The specs, the benchmarks, the pricing, the origin story, and where to run it now.
GLM-5.3 Flash is Z.ai's (Zhipu AI's) fast, cheap, natively multimodal model: a 320B-total, 18B-active Mixture-of-Experts with a 1,048,576-token (1M) context window and text, image, and video input. In late August 2026 it stepped out of stealth: it was the anonymous “Ox Alpha” model that developers had wired straight into their coding agents while nobody knew who built it, exactly where the community's fingerprinting had pointed.
Then Z.ai went further and published the weights on Hugging Face under a plain, unmodified MIT License, so anyone can now download, self-host, fine-tune, and commercially deploy the same model. Below is the honest spec sheet, the current benchmark picture, where to run it, and the origin story. Last verified September 4, 2026. It first appeared, free, on August 20, 2026.
1M
token context window
131K
max output tokens
$0
during the stealth preview
#2
on OpenCode at launch
Specifications
Confirmed against Z.ai's open-weight release: the architecture that used to be third-party guesswork is now published. Every row carries its confidence and links to its source.
Benchmarks
Post-reveal, a coherent picture replaced the messy pre-reveal community evals: a high-value coder that trails the leaders on raw agentic coding but wins on price and stands out on document vision.
57
Artificial Analysis Intelligence Index v4.1.1
63.4
DeepSWE v1.1 (coding / agentic SWE)
84.3
Terminal-Bench 2.1 (GPT-5.6 Terra 87.4, Gemini 3.7 Flash 85.8)
78.0
Chartography with Tools (vision, leads Gemini's 65.0)
Where to run it
Two real paths: call an OpenAI-compatible API at cents per million tokens (Z.ai's own, or any of the hosts on OpenRouter), or self-host the MIT weights for the cost of your own GPUs. There is no free API tier of the model itself. For the self-hosting side, FastGPU, a GPU pricing site run by the same founder as Ox Alpha, has a model-agnostic breakdown of what it costs to run an LLM in 2026.
Pricing
Z.ai lists it at $0.15 per 1M input tokens and $0.50 per 1M output tokens. Because the weights are MIT-licensed, anyone can serve it: 29 hosts list it on OpenRouter, and 11 of them charge less than Z.ai. Sorted cheapest first by blended price (three parts input to one part output, the convention Artificial Analysis uses).
| Host | Input /1M | Output /1M | Cached input /1M | Precision | Context | Max output |
|---|---|---|---|---|---|---|
DeepInfra | $0.075 | $0.25 | $0.015 | fp4 | 1,048,576 | 131,072 |
Novita | $0.084 | $0.28 | $0.0168 | fp8 | 1,048,576 | 131,072 |
StreamLake | $0.087 | $0.29 | $0.0174 | fp8 | 1,024,000 | 128,000 |
GMICloud | $0.09 | |||||
FAQ
You have seen the specs, the scores, and the price. The next step is to build something real with it, without wiring up any of the infrastructure yourself.
Prices in USD per 1M tokens from OpenRouter's public API, fetched October 8, 2026 at 23:13 UTC and refreshed every 6 hours.
© 2026 Ox Alpha. Independent of Z.ai.
Facts sourced from model cards, public reporting and community forensics. Every benchmark is cited to its source.