GLM API pricing
Z.ai's own price next to the cheapest third-party host for each GLM model, in US dollars per 1M tokens. Pulled from OpenRouter's public API and refreshed every 6 hours, so these are today's prices, not a launch-day snapshot.
16
GLM models priced
14
sold by Z.ai on OpenRouter
29
hosts serving GLM-5.3 Flash
$0.039
lowest input price per 1M (GLM-5.3)
The table
Newest model first. The Z.ai column is Z.ai's own endpoint, which charges its published list price. The cheapest host is the lowest blended price (three parts input to one part output) among every host OpenRouter lists for that model. GLM-5.3 Flash, the model that ran as Ox Alpha, is highlighted.
| Model | Z.ai price (in / out) | Cheapest host (in / out) | Cheapest host | Hosts | Max context |
|---|---|---|---|---|---|
GLM-5.3 Prime | Not sold by Z.ai on OpenRouter | $2.80 / $8.80 | Alibaba precision not stated | 1 | 1,000,000 |
GLM-5.3 FlashX | $0.37 / $1.25 | $0.37 / $1.25 | |||
Reading the prices
Z.ai sets a list price for each model it sells. Because many GLM models have open weights, other companies serve the same models and set their own prices, often below Z.ai's.
Hosts state the number format they run. FP4 builds are cheaper to serve than FP8 but are quantized further, so answers can differ slightly from Z.ai's. Each model's host table shows it.
Z.ai's GLM Coding Plan is a monthly subscription for using GLM inside coding tools, priced separately from these per-token API rates. z.ai/subscribe
FAQ
Prices tell you what a model costs to call. Turning it into a product people use is the harder part, and it is the part you can skip wiring up yourself.
Prices in USD per 1M tokens from OpenRouter's public API, fetched October 8, 2026 at 23:13 UTC and refreshed every 6 hours.
© 2026 Ox Alpha. Independent of Z.ai.
Facts sourced from model cards, public reporting and community forensics. Every benchmark is cited to its source.