The GLM family
GLM (General Language Model) is the line of AI models built by Z.ai, formerly Zhipu AI. It runs from the small open ChatGLM models to the frontier-scale GLM-5 generation: GLM-5, GLM-5.1, GLM-5.2 and the current GLM-5.3. Here is the whole family, honestly, and how to run each one.
GLM stands for General Language Model, an architecture first published out of Tsinghua University in 2021. The company behind it, Z.ai (formerly Zhipu AI), has made GLM one of the most openly available frontier model families: the recent releases are published under a plain MIT license, so you can download, self-host, fine-tune, and commercially deploy them.
The family is positioned around agentic, reasoning, and coding work. Below is the modern lineup compared side by side, the full lineage from 2021 to today, and every way to run a GLM model. This is an independent reference, not affiliated with Z.ai. Last verified October 4, 2026.
2021
First GLM paper
MIT
License on recent models
1M
Context (GLM-5.2, GLM-5.3)
744B
GLM-5 generation parameters
The lineup
The recent, fully-specced members of the family, newest first. Every row is provider-published and every model has open weights; all but GLM-5.3 are MIT. The older models are in the timeline below.
| Model | Released | Parameters | Context | Modalities | License |
|---|---|---|---|---|---|
| GLM-5.3 Flash Cheap, fast, 1M-context multimodal (the model behind Ox Alpha) | Aug 2026 | 320B total / 18B active | 1M | Text + image + video | MIT |
| GLM-5.3 Z.ai's current flagship: the strongest GLM for coding and agents, 1M context | Aug 2026 | 753B total (GLM-5.2 base) | 1M | Text | GLM-5.3 License |
| GLM-5.2 The newest GLM flagship under plain MIT: 1M context, long-horizon coding | Jun 2026 | 744B / 40B active | 1M | Text | MIT |
| GLM-5.1 GLM-5 retrained for long-running agentic engineering, 200K context | Apr 2026 | 744B / 40B active | 200K | Text | MIT |
| GLM-5 The 744B model that opened the GLM-5 generation, 200K context | Feb 2026 | 744B / 40B active | 200K | Text | MIT |
| GLM-4.7-Flash A 30B-A3B model for one machine, free on Z.ai's API | Jan 2026 | 30B / ~3B active | 200K | Text | MIT |
| GLM-4.7 The last GLM-4.5-architecture flagship: a 355B coding model, 200K context | Dec 2025 | 355B / 32B active | 200K | Text | MIT |
GLM-4.6V Vision model with native multimodal function calling | Dec 2025 | 106B total / 12B active | 128K | Text + vision (images, video, documents) | MIT |
GLM-4.6 Real-world coding and long-context agents (the prior open flagship) | Sep 2025 | ~355B total / 32B active | 200K | Text | MIT |
GLM-4.5V Multimodal document, screen, and video understanding | Aug 2025 | 106B total / 12B active | Not disclosed | Text + vision (images, video, PDFs, GUI) | MIT |
GLM-4.5 Open-weight agentic, reasoning, and coding flagship | Jul 2025 | 355B total / 32B active | 128K | Text | MIT |
GLM-4.5-Air Compact, cheaper-to-run sibling of GLM-4.5 | Jul 2025 | 106B total / 12B active | 128K | Text | MIT |
The lineage
How the family grew from a 2021 research architecture into an open, frontier-scale model line. Each step links to its source.
March 18, 2021
The GLM architecture is introduced in the paper 'GLM: General Language Model Pretraining with Autoregressive Blank Infilling', out of Tsinghua University's KEG lab. GLM stands for General Language Model.
August 2022
GLM-130B, a 130B open bilingual (Chinese and English) base model, is released and later presented at ICLR 2023. One of the first openly available models at the 100B scale.
March 2023
ChatGLM-6B, a small open bilingual chat model, is released and becomes a breakout in the open-source community. ChatGLM2-6B and ChatGLM3-6B follow later in 2023.
June 18, 2024
The GLM-4 generation arrives: a proprietary API flagship (GLM-4, later GLM-4-Plus) plus the open GLM-4-9B (with an 8K base and a 1M-context variant) and the vision model GLM-4V-9B.
July 2025
GLM-4.5 and GLM-4.5-Air ship as open-weight Mixture-of-Experts models under the MIT license (355B/32B and 106B/12B), branded as 'Agentic, Reasoning, and Coding' foundation models. Independent coverage rates GLM-4.5 among the strongest open-weight models at launch.
August 2025
GLM-4.5V, a multimodal vision-language model built on GLM-4.5-Air, is released under MIT, handling images, video, PDFs, and on-screen content.
September 30, 2025
GLM-4.6 lands under MIT with a 200K context window (up from 128K), stronger real-world coding and agentic tool use, and roughly 15% fewer tokens per task than GLM-4.5.
December 2025
GLM-4.6V (106B) adds native multimodal function calling to the vision line, and GLM-4.7 (December 22, MIT) brings large agentic-coding gains over GLM-4.6 plus Preserved and Turn-level Thinking. Its 30B sibling GLM-4.7-Flash follows on January 19, 2026.
February 11, 2026
GLM-5 doubles the family's scale to 744B parameters (40B active), trained on 28.5T tokens with DeepSeek Sparse Attention, under MIT. Z.ai had previewed it anonymously on OpenRouter as 'Pony Alpha'.
March to April 2026
GLM-5 Turbo, a faster API-only model for agent workflows, then GLM-5V-Turbo, Z.ai's first native multimodal agent model. In April, GLM-5.1 retrains GLM-5 for long-horizon agentic engineering and tops Z.ai's SWE-Bench Pro comparison.
June 16, 2026
GLM-5.2 extends the context window to a full 1M tokens and lifts coding sharply (Terminal Bench 2.1 rises from 63.5 to 81.0), still under MIT.
August 14, 2026
GLM-5.3, the GLM-5.2 base with much more post-training, becomes Z.ai's flagship. Its open weights follow on August 28 under a new GLM-5.3 License.
August 2026
GLM-5.3 Flash (320B/18B, 1M context, MIT) debuts anonymously as the stealth 'Ox Alpha' model, then is revealed as Z.ai's and open-weighted, the model this site is built around.
September 2026
Z.ai adds faster-serving variants: GLM-5.3 FlashX (up to 200 tokens per second) and GLM-5.3 Prime (1.5 to 2 times GLM-5.3's output throughput).
Access
Use Z.ai's own API or its GLM Coding Plan, run it through a third-party host, or self-host the open weights. The recent models are MIT, so nothing locks you in.
FAQ
You have seen the whole family. The next step is to build something real with it, without wiring up any of the infrastructure yourself. For a hands-on start, O-mega.ai, a site from the same team behind Ox Alpha, has a practical guide to building cheap vision agents with GLM-5.3 Flash, the exact model behind this site.