It appeared overnight, free, with the largest context window in its class, then split the coding benchmarks down the middle. The community spent a week fingerprinting it, and the case is now closed: it was Z.ai's GLM-5.3 Flash. Here is the honest data, the trail that led there, and where to run it now.
Run it now
The free chat on this site closed on September 30, 2026. To try the model now, open GLM-5.3 Flash in GenMagic, a multi-model studio made by the team behind Ox Alpha: verify your email there and you start with free credits.
At a glance
No hype. Provider-published where available, community-measured where not, every claim traceable to a source on the pages below. Last verified September 4, 2026. It launched free on OpenRouter on August 20, 2026; when the preview closed, the host that ran it retired the stealth slug and revealed the model as Z.ai's GLM-5.3 Flash.
The GLM family
GLM-5.3 Flash is one model in a fast-moving open-weight family. Each model below has its own page: a sourced spec sheet, Z.ai's published benchmarks, how to run it, and what it costs on every host. Prices are Z.ai's own, per 1M input / output tokens.
| Model | Released | Parameters | Context | License | Z.ai price (in / out) |
|---|---|---|---|---|---|
GLM-5.3 Flash Cheap, fast, 1M-context multimodal (the model behind Ox Alpha) | Aug 2026 | 320B total / 18B active | 1M | MIT | $0.15 / $0.50 |
GLM-5.3 Z.ai's current flagship: the strongest GLM for coding and agents, 1M context | Aug 2026 | 753B total (GLM-5.2 base) | 1M | GLM-5.3 License | $1.40 / $4.40 |
GLM-5.2 The newest GLM flagship under plain MIT: 1M context, long-horizon coding | Jun 2026 | 744B / 40B active | 1M | MIT | $1.40 / $4.40 |
The investigation
A jailbroken system prompt confirmed Ox Alpha was instructed to hide its maker, which set off an open-source effort to fingerprint it. The evidence pointed overwhelmingly at one lab, and when the preview closed the host retired the stealth slug and confirmed it: Z.ai's GLM-5.3 Flash. Here is the trail that led there, ranked by the weight of the evidence, with market odds shown where a source published one.
CONFIRMED, and it was the dominant theory all along. When the preview closed, OpenRouter retired the stealth slug and resolved it to z-ai/glm-5.3-flash, confirming this attribution at the host layer. The forensics that pointed here: tokenizer probes matched GLM-5.3 exactly (30/30 and separately 95/95, +75-token wrapper); video encoder matched GLM-5V-Turbo token-for-token; rejects audio like GLM-5V; shares GLM error codes 1214 and 1301; a leaked Java stack trace exposed 'com.wd.paas.api.domain.v4.chat.ChatCompletionRequest' matching Zhipu's /api/paas/v4 route (Chetaslua rated operator-layer ID at 0.98); Ben Davis said he was '99% certain'; Manifold priced Z.ai ~80%. Z.ai previously ran GLM-5 anonymously as 'Pony Alpha.' (Z.ai has not published a separate formal statement.)
Benchmarks
Ox Alpha's reputation was made by a hand-picked 10-task run and complicated by the full one. We show every score with its source, so you judge the model, not the headline.
~80%
DeepSWE, 10-task subset (the viral number)
~63%
DeepSWE, full 113-task run (corrected)
28%
LiveCodeBench v6, independent & reproducible
57
Artificial Analysis Intelligence Index (v4.1.1)
FAQ
You have seen the data. When you want a working product rather than an API call, Founden builds it: a company, a product, a 3D world, whatever you have in mind, without wiring up any of the infrastructure yourself.
1M
token context window
131K
max output tokens
$0
during the stealth preview
#2
on OpenCode at launch
Prices in USD per 1M tokens from OpenRouter's public API, fetched October 8, 2026 at 23:13 UTC and refreshed every 6 hours.
Weighed and ruled out. The main non-Chinese theory: Wccftech initially suggested GLM, then updated to say Ox Alpha 'could be an unreleased version of Microsoft's MAI.' It was never backed by the infrastructure fingerprints that supported the Z.ai case, and the host reveal (GLM-5.3 Flash) settled it against MAI.
Weighed and ruled out. Early community sentiment thought it was a Gemini model, but this was dismissed as 'community sentiment, not infrastructure proof' and rejected by the fingerprinting crowd (one skeptic: it 'can't score below 3.7 Flash to be next gen'). The reveal (GLM-5.3 Flash) confirmed it was not Gemini.
© 2026 Ox Alpha. Independent of Z.ai.
Facts sourced from model cards, public reporting and community forensics. Every benchmark is cited to its source.