Ox Alpha, honestly: what the data actually said, and who it turned out to be
August 24, 2026 · 8 min read · Ox Alpha
On August 20, 2026, a model called Ox Alpha appeared with no announcement, no maker attached, and a price tag of zero. Within three days it was the No. 2 model on OpenCode by usage, developers had pushed trillions of tokens through it, and the internet had split into two camps: the people calling it a frontier-class coder, and the people who ran the full test suite and found something more ordinary.
Then, when the preview closed in late August 2026, the mystery closed with it. The community called it correctly. Ox Alpha was Z.ai's (Zhipu AI's) GLM-5.3 Flash. This is the honest version, with every number traceable to a source. No hype, because the interesting thing about Ox Alpha is that it never needed any: the evidence pointed to one answer, and the answer held up.
The specs that actually matter
Ox Alpha is a reasoning model built for coding and sustained agentic work. The provider-published specs are unusually strong on two axes and unremarkable on the rest:
| Spec | Value |
|---|---|
| Context window | 1,048,576 tokens (a full 1M) |
| Max output | 131,072 tokens (includes reasoning tokens) |
| Input modalities | Text, image, video |
| Output | Text (plus tool/function calls) |
| Reasoning | Yes, with adjustable effort |
| Price (during preview) | $0 in / $0 out |
| Interface | OpenAI-compatible (stealth slug ox-alpha, now retired to z-ai/glm-5.3-flash) |
The 1M context is the headline. It is large enough to hold an entire codebase, a long paper, or a full transcript in working memory at once, which is exactly the kind of task the model is positioned for. The free preview price was the other headline, and the two together explain the adoption spike far better than any benchmark does. That exact 1,048,576-token figure is also one of the tells: it matches GLM-5.3 Flash's published context window to the token.
One honest footnote on provenance, now resolved: for a week nobody had officially claimed Ox Alpha, and independent fingerprinting (tokenizer counts, video-token accounting, shared error codes) pointed strongly at Zhipu's GLM family. The analysts were careful to say shared serving infrastructure is not proof of identity. It turned out not to matter here. When the preview closed, OpenRouter (the host that ran the stealth preview) retired the stealth/ox-alpha slug and identified the model as Z.ai's GLM-5.3 Flash, exactly where the forensics pointed. Z.ai has not issued a separate formal statement of its own, so the confirmation rests on the preview host's slug graduation, not a Zhipu press release.
The benchmark reality
Here is where the two camps come from. Ox Alpha's reputation was made by a single number and complicated by the next one.
- The viral result was roughly 80% on a hand-picked 10-task slice of DeepSWE, a coding-agent benchmark. That is the number that launched the headlines.
- The same tester's full run on all 113 DeepSWE tasks corrected it down to about 63%: mid-tier, near GPT-5.6 Sol and Claude Opus 4.8, and honestly a fine result, just not a record.
- An independent, reproducible test on LiveCodeBench v6 (greedy, temperature 0, single attempt) put it at 28% pass@1: well below the frontier coders, and a direct contradiction of the "frontier" framing.
There is one more thing worth knowing: through late August 2026, Ox Alpha had no entry on Artificial Analysis or LMArena, and its stealth listing carried no capability benchmark. Every score in circulation was community-run and unverified. Its single most defensible ranking was not a capability score at all, it was adoption: No. 2 on OpenCode within three days, which is a story about being free and available, not about being the best.
The pattern underneath the noise: Ox Alpha rates far higher inside a real agent harness (with tools and a terminal) than it does in plain chat. If you judge it by direct code generation, it looks middling. If you wire it into an agent loop, people who did report genuinely strong results.
How it compares
Because the community benchmarks are a mess of different suites run by different people, the fairest way to place Ox Alpha is on the axes that are objective (context, price, and licensing) alongside each model's own best-known coding number for rough orientation. (Coding scores below come from different benchmarks and are not strictly comparable; treat them as directional.)
| Model | Context | Price /1M (in / out) | Best-known coding | Open weights |
|---|---|---|---|---|
| Ox Alpha (GLM-5.3 Flash) | 1M | $0.15 / $0.50 | ~63% DeepSWE (community) | No |
| GLM-5.3 | 1M | $1.40 / $4.40 | 66.9% DeepSWE | Not yet |
| Gemini 3.7 Flash | 1M | $0.75 / $3.75 | 65.3% DeepSWE | No |
| GPT-5.6 Sol | 1.05M | $4 / $20 | 82.2% SWE-bench Verified | No |
| Claude Opus 4.8 | 1M | $5 / $25 | 88.6% SWE-bench Verified | No |
| Claude Fable 5 | 1M | $10 / $50 | 95% SWE-bench Verified | No |
| DeepSeek V4 Pro | 1M | ~$0.66 / $1.98 | 80.6% SWE-bench Verified | Yes (MIT) |
Read it and the picture is clear, and the reveal makes it clearer. On price and context, GLM-5.3 Flash is still one of the cheapest 1M-context models here ($0.15 in / $0.50 out per 1M tokens at Z.ai's list price, far under the GLM-5.3 full model's $1.40 / $4.40). On measured coding, it sits with the mid-pack (near its own GLM-5.3 sibling and Gemini 3.7 Flash) and well behind the SWE-bench Verified leaders. That is a coherent story now: a cheap, fast, high-throughput Flash-tier model, not a stealth frontier coder. For the full, sourced, model-by-model matrix, see the comparison.
What it is good at, and what it is not
Strengths, from people running it in production:
- Whole-repository reasoning across a 1M-token working set, without constant context compaction.
- Sustained agentic operation: one documented run did 69 tool calls with a single error and no retry loop.
- Genuine bug-finding: reports of it catching real issues other tools missed.
- Multimodal input, so you can combine code or text with screenshots and video.
Weaknesses, from the same people:
- It gets stuck in loops or stalls on long tasks until you nudge it. This is the single most common complaint.
- Weak on frontend and visual reasoning, and it tends toward verbose output.
- Middling interactive latency and occasional errors under viral load.
Should you build on it?
The reveal simplifies this. Ox Alpha was a time-boxed stealth preview of GLM-5.3 Flash, and now that the preview has closed, the honest guidance is to build on the named, documented model, not the ghost. Three things follow, and they matter more than any single score:
- Build against GLM-5.3 Flash directly, not the retired slug. The
stealth/ox-alphaslug is gone. The model now lives atz-ai/glm-5.3-flashon OpenRouter (and on Z.AI's own API), with a real model card, real pricing ($0.15 in / $0.50 out per 1M tokens at Z.ai's list price, less on some hosts), and an attributable operator you can actually contract with. Wire that in, with a fallback model for anything load-bearing. - Mind the data trail from the preview period. During the anonymous window the operator, now known to be Z.ai, retained prompts and completions, and the preview terms were ambiguous about training. That is history now, but the general rule stands: do not send secrets, credentials, private source code, personal data, or anything regulated to any endpoint without a data agreement you can rely on. Use sanitized test data, and read GLM-5.3 Flash's current terms before you send anything real.
- The "free" was always a subsidy, and it has ended. The provider preview across OpenRouter, OpenCode, and others was a launch promotion, not a commitment, and it closed. This site's own subsidized free chat and API were retired too, on September 30, 2026. GLM-5.3 Flash is a paid model now, though a cheap one. Price your real workloads against it accordingly.
The bottom line
Ox Alpha was the most interesting free thing in AI for a week, and it earned the attention: a cheap, fast, 1M-context model that developers wired straight into their agents before anyone knew who built it. The community ran the forensics, called the maker, and was proven right when the preview closed and the slug resolved to Z.ai's GLM-5.3 Flash. It is not a frontier coder by the reproducible numbers (that framing never survived the full benchmark runs), and now it does not need to hide behind a codename either. It is exactly what the evidence said: a solid, mid-tier, high-throughput Flash model.
To judge it for yourself, see where to run GLM-5.3 Flash now (this site's own free chat was retired on September 30, 2026). When you want the receipts, the full spec sheet, the charted benchmarks, and the FAQ have every claim above with its source.