What Is Ox Alpha? The Stealth Model, Revealed
August 28, 2026 · 22 min read · Ox Alpha
The plain-English explainer for the anonymous 1M-context model that developers wired into their agents, and the community traced all the way back to Z.ai's GLM-5.3 Flash.
On August 20, 2026, a reasoning model called Ox Alpha appeared, free of charge, with no maker attached to it. Within three days it was the No. 2 model on OpenCode by usage, developers had pushed trillions of tokens through it, and the entire AI community was running the same forensic experiment: taking the model apart to figure out who built it. Then, when the free preview closed in late August, the answer arrived. The host that ran the preview, OpenRouter, retired the stealth/ox-alpha slug and identified the model as Z.ai's (Zhipu AI's) GLM-5.3 Flash, exactly where the community's fingerprints had pointed. The most talked-about release of the week ended with the crowd proven right.
This is the encyclopedic version: what Ox Alpha actually is, the specs that are genuinely confirmed, how the identity was resolved, how ordinary people access it now, whether it is any good at the coding it was built for, and whether it is safe to point at your real work. For the deep investigation into who built it, we keep a separate, sourced piece: who built Ox Alpha, the evidence ranked. For the hype-free benchmark autopsy, see Ox Alpha, honestly. This guide is the map that sits above both.
Contents
- What Ox Alpha is, in one paragraph
- Where it came from: a stealth launch, not a leak
- The confirmed specs, in a table
- Confirmed vs reported vs speculative: how to read the facts
- The identity, resolved: it was Z.ai's GLM-5.3 Flash
- How people actually access Ox Alpha
- Is it any good? The honest capability picture
- Is it safe? The privacy and anonymity reality
- What Ox Alpha is, and what it is not
What Ox Alpha is, in one paragraph
Ox Alpha was an anonymous "stealth" reasoning model that appeared, free of charge, on August 20, 2026, and was immediately wired into real coding agents by hundreds of thousands of developers. Its public listing described it verbatim as "a reasoning model designed for coding, sustained agentic work, and production workloads," and its provider-published specs are genuinely strong on two axes: a 1,048,576-token context window (a full 1M) and multimodal input that accepts text, images, and video - OpenRouter. It outputs text plus tool calls, supports streaming and JSON, and exposes an adjustable reasoning mode. In shape, it is an ordinary modern coding model. In circumstance, it was not ordinary at all.
The defining feature of Ox Alpha was not a number on its spec sheet but a deliberate absence. During the preview, its listing stated only that the model was "developed and operated by a third-party provider who has chosen to remain anonymous during this preview," and a jailbroken system prompt confirmed the model was explicitly instructed to identify itself strictly as "ox-alpha" from "an undisclosed organization" - Pliny the Liberator. That absence has since been filled. When the preview closed, OpenRouter retired the cover name and resolved it to Z.ai's GLM-5.3 Flash, so the honest one-line answer to "what is Ox Alpha" is now this: it was Z.ai's (Zhipu AI's) GLM-5.3 Flash, a capable, cheap, 1M-context coding model, run under a cover name for a week of clean public testing, that the open-source community fingerprinted correctly before the host confirmed it. Everything else in this guide is detail hung on that frame. If you want the live, continuously verified numbers rather than the narrative, the full spec sheet carries every figure with its source.
Where it came from: a stealth launch, not a leak
It is tempting to read a nameless model as a mistake, a leak, or an abandoned experiment. It was none of those. Ox Alpha was a stealth launch, a recognised and increasingly common strategy where a lab ships a model under a cover name to collect clean, unbiased public testing before attaching its brand. A stealth release buys three things a named release cannot: reception that is not distorted by brand-driven hype or brand-driven hate, a week of real agentic workloads running against production traffic, and the option to claim credit later only if the results are good. Read that way, the anonymity was not a bug in the story, it was the entire point of it, and Z.ai had run the exact play before, previewing GLM-5 anonymously as "Pony Alpha" before shipping GLM-5.3 on August 14, six days before Ox Alpha appeared.
The timeline supports the deliberate reading. Ox Alpha appeared on OpenRouter as stealth/ox-alpha on August 20, 2026, priced at $0 in and $0 out, and OpenCode launched it the same day, promising it "free for the next week" and advertising serving capacity on the order of 100 trillion tokens per day - Cryptobriefing. Within 24 hours the fingerprinting had begun. By August 23 the coverage peaked, with TechCrunch and Bloomberg both asking who was behind it, and Stripe CEO Patrick Collison testing it and calling it "very impressive" - The Cherry Creek News. The provider projected a roughly one-week free window, and on August 28, 2026 the preview closed exactly as planned: OpenRouter retired the stealth/ox-alpha slug, which now resolves to z-ai/glm-5.3-flash, with a reveal message thanking testers for "participating in the Stealth Ox Alpha testing period" and confirming "This model was ZAI's GLM-5.3 Flash" - OpenRouter.
That resolution is the practical takeaway of this whole section. The week-long experiment did exactly what a stealth launch is designed to do: it gathered clean, brand-blind testing, and then the maker's identity graduated into the open. The cover name is retired, the model lives on under its real slug, and the honest read of the whole episode is that the community's forensics called it before the host confirmed it. If you wired Ox Alpha into an agent during the preview, the model you were using was GLM-5.3 Flash all along, and the guidance from here is simply to point your integration at its real slug (z-ai/glm-5.3-flash) or pick another of the current ways to run it.
The confirmed specs, in a table
Before any of the disputed benchmark drama, it is worth pinning down the part that is not in dispute: the provider-published specifications. These are the figures that appear on the model's own listings across multiple independent platforms, which is the highest confidence tier this site recognises. They tell you what the model can physically accept and emit, independent of how well it performs, and they are the same wherever Ox Alpha is served. Getting these straight first means the rest of the guide can argue about quality without also arguing about facts.
Two numbers dominate. The 1M context window is large enough to hold an entire codebase, a long specification, or a full transcript in working memory at once, which is precisely the long-horizon engineering work the model is positioned for. The 131,072-token maximum output is the ceiling on a single completion, and crucially it includes reasoning tokens, so a heavy chain of thought eats into the same budget as the final answer. Everything else in the table below is the standard modern agentic toolkit, present and confirmed, but unremarkable next to those two headline capacities.
| Spec | Value | Confidence |
|---|---|---|
| Model slug | stealth/ox-alpha (retired), now z-ai/glm-5.3-flash | Confirmed |
| Context window | 1,048,576 tokens (exactly 1M / 2^20) | Confirmed |
| Max output | 131,072 tokens (includes reasoning tokens) | Confirmed |
| Input modalities | Text, image (JPG/PNG/GIF/WEBP), video (base64); audio rejected | Confirmed |
| Output modality | Text only, plus tool/function calls | Confirmed |
| Tool / function calling | Supported (tool_choice none/auto/required/named) | Confirmed |
| Streaming | SSE supported (stream: true) | Confirmed |
| Structured output | JSON object mode, no schema enforcement | Confirmed |
| Reasoning control | reasoning_effort supported | Reported |
| Price | Preview was $0 in / $0 out; revealed model (GLM-5.3 Flash) now $0.15 in / $0.50 out per 1M at Z.ai's list price (launch price $0.075 / $0.25) | Confirmed |
Every row above marked "Confirmed" traces to a provider listing: the context window, output ceiling, output modality, and JSON behavior come from the OpenRouter model page, the input modalities, tool-calling, and streaming details come from the AI/ML API documentation, and the retired slug plus the revealed model's price come from the OpenRouter z-ai/glm-5.3-flash listing. The one row marked "Reported" is the reasoning control, and it earns the lower grade for a specific reason: AI/ML API documents effort levels of low, medium, and high, while an independent fingerprint report lists max, high, and low with no "medium" level at all - DigitalApplied. That small disagreement is a useful illustration of the honesty ladder this whole site runs on, which the next section makes explicit.
One note on that price row: $0.15 / $0.50 per 1M is what GLM-5.3 Flash costs on Z.ai's own endpoint now that both the free preview and the half-price launch have ended; some third-party hosts charge less ( live prices for every host). It is a dossier fact about the real model. This site's own free chat, which it subsidized, was retired on September 30, 2026.
Confirmed vs reported vs speculative: how to read the facts
The most important skill for understanding Ox Alpha is not technical, it is epistemic: knowing which claims are solid and which are inference. Throughout the preview the operator was anonymous, with no model card, no official benchmark, and no support channel to appeal to, so every fact about Ox Alpha sat on one of three rungs. Confirmed means provider-published, appearing on the model's own listings. Reported means community-observed, seen consistently by testers but not published by the maker. Speculative means inferred or disputed, a reasonable guess with real evidence behind it but no confirmation and sometimes active contradiction. Collapsing those three rungs into one is exactly how the hype got out of hand. That ladder is also why the identity call held up: the fingerprints below were flagged as reported inference, not fact, right up until the host reveal moved the identity itself onto the confirmed rung.
The distinction has teeth because the rungs disagree with each other on some of the most quoted facts. The context window and free pricing are confirmed and stable. The throughput (roughly 25 to 50 tokens per second) and latency (a median multi-step agent turn around 11.6 seconds) are reported, meaning real but variable by load and snapshot - explainx.ai. And the juiciest details, the 744B-parameter Mixture-of-Experts architecture and the November 2025 knowledge cutoff, are purely speculative: a third-party analysis and a fingerprint, respectively, with no official disclosure behind either - Local AI Zone. Treating that architecture figure as fact, which many write-ups do, is the single most common error in Ox Alpha coverage.
The practical way to apply this: when you read any Ox Alpha claim, including on this site, ask which rung it sits on before you act on it. A confirmed spec you can plan around. A reported behavior you should verify in your own harness before trusting. A speculative figure you should quote only with the caveat attached, or not at all. This is not pedantry, it is the difference between an honest resource and a rumor mill, and it is the reason our benchmarks page tags every number with its confidence level rather than presenting them as a flat leaderboard.
The identity, resolved: it was Z.ai's GLM-5.3 Flash
The question everyone arrived with was "who made it," and the case is now closed: Ox Alpha was Z.ai's (Zhipu AI's) GLM-5.3 Flash. During the preview the host positioned itself only as a router, disclaimed being the developer, and stated the provider had chosen anonymity - TechCrunch. That silence launched a weekend-long open-source investigation whose evidence pointed overwhelmingly at one lab, China's Zhipu AI (Z.ai) and its GLM-5.x family, and when the preview closed the host reveal confirmed exactly that call. We investigate the full evidence trail, with every fingerprint ranked, in who built Ox Alpha. Here is only the compressed version you need to understand what Ox Alpha is.
Four independent classes of fingerprint converged on Zhipu, and all four turned out to be right. Tokenizer probes matched GLM-5.3 exactly, down to a constant hidden 75-token wrapper on every request. Video-token accounting matched GLM-5V-Turbo token-for-token. The model rejects audio exactly as GLM-5V does and shares GLM's 1214 and 1301 error codes. And a deliberately malformed request leaked a Java class name mapping onto Zhipu's own documented API route - OrcaRouter. Confidence from the people who looked hardest ran high: Ben Davis said he was 99% certain, one researcher rated operator-layer identity at 0.98, and a Manifold prediction market priced Z.ai at roughly 80% - Manifold Markets. Z.ai also had a precedent, having previewed GLM-5 anonymously as "Pony Alpha" before shipping GLM-5.3 on August 14, six days before Ox Alpha appeared.
The reveal came from the host, not from a Zhipu press release, and that distinction is worth keeping honest. When the free preview closed on August 28, 2026, OpenRouter (the platform that ran the stealth preview) retired the stealth/ox-alpha slug and identified the model as Z.ai's GLM-5.3 Flash, pointing users to z-ai/glm-5.3-flash - OpenRouter. Z.ai itself has not issued a separate formal statement of its own, so the identification rests on the preview host's slug graduation. Throughout the hunt the load-bearing caveat had been that shared infrastructure proves shared serving, not shared identity: a tokenizer match, an error code, and a leaked API path all live in the operator's serving stack, not in the model weights, and as the builder of one fingerprinting tool put it, a lab "can serve two different models on the same stack" - SiliconANGLE. That discipline turned out not to matter here, the trail led to the right answer, but it is exactly why the conclusion reads as credible rather than lucky. The competing theories (an unreleased Microsoft MAI model, an early Google Gemini guess, Xiaomi's MiMo ruled out because it accepts audio, an xAI/Cursor fine-tune) were all weighed and ruled out. So the verdict is no longer "fingerprint-strong but officially unconfirmed"; it is confirmed at the host layer, and the community called it.
How people actually access Ox Alpha
Understanding what Ox Alpha is only becomes useful when you can touch it, and the good news is that access is unusually easy because the model is cheap and OpenAI-compatible. There are three broad routes, and they suit three different kinds of user. The simplest is to call the API from your own code, through Z.ai's own endpoint or OpenRouter, which is also how the model shows its real strength inside an agent loop. The most independent is to self-host the open weights, which Z.ai published on Hugging Face (zai-org) under the MIT license, with a runtime like vLLM or SGLang. And the third is the grab-bag of third-party platforms that listed it under their own aliases during the preview, useful mainly if you already live inside one of them. (Until September 30, 2026 this site also ran a free browser chat on it; that chat is now retired.)
The API route is the one worth understanding in detail, because Ox Alpha's whole design points at it. Now that the preview has closed, the stealth ox-alpha slug has retired and the model lives on under its real name: z-ai/glm-5.3-flash on OpenRouter, or Z.AI's own API. Every serving endpoint is OpenAI-compatible, which means most existing SDKs work simply by swapping the base URL and pointing at the provider you choose, and other hosts (AI/ML API, Felo, AIHubMix) expose their own compatible base URLs, some also offering an Anthropic-compatible shape - Felo. Our API page is now a migration guide for code that used this site's retired endpoint; wherever you land, a backup model is always worth wiring in for any single-provider dependency.
The routes below are the main ones, each with its own quirk worth knowing before you commit to it.
- The open weights, MIT-licensed on Hugging Face (zai-org), which you can self-host with a runtime like vLLM or SGLang.
- OpenRouter at
z-ai/glm-5.3-flash(the retiredstealth/ox-alphaslug now redirects here), the canonical listing, where Z.ai's own endpoint charges $0.15 in and $0.50 out per 1M tokens and other hosts set their own prices - OpenRouter. - Z.AI's own API, the model's first-party home now that the maker is known, OpenAI-compatible.
- AI/ML API and Felo, OpenAI-compatible hosts for dropping it into an existing codebase - AI/ML API.
- AIHubMix and independent UIs, which listed the preview at $0 and now bill at standard GLM-5.3 Flash rates, so read the fine print - AIHubMix.
One habit carried over from the stealth period is still useful: the model wore different ids on different platforms during the preview (OpenCode Zen called it opencode/x-preview-f-free, OpenCode Go called it ox-alpha-free), so listing available models and selecting by name rather than hardcoding an alias remains the safe pattern - Kingy AI. For a full walkthrough of setup, harness choice, and prompting, our companion guide covers how to use Ox Alpha end to end. The theme is simply that access is now stable under a known name, and where to run GLM-5.3 Flash now lists every current option.
Is it any good? The honest capability picture
Now the question the specs cannot answer: is Ox Alpha actually good? Here the honest answer is genuinely mixed and genuinely disputed, and the reason is instructive. Ox Alpha's entire benchmark reputation was built on one developer's unaudited coding evaluation, then complicated by the same developer's more careful follow-up, then contradicted outright by an independent reproducible test. The model became famous for a number that its own tester later walked back, which is why "how good is it" has no clean one-line answer and why anyone quoting a single figure is telling you less than half the story. The chart below is the whole controversy in three bars.
Read left to right, the three bars are a lesson in sample size. The viral 80% came from a hand-picked 10-task slice of DeepSWE, where Ox Alpha beat Claude Fable 5, GLM-5.3, Grok 4.6, and GPT-5.6 Sol, and where the tester himself warned that "10 tasks can have high variance" and to treat it as directional - oxalpha.com. The corrected 63% came when the same tester ran all 113 DeepSWE tasks, landing mid-tier, near GPT-5.6 Sol and Claude Opus 4.8, a perfectly respectable result but not a record - The Cherry Creek News. The independent 28% came from a fully reproducible LiveCodeBench v6 run (greedy, temperature 0, single attempt, zero generation failures), well below the frontier and a direct contradiction of the "frontier coder" framing - xnasarx on GitHub.
There is one more fact that reframes all three numbers: Ox Alpha has no entry on Artificial Analysis or LMArena, and its own listing carries no capability benchmark, so there is no authoritative aggregator score to anchor against - Local AI Zone. Its single most defensible ranking is not a capability score at all, it is adoption: No. 2 on OpenCode within three days, which is a story about being free, not about being best. Underneath the noise sits the one pattern everyone agrees on, and it is the practical guidance to take away: users who run it inside a real agent harness with tools rate it far higher than chat-only users, who cite loops on long tasks, weak frontend work, and verbosity. For the full model-by-model placement against the field, see the comparison and our companion piece on Ox Alpha versus the field.
Is it safe? The privacy and anonymity reality
Capability is only half of whether you should use a model. The other half is safety, and on that axis Ox Alpha carries a caveat the reveal only partly retires. Throughout the preview the operator was anonymous and retained every prompt and completion, which meant you could not identify, audit, or contract with the party holding your inputs - TechTimes. We now know that party was Z.ai (Zhipu AI), which closes the "you cannot name the operator" gap but does not erase the deeper point: anything sent during the stealth window went to a provider you could not contract with at the time, and GLM-5.3 Flash today still comes with the ordinary due-diligence a named provider deserves rather than a data-processing agreement you have already signed.
The preview data terms compounded the problem by conflicting with each other. One host said prompts were retained but not used for training, another host marketed "zero data retention," and the broader Stealth Program agreement reportedly permitted collection for training - OpenRouter. When terms conflict, the only safe reading is the strictest one, so the working assumption for anything sent during the preview should be that it may have been logged and possibly trained on. There are geopolitical wrinkles too, now that the operator is confirmed as a Chinese lab: language-dependent censorship has been observed (refusing some questions in Chinese but answering them in English), and supply-chain considerations follow for any organization with a China-sourcing policy - Undercode Testing.
The practical rule that falls out of this is simple and non-negotiable, and it is the same one the community reached independently during the preview. As one Hacker News commenter put it, using an anonymous provider for free was like "getting free steak smuggled out of a grocery store inside somebody's pants," and others warned plainly to "never paste secrets into an anonymous host" - Hacker News. The generalized version, still correct now that the maker is known: send GLM-5.3 Flash only sanitized, non-sensitive data unless you have your own data agreement with Z.AI, never personal data, credentials, private source code, customer records, or anything regulated, and route anything you would not want logged through a provider whose terms you have actually reviewed. That is the correct default for any model, and the FAQ restates it for the specific questions people keep asking.
What Ox Alpha is, and what it is not
Pulling the whole picture together, Ox Alpha resists a single label because it genuinely is several things at once, and the honest verdict has to hold all of them. It is a real, capable reasoning model with a legitimate 1M-token context and strong multimodal input, one that people running it inside an agent harness report doing genuinely useful work: whole-repository reasoning, sustained tool use (one documented run did 69 tool calls with a single error and no retry loop), and catching real bugs other tools missed - OpenRouter. It is, now confirmed, Z.ai's GLM-5.3 Flash, a cheap, fast, high-throughput model, and it is open weight under the MIT license, so you can host it yourself. And it is a rare case of the community's forensics being fully vindicated by the reveal.
Equally, the guide would be dishonest if it stopped there, because what Ox Alpha is not is just as load-bearing. It is not a confirmed frontier coder by the reproducible numbers, which put it mid-tier at best and well below the leaders on independent tests. It is not ranked on any authoritative aggregator, so every circulating capability score is community-run and unverified. It is not the free-forever upstream preview it briefly was: that window closed on August 28, and the real model is now a cheap paid model everywhere; even the free chat this site subsidized was retired on September 30, 2026. And it is not a model for sensitive data without your own agreement in place, because the maker is a Chinese lab and the preview terms were ambiguous about training. Those four "is nots" are not hedges bolted onto the praise, they are the other half of an accurate description.
So the decision framework writes itself, and it is the note to leave on. If you want to explore a capable, cheap, high-context coding model inside an agent loop with sanitized data, GLM-5.3 Flash (the model that was Ox Alpha) is an easy yes, and where to run GLM-5.3 Flash now lists the quickest ways to judge it for yourself. If you want to bet a product on frontier-level coding accuracy, look higher up the leaderboard, and no benchmark, viral or corrected, changes that. Treat it as exactly what it is: a capable mid-tier model whose stealth debut was one of the more entertaining whodunits in AI, now solved. For the receipts behind every claim here, the spec sheet, the charted benchmarks, the comparison, and the two deeper dives on who built it and what the data says carry every number with its source.
The natural next step, once you understand what Ox Alpha is, is to build something real and see for yourself: start at founden.ai/app/build, a builder run by the same founder as this site.
This guide reflects the resolved picture as of August 29, 2026: the stealth preview closed on August 28 and OpenRouter identified Ox Alpha as Z.ai's GLM-5.3 Flash. This site's own free chat was retired on September 30, 2026.