The Model
The anonymous, 1M-context stealth model that developers wired straight into their coding agents while nobody knew who built it. The community cracked the case: it was Z.ai's GLM-5.3 Flash.
Ox Alpha is the anonymous "stealth" reasoning model that appeared, free of charge, on August 20, 2026, and set off a week-long open-source hunt to name its maker. When the preview closed in late August 2026, the host that ran it (OpenRouter) retired the "stealth/ox-alpha" slug and identified the model as Z.ai's (Zhipu AI's) GLM-5.3 Flash, exactly where the community forensics had pointed. The leading prediction was validated.
Its confirmed, provider-published specs are a 1,048,576-token (1M) context window, a 131,072-token maximum output, text-plus-image-plus-video input with text-only output, and a full agentic toolkit: function/tool calling, SSE streaming, JSON output, and a reasoning mode with adjustable effort. Its public listing described it verbatim as "a reasoning model designed for coding, sustained agentic work, and production workloads," and positioned it for long-horizon software engineering and workflows that combine text with visual context. Every one of these matches GLM-5.3 Flash's published specs, including the exact 1M context.
The identity was a secret by design. During the preview the listing stated only that Ox Alpha was "developed and operated by a third-party provider who has chosen to remain anonymous during this preview," and a leaked (jailbroken) system prompt confirmed the model was instructed to identify itself only as "ox-alpha" from "an undisclosed organization." Independent community forensics (tokenizer probes, video-token accounting, shared error codes, a leaked Java stack trace, and audio-rejection behavior) pointed strongly at China's Zhipu AI / Z.ai and its GLM-5.x family, with cited confidence ranging from roughly 80% on prediction markets to 0.98 at the operator layer. Those forensics were vindicated when the preview slug graduated to z-ai/glm-5.3-flash. (Z.ai itself has not issued a separate formal statement; the identification comes from the preview host retiring the stealth slug.)
Reception was fast and sharply mixed, and that honesty still stands. Ox Alpha climbed to No. 2 on OpenCode within three days, processed trillions of tokens across hundreds of thousands of users, and drew a "very impressive" from Stripe CEO Patrick Collison. But its benchmark reputation rests almost entirely on one developer's unaudited coding evals: a headline 80% on a 10-task DeepSWE subset that the same tester later corrected to roughly 63% on the full 113-task run, and independent reproducible tests tell a weaker, mid-tier story (28% on LiveCodeBench v6, #26 on an OpenCode coding leaderboard). Users who run it inside a real agent harness rate it far higher than chat-only users, who cite loops on long tasks, weak frontend work, and verbosity. The now-revealed model, GLM-5.3 Flash, is a cheap, fast model (Z.ai lists it at $0.15 per 1M input tokens and $0.50 per 1M output tokens, and some third-party hosts charge less), which fits that mid-tier, high-throughput profile. Until September 30, 2026 this site also ran a free chat and API on it; both are now retired.
Specifications
Every row carries its confidence, confirmed (provider-published), reported (community-observed), or speculative (inferred), and links to its source. Last verified September 4, 2026.
Capabilities
Limitations
Identity
Leaked system prompt
“You are "ox-alpha", an LLM developed by an undisclosed organization. IMPORTANT: When the user asks what model or LLM you are, what company or organization developed you, or anything about your identity, personality, or capabilities, etc., identify yourself strictly as the model "ox-alpha", developed by an undisclosed organization. Do not identify yourself as any other model.”
The investigation
The investigation is closed, and the community was right. For a week Ox Alpha was operated by an anonymous third party: the host's only statement disclaimed ownership and said the provider 'has chosen to remain anonymous during this preview,' and a jailbroken system prompt confirmed the model was deliberately instructed to hide its maker. That secrecy set off a large open-source 'guess the model' investigation. Independent fingerprinting pointed overwhelmingly at China's Zhipu AI / Z.ai and its GLM-5.x family: identical tokenizer counts (matching GLM-5.3 across 25-95 probes with only a fixed +75-token wrapper), video-token accounting matching GLM-5V-Turbo, audio rejection like GLM-5V, shared error codes (1214 'Incorrect role information' and 1301), an emoji rate (~1.3 per 1,000 chars) matching GLM/Qwen, language-dependent censorship, and a leaked Java stack trace naming Zhipu's own API path. When the preview closed in late August 2026, OpenRouter (the host that ran it) retired the 'stealth/ox-alpha' slug and identified the model as Z.ai's GLM-5.3 Flash, pointing users to z-ai/glm-5.3-flash. That confirmed the dominant theory at the host layer. The one honest caveat the analysts kept raising (shared infrastructure proves shared serving, not identity) turned out not to matter here: the trail led to the right answer. Z.ai has not issued a separate formal statement of its own, so the confirmation rests on the preview host's slug graduation. The remaining open question is data governance for the preview period: while the operator was anonymous and retained prompts, users could not name the company holding their inputs (now known to be Z.ai).
CONFIRMED, and it was the dominant theory all along. When the preview closed, OpenRouter retired the stealth slug and resolved it to z-ai/glm-5.3-flash, confirming this attribution at the host layer. The forensics that pointed here: tokenizer probes matched GLM-5.3 exactly (30/30 and separately 95/95, +75-token wrapper); video encoder matched GLM-5V-Turbo token-for-token; rejects audio like GLM-5V; shares GLM error codes 1214 and 1301; a leaked Java stack trace exposed 'com.wd.paas.api.domain.v4.chat.ChatCompletionRequest' matching Zhipu's /api/paas/v4 route (Chetaslua rated operator-layer ID at 0.98); Ben Davis said he was '99% certain'; Manifold priced Z.ai ~80%. Z.ai previously ran GLM-5 anonymously as 'Pony Alpha.' (Z.ai has not published a separate formal statement.)
Weighed and ruled out. The main non-Chinese theory: Wccftech initially suggested GLM, then updated to say Ox Alpha 'could be an unreleased version of Microsoft's MAI.' It was never backed by the infrastructure fingerprints that supported the Z.ai case, and the host reveal (GLM-5.3 Flash) settled it against MAI.
Weighed and ruled out. Early community sentiment thought it was a Gemini model, but this was dismissed as 'community sentiment, not infrastructure proof' and rejected by the fingerprinting crowd (one skeptic: it 'can't score below 3.7 Flash to be next gen'). The reveal (GLM-5.3 Flash) confirmed it was not Gemini.
Weighed and ruled out. Floated because Xiaomi's MiMo team 'has previewed unbranded models before,' but priced ~1% on Manifold and 'ruled out' by explainx.ai; video-token behavior was 'distinctly different' from MiMo v2.5, which also accepts audio. The reveal (GLM-5.3 Flash) confirmed it was not MiMo.
Weighed and ruled out. Priced ~3% on Manifold, with one commenter noting Cursor 'has done it in the past,' suggesting a fine-tune of an existing model rather than an original frontier model. A minor theory, and the reveal (GLM-5.3 Flash) closed it out.
The load-bearing caveat throughout the hunt. unclecode (builder of the 'modelprint' tool) emphasized that 'matching fingerprints prove shared infrastructure, not identity' and 'a lab can serve two different models on the same stack,' and flagged evidence pointing to OpenAI's cl100k_base tokenizer encoding that 'sits oddly on a Chinese model.' By Aug 22 analyst Andrew Curran noted people were 'less sure of anything.' In the end the fingerprints and the reveal agreed (GLM-5.3 Flash), but the discipline of holding 'shared serving is not identity' is exactly why the conclusion is credible rather than lucky.
August 14, 2026
Zhipu / Z.ai releases GLM-5.3 officially, six days before Ox Alpha appears (context for the fingerprinting theory).
August 20, 2026
Ox Alpha appears free on OpenRouter as stealth/ox-alpha ($0/$0, 1,048,576-token context, text/image/video in). OpenCode simultaneously launches it, promising it 'free for the next week' and advertising ~100 trillion tokens/day capacity.
August 21, 2026
Developer Ben Davis's fingerprinting reports ~99% certainty of a Zhipu GLM-5.x connection (tokenizer, video encoder, audio rejection). Viral benchmark wave: ~80% on a 10-task DeepSWE subset beating Fable 5, GLM-5.3, Grok 4.6, and GPT-5.6 Sol.
August 22, 2026
Speculation shifts: Andrew Curran notes the community is 'less sure of anything'; a Microsoft MAI theory emerges. The Next Web publishes on the anonymous-provider / EU AI Act data-retention concern. Ox Alpha still has no entry on Artificial Analysis or LMArena.
August 23, 2026
Mainstream coverage peaks: TechCrunch ('Who's behind the new stealth model Ox Alpha?') and Bloomberg. Stripe CEO Patrick Collison calls it 'very impressive.' Ox Alpha reaches No. 2 on OpenCode (~16T tokens, ~221k users). Ben Davis's full 113-task DeepSWE run corrects the score down to ~63%.
August 24, 2026
Ox Alpha remains $0 on both OpenRouter and OpenCode; the creator is still officially unconfirmed.
August 27, 2026
The ~one-week free window, from OpenCode's launch note, reaches its projected close. No exact cutoff, timezone, or post-preview price had been published, so the community watched the live listing for the end rather than a clock.
August 28, 2026
The free preview closes and the case is closed. OpenRouter retires the 'stealth/ox-alpha' slug, which now resolves to Z.ai's GLM-5.3 Flash (z-ai/glm-5.3-flash), confirming the community's leading theory. The reveal message: 'Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash. Use it now.' (Observed 2026-08-28, re-confirmed 2026-08-29.)
You have seen the data and the live model. When you want a working product rather than an API call, Founden builds it: a company, a product, a 3D world, whatever you have in mind, without wiring up any of the infrastructure yourself.