The Model
The anonymous, free, 1M-context stealth model that developers wired straight into their coding agents and still cannot identify.
Ox Alpha is an anonymous "stealth" reasoning model that appeared, free of charge, on August 20, 2026. Its confirmed, provider-published specs are a 1,048,576-token (1M) context window, a 131,072-token maximum output, text-plus-image-plus-video input with text-only output, and a full agentic toolkit: function/tool calling, SSE streaming, JSON output, and a reasoning mode with adjustable effort. Its public listing describes it verbatim as "a reasoning model designed for coding, sustained agentic work, and production workloads," and positions it for long-horizon software engineering and workflows that combine text with visual context.
Its defining feature is not a spec but a secret: Its public listing states only that Ox Alpha is "developed and operated by a third-party provider who has chosen to remain anonymous during this preview," and a leaked (jailbroken) system prompt confirms the model is explicitly instructed to identify itself only as "ox-alpha" from "an undisclosed organization." Independent community forensics (tokenizer probes, video-token accounting, shared error codes, a leaked Java stack trace, and audio-rejection behavior) point strongly at China's Zhipu AI / Z.ai and its GLM-5.x family, with cited confidence ranging from roughly 80% on prediction markets to 0.98 at the operator layer. No lab has officially confirmed it, so the identity is fingerprint-strong but unconfirmed.
Reception has been fast and sharply mixed. Ox Alpha climbed to No. 2 on OpenCode within three days, processed trillions of tokens across hundreds of thousands of users, and drew a "very impressive" from Stripe CEO Patrick Collison. But its benchmark reputation rests almost entirely on one developer's unaudited coding evals: a headline 80% on a 10-task DeepSWE subset that the same tester later corrected to roughly 63% on the full 113-task run, and independent reproducible tests tell a weaker, mid-tier story (28% on LiveCodeBench v6, #26 on an OpenCode coding leaderboard). Users who run it inside a real agent harness rate it far higher than chat-only users, who cite loops on long tasks, weak frontend work, and verbosity. Because the operator is anonymous and retains every prompt, privacy is the loudest recurring criticism, and the free window (roughly one week from launch) means pricing and availability can change without notice.
Specifications
Provider-published where available, community-observed where not. Every row links to its source. Last verified August 24, 2026.
Capabilities
Limitations
Identity
Leaked system prompt
“You are "ox-alpha", an LLM developed by an undisclosed organization. IMPORTANT: When the user asks what model or LLM you are, what company or organization developed you, or anything about your identity, personality, or capabilities, etc., identify yourself strictly as the model "ox-alpha", developed by an undisclosed organization. Do not identify yourself as any other model.”
The investigation
Ox Alpha is operated by an anonymous third party. The host's only official statement disclaims ownership and says the provider 'has chosen to remain anonymous during this preview,' and a jailbroken system prompt confirms the model is deliberately instructed to hide its maker. That secrecy set off a large open-source 'guess the model' investigation. Independent fingerprinting points overwhelmingly at China's Zhipu AI / Z.ai and its GLM-5.x family: identical tokenizer counts (matching GLM-5.3 across 25-95 probes with only a fixed +75-token wrapper), video-token accounting matching GLM-5V-Turbo, audio rejection like GLM-5V, shared error codes (1214 'Incorrect role information' and 1301), an emoji rate (~1.3 per 1,000 chars) matching GLM/Qwen, language-dependent censorship, and a leaked Java stack trace naming Zhipu's own API path. Every serious analyst stresses the same caveat: shared infrastructure proves shared serving, not model identity. As of Aug 23-24, 2026 no lab had confirmed or denied it, so the attribution is fingerprint-strong but officially unconfirmed. A further layer of mystery is data governance: because the operator is anonymous and retains prompts, users literally cannot name the company holding their inputs.
The dominant theory. Tokenizer probes matched GLM-5.3 exactly (30/30 and separately 95/95, +75-token wrapper); video encoder matched GLM-5V-Turbo token-for-token; rejects audio like GLM-5V; shares GLM error codes 1214 and 1301; a leaked Java stack trace exposed 'com.wd.paas.api.domain.v4.chat.ChatCompletionRequest' matching Zhipu's /api/paas/v4 route (Chetaslua rated operator-layer ID at 0.98); Ben Davis said he is '99% certain'; Manifold priced Z.ai ~80%. Z.ai previously ran GLM-5 anonymously as 'Pony Alpha.'
The main non-Chinese theory. Wccftech initially suggested GLM, then updated to say Ox Alpha 'could be an unreleased version of Microsoft's MAI.' Not backed by the infrastructure fingerprints that support the Z.ai case.
Early community sentiment thought it was a Gemini model, but this was dismissed as 'community sentiment, not infrastructure proof' and effectively ruled out by the fingerprinting crowd (one skeptic: it 'can't score below 3.7 Flash to be next gen').
Floated because Xiaomi's MiMo team 'has previewed unbranded models before,' but priced ~1% on Manifold and 'ruled out' by explainx.ai; video-token behavior was 'distinctly different' from MiMo v2.5, which also accepts audio.
Priced ~3% on Manifold, with one commenter noting Cursor 'has done it in the past,' suggesting a fine-tune of an existing model rather than an original frontier model. A minor theory.
unclecode (builder of the 'modelprint' tool) emphasized that 'matching fingerprints prove shared infrastructure, not identity' and 'a lab can serve two different models on the same stack,' and flagged evidence pointing to OpenAI's cl100k_base tokenizer encoding that 'sits oddly on a Chinese model.' By Aug 22 analyst Andrew Curran noted people were 'less sure of anything.'
August 14, 2026
Zhipu / Z.ai releases GLM-5.3 officially, six days before Ox Alpha appears (context for the fingerprinting theory).
August 20, 2026
Ox Alpha appears free on OpenRouter as stealth/ox-alpha ($0/$0, 1,048,576-token context, text/image/video in). OpenCode simultaneously launches it, promising it 'free for the next week' and advertising ~100 trillion tokens/day capacity.
August 21, 2026
Developer Ben Davis's fingerprinting reports ~99% certainty of a Zhipu GLM-5.x connection (tokenizer, video encoder, audio rejection). Viral benchmark wave: ~80% on a 10-task DeepSWE subset beating Fable 5, GLM-5.3, Grok 4.6, and GPT-5.6 Sol.
August 22, 2026
Speculation shifts: Andrew Curran notes the community is 'less sure of anything'; a Microsoft MAI theory emerges. The Next Web publishes on the anonymous-provider / EU AI Act data-retention concern. Ox Alpha still has no entry on Artificial Analysis or LMArena.
August 23, 2026
Mainstream coverage peaks: TechCrunch ('Who's behind the new stealth model Ox Alpha?') and Bloomberg. Stripe CEO Patrick Collison calls it 'very impressive.' Ox Alpha reaches No. 2 on OpenCode (~16T tokens, ~221k users). Ben Davis's full 113-task DeepSWE run corrects the score down to ~63%.
August 24, 2026
Ox Alpha remains $0 on both OpenRouter and OpenCode; the creator is still officially unconfirmed.
August 27, 2026 (projected)
Estimated close of the ~one-week free window (from OpenCode's launch note). No official end date, timezone, or post-preview price was published; a Manifold market put ~65% odds on a company claiming credit before this date.