The questions everyone asks about the stealth model the community cracked as Z.ai's GLM-5.3 Flash, answered honestly and sourced from public reporting.
What is Ox Alpha?
Ox Alpha was an anonymous 'stealth' reasoning model that appeared free on August 20, 2026. When the preview closed in late August 2026, the host that ran it (OpenRouter) revealed it as Z.ai's (Zhipu AI's) GLM-5.3 Flash. It has a 1,048,576-token (1M) context window, accepts text, images, and video, outputs text, and is positioned for coding, sustained agentic work, and production workloads.
Who made Ox Alpha?
Z.ai (Zhipu AI). Ox Alpha was Z.ai's GLM-5.3 Flash, run anonymously during a stealth preview. When the preview closed, OpenRouter (the host that ran it) retired the 'stealth/ox-alpha' slug and resolved it to z-ai/glm-5.3-flash, confirming the community's leading fingerprinting theory. Z.ai has not published a separate formal statement of its own, so the identification comes from the preview host's slug graduation.
Is Ox Alpha really free?
Not anymore. The upstream stealth preview that was $0 across several providers has ended, and this site's own free chat and API, which we subsidized, were retired on September 30, 2026. The revealed model, GLM-5.3 Flash, is now a cheap paid model (Z.ai lists it at $0.15/1M input and $0.50/1M output), and its MIT-licensed open weights are free to download and self-host. For every current option, see where to run GLM-5.3 Flash now.
How good is Ox Alpha at coding, really?
Mixed and disputed. Its viral headline was 80% on a hand-picked 10-task DeepSWE subset, but the same tester corrected that to roughly 63% on the full 113-task set (mid-tier, near GPT-5.6 Sol mid and Opus 4.8). Independent reproducible tests are weaker still: 28% on LiveCodeBench v6 and #26 on an OpenCode coding leaderboard. It performs notably better inside a real agent harness with tools than in plain chat.
Is it safe to use Ox Alpha with sensitive data?
Treat it with care. The operator, now identified as Z.ai (Zhipu AI), retained prompts and completions during the stealth preview, and the preview terms were ambiguous about training. As a general rule, do not send personal data, secrets, credentials, private source code, customer data, or regulated data to any model without a data agreement you can rely on. Use sanitized test data, and route anything you would not want logged through a separate, attributable model with clear terms.
Does Ox Alpha train on my data?
The model's public listing states the provider retains prompts and completions but does not use them for training. However, this conflicts with the broader Stealth Program agreement (which reportedly permits collection for training) and with a competing 'zero data retention' claim on another host. The terms are ambiguous, so follow the strictest boundary.
What is the context window and output limit?
The context window is exactly 1,048,576 tokens (1M), a provider-published figure confirmed across multiple platforms. The maximum output is 131,072 tokens (~128K), and that cap includes reasoning tokens. The 1M input is separate from and much larger than the 128K output cap.
How do I access Ox Alpha?
This site's own free chat and API were retired on September 30, 2026. To call it from your own code, the model is now Z.ai's GLM-5.3 Flash (slug z-ai/glm-5.3-flash on OpenRouter, or Z.AI's own API); the stealth 'ox-alpha' slug has retired. You can also self-host the MIT-licensed open weights (zai-org on Hugging Face) with a runtime like vLLM or SGLang. Most SDKs work by just swapping the base URL. For every current option, see where to run GLM-5.3 Flash now.
Is there a free GLM-5.3 Flash API?
Not on this site anymore. Until September 30, 2026 this site ran a free, OpenAI-compatible GLM-5.3 Flash API (model id ox-alpha); it is now retired, and its endpoint returns HTTP 410 Gone. To call GLM-5.3 Flash today, use Z.ai's own OpenAI-compatible API or OpenRouter (z-ai/glm-5.3-flash), or self-host the MIT-licensed open weights, which are free to download. If your code called this site's endpoint, see the API migration guide.
Is GLM-5.3 Flash open source, and can I run it myself?
Yes. After the reveal, Z.ai released GLM-5.3 Flash as open-weight under the permissive MIT license (weights published on Hugging Face as zai-org/GLM-5.3-Flash). It is a 320B-total, 18B-active (320B-A18B) mixture-of-experts model with a 1M-token context, so you can self-host it on your own hardware or run it through any provider that hosts it. Note this is the Flash model specifically; the larger GLM-5.3 is a separate model released under a different license. Before committing to hardware, see how 2026's top open-weight models compare on GPU cost and license on Founden.ai, an AI builder that Ox Alpha's owner also operates.
Why did people think Ox Alpha was a GLM (Zhipu) model, and were they right?
They were right. Multiple independent tells aligned with Zhipu's GLM: tokenizer counts matched GLM-5.3 exactly (with only a fixed +75-token wrapper), video-token consumption matched GLM-5V-Turbo, it rejects audio like GLM-5V, it shares GLM's 1214 and 1301 error codes, its emoji rate matches GLM/Qwen, and a leaked Java stack trace named Zhipu's own API path. Z.ai also had a precedent of anonymous previews (it ran GLM-5 as 'Pony Alpha'). When the preview closed, OpenRouter confirmed the model as Z.ai's GLM-5.3 Flash, vindicating the forensics.
What are Ox Alpha's main weaknesses?
The most common complaint is getting stuck in loops or stalling on long tasks. Others include weak frontend/CSS and visual reasoning, verbose output, overthinking at high reasoning effort, middling interactive latency (~11.6s median agent turn), a ~4.45% tool-call error rate, and 503 errors or dropped streams under viral load. It also has no audio input and no schema-enforced JSON.
Is Ox Alpha ranked on official leaderboards like Artificial Analysis or LMArena?
Now, yes. During the stealth preview it had no entry, but after the reveal and open-weight release Artificial Analysis lists GLM-5.3 Flash at an Intelligence Index of 57 (v4.1.1, about 46.7 tokens/sec), which sits just below the frontier band (60-63) and level with Claude Opus 4.8 at a fraction of the cost. Its strongest signal is still adoption: No. 2 on OpenCode within three days of launch.
Does Ox Alpha support tool calling and reasoning?
Yes. It supports function/tool calling (tools and tool_choice), SSE streaming, JSON output via response_format (without strict schema enforcement), and a reasoning mode with adjustable effort. It is OpenAI-compatible, so most SDKs work by swapping the base URL. Note the reasoning-effort levels are reported inconsistently: AI/ML API documents low/medium/high, while one fingerprint report lists max/high/low with no 'medium.'
Was Ox Alpha's creator ever revealed?
Yes. When the preview closed in late August 2026, OpenRouter (the host that ran it) retired the 'stealth/ox-alpha' slug and identified the model as Z.ai's GLM-5.3 Flash, pointing users to z-ai/glm-5.3-flash. That confirmed the community's leading theory, which prediction markets had priced around ~80% for Z.ai. Z.ai has not issued a separate formal announcement of its own, so the confirmation rests on the preview host's slug graduation.
You have seen the data and the live model. When you want a working product rather than an API call, Founden builds it: a company, a product, a 3D world, whatever you have in mind, without wiring up any of the infrastructure yourself.