How to Use Ox Alpha: API, Open Weights, and Build
August 28, 2026 · 24 min read · Ox Alpha
The working manual for the 1M-context model the community cracked as Z.ai's GLM-5.3 Flash: try it in a minute, wire it into your code, and know exactly where it breaks.
Ox Alpha appeared free on August 20, 2026, with a 1,048,576-token context window and an OpenAI-compatible API, and within three days developers had pushed trillions of tokens through it - OpenRouter. That combination (a full 1M context, a real agentic toolkit, and a price of $0 during the preview) is why so many people wired it straight into their coding agents before anyone knew who made it. When the preview closed in late August 2026, OpenRouter (the host that ran it) retired the "stealth/ox-alpha" slug and identified the model as Z.ai's (Zhipu AI's) GLM-5.3 Flash, exactly where the community forensics had pointed. This guide is the practical companion to that story: not "what is it" and not "who built it," but how you actually use it well, and how you turn an afternoon of testing into something real.
The honest framing up front, because it changes how you should use every feature below: the upstream stealth preview was a no-SLA experiment, and it has now closed. The revealed model, GLM-5.3 Flash, is a cheap paid model going forward (Z.ai lists it at $0.15 per 1M input and $0.50 per 1M output after launching at half that, and some OpenRouter hosts charge less; see every host's current price). This site's own free chat and API on the model were retired on September 30, 2026, so this guide now points you to where to run GLM-5.3 Flash now instead. The API is how you evaluate. It is not, on its own, how you ship. When you want to go from a promising test to a real product, a real company back-office, or a real 3D experience without wiring the infrastructure yourself, the fastest path is to build a product with Ox Alpha on a platform that owns the plumbing. We will come back to that, but hold it in mind as you read, because it is the difference between a weekend experiment and something people can use.
Contents
- The fastest ways to try Ox Alpha now
- The API quickstart: base URL, model id, and three code examples
- Prompting a 1M context (and the 131K output cap that trips people up)
- Where it actually shines: inside a real agent harness
- Set expectations before you build: the honest coding numbers
- The limits to plan around, and the fallback-model rule
- The safety rules that are not optional
- Build something real with Ox Alpha
- The bottom line
1. The fastest ways to try Ox Alpha now
Before you write a line of code, the goal is to form your own opinion of the model with the least possible setup, because the community reviews are so split that borrowing someone else's verdict is a mistake. Whichever route you pick below, judge it by how it behaves under tools and automation, where reviews consistently rate it higher, because chat-only impressions systematically undersell it.
Until September 30, 2026, the single fastest way was a free chat on this site's homepage, with a free API beside it; both are now retired. Today there are three routes in, all listed at where to run GLM-5.3 Flash now: Z.ai's own OpenAI-compatible API (base URL https://api.z.ai/api/paas/v4/) or its GLM Coding Plan inside coding tools; OpenRouter, where many hosts serve it under the model id z-ai/glm-5.3-flash; and the MIT-licensed open weights on Hugging Face (zai-org), which you can self-host with runtimes like vLLM or SGLang. The model's most famous quirk is history now: during the stealth preview, if you asked what model it was, it would identify itself only as "ox-alpha, developed by an undisclosed organization," because it was explicitly instructed to - OpenRouter. For the story behind that deliberate anonymity and the fingerprints that cracked the case (it was Z.ai's GLM-5.3 Flash), read who built Ox Alpha.
The route this guide is really about is the API. Ox Alpha is OpenAI-compatible, which means most existing SDKs and tools work by swapping the base URL and the model name, with nothing else changed - AI/ML API docs. That single design decision is why adoption spiked so fast: a developer already using the OpenAI client library could point it at Ox Alpha in under a minute. During the preview the model id was simply ox-alpha (and stealth/ox-alpha on OpenRouter); now that the preview has closed, OpenRouter's slug resolves to z-ai/glm-5.3-flash. The full spec sheet with every provider-published number lives on the model page. If you have not yet decided whether it is worth your time at all, the sober version of "what it is and whether to trust it" is in Ox Alpha, honestly.
2. The API quickstart: base URL, model id, and three code examples
The API is where Ox Alpha stops being a curiosity and starts being usable, and the good news is that there is almost nothing new to learn. Because the endpoint follows the OpenAI chat-completions schema, the three things you configure are the same three you configure for any OpenAI-compatible provider: a base URL that points at your chosen host, a Bearer token for auth, and the model id. During the preview the model id was ox-alpha (or stealth/ox-alpha); now that the preview has closed and the model is revealed as GLM-5.3 Flash, OpenRouter's slug is z-ai/glm-5.3-flash, so the only real decision is which host to route through. Our API page is now the migration guide for code that called this site's retired endpoint. The examples below use a placeholder base URL (openrouter.ai/api/v1 on OpenRouter, api.z.ai/api/paas/v4 on Z.ai's own API) and OpenRouter's model id, so on another host swap in the model id that host lists.
Several hosts served the same model, and now that the stealth slug has retired, the current identity is GLM-5.3 Flash. The canonical listing is on OpenRouter at https://openrouter.ai/api/v1, where stealth/ox-alpha graduated to z-ai/glm-5.3-flash - OpenRouter. AI/ML API served it at https://api.aimlapi.com/v1 behind an account - AI/ML API docs. Felo exposed both an OpenAI-compatible and an Anthropic-compatible base URL - Felo. Here is how the access routes shook out, with the preview pricing that applied during the free window and where each one lands now:
| Route | Model id (now) | Preview price | Best for |
|---|---|---|---|
| OpenRouter | z-ai/glm-5.3-flash | $0 / $0 in preview | The canonical listing; now $0.15 / $0.50 per 1M on Z.ai's own endpoint, less on some hosts |
| AI/ML API | stealth/ox-alpha | Account-gated | A managed OpenAI-compatible endpoint |
| AIHubMix | ox-alpha | $0 input/output/cache in preview | A cheap drop-in during the free window |
| Felo | ox-alpha | Free in launch window | OpenAI and Anthropic base URLs |
| OpenCode Zen | opencode/x-preview-f-free | Free | Using it inside the OpenCode CLI agent |
| ModelsLab | stealth-ox-alpha | From $21/mo | Note: no real free tier despite $0 listing |
A caveat that belongs right next to that table: the $0 columns were a preview subsidy, not a commitment, and that subsidy has ended. When the preview closed, OpenRouter retired stealth/ox-alpha and the slug now resolves to z-ai/glm-5.3-flash, a cheap paid model. Z.ai's own endpoint launched it at $0.075 / 1M input and $0.25 / 1M output, then moved to its $0.15 / $0.50 list price in September 2026, while several hosts charge less. That move is exactly why you should treat the live listing (or our live GLM price table) as the only reliable signal, and never hardcode a per-token price into a budget. A fuller pricing roundup, including the independent hosted UIs, is tracked at oxalpha.online and analyzed at CellCog.
Now the code. The cURL example is the fastest sanity check that your key and base URL work, before you pull in any SDK:
curl https://YOUR-BASE-URL/chat/completions \
-H "Authorization: Bearer $OX_ALPHA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/glm-5.3-flash",
"messages": [
{ "role": "user", "content": "Refactor this function for readability and explain why." }
],
"reasoning_effort": "medium",
"stream": false
}'
Once that returns a completion, the Node version using the official openai SDK is a two-line change from any existing OpenAI setup: swap baseURL and model. Everything else, streaming, tool definitions, message shape, is identical - AI/ML API docs:
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.OX_ALPHA_API_KEY,
baseURL: "https://YOUR-BASE-URL",
});
const res = await client.chat.completions.create({
model: "z-ai/glm-5.3-flash",
messages: [
{ role: "user", content: "Read this repo and list the three most likely bugs." },
],
reasoning_effort: "high",
});
console.log(res.choices [0].message.content);
The Python version is the same idea, and it is the one most agent frameworks sit on top of. Note the explicit max_tokens: the completion cap includes reasoning tokens, so on a heavy task you want headroom, which we cover in the next section:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ ["OX_ALPHA_API_KEY"],
base_url="https://YOUR-BASE-URL",
)
resp = client.chat.completions.create(
model="z-ai/glm-5.3-flash",
messages= [
{"role": "user", "content": "Summarize this 150-page spec into an implementation plan."},
],
max_tokens=16384,
stream=True,
)
for chunk in resp:
print(chunk.choices [0].delta.content or "", end="")
Those three snippets cover the vast majority of real usage, and they all rely on the same confirmed capabilities: function and tool calling with tool_choice control, SSE streaming with optional in-stream usage reporting, and JSON output via response_format - AI/ML API docs. One implementation detail matters in production: the JSON mode is json_object without strict schema enforcement, so you must validate structured responses in your own code rather than trusting the model to honor a schema. If you skip that validation, a single malformed object will crash a pipeline that looked fine in testing.
3. Prompting a 1M context (and the 131K output cap that trips people up)
The defining spec is the 1,048,576-token context window, exactly one million tokens (2 to the power of 20), and using it well is the single biggest lever you have on output quality - OpenRouter. Most models force you to choose what to show them; Ox Alpha lets you show it the whole thing. That changes the shape of a good prompt. Instead of hand-selecting the three files you think are relevant and hoping you guessed right, you can load an entire repository, a full specification, or a complete transcript into working memory and let the model do the retrieval. In practice this is where the strongest reviews come from: one user cloned and reasoned over the entire Kanidm repository in a single session - Elixir Forum.
The structural reason this works is worth internalizing, because it tells you when to reach for Ox Alpha over a smaller-context model. A 1M window collapses a class of problems that otherwise require an external retrieval system. Whole-repository reasoning, large-project migration across 100K to 200K token working sets, and cross-file bug hunting all become single-prompt tasks instead of orchestration problems - OrcaRouter. When your question genuinely depends on the interaction between many files (a refactor that ripples, a bug that only appears across a boundary), giving the model everything beats giving it a curated slice, because you are not pre-deciding what matters. For the deeper explanation of what the model is and where that context window came from, see what is Ox Alpha.
The trap that catches people is the output cap, and it is a different number from the context. The maximum completion is 131,072 tokens (about 128K), and, critically, that ceiling includes reasoning tokens - OpenRouter. So the 1M is how much the model can read; the 128K is how much it can write in one response, and the model's own internal reasoning eats into that same budget. If you ask a high-reasoning-effort question that expects a very long answer, the reasoning trace can consume a large share of the cap before the visible answer even starts. The fix is to prompt for scoped outputs: ask for a plan, then the code for one module, then the next, rather than "rewrite the whole thing in one shot." You get more reliable results and you never hit the ceiling mid-answer.
- Load whole, retrieve narrow. Put the entire codebase in context, then ask a specific question about it.
- Scope the output, not the input. A 1M input is fine; a 128K single response is a hard wall, so decompose long generations.
- Watch reasoning effort. High effort improves hard problems but consumes output budget and can cause overthinking.
- Anchor with the real files. Because it can ignore existing code, paste the actual current version rather than describing it.
Those four habits map directly onto the model's documented behavior, and skipping them is the source of most "it did something weird" complaints. Loading whole and retrieving narrow plays to the 1M strength; scoping the output avoids the reasoning-plus-answer collision at the 128K cap; moderating reasoning effort sidesteps the documented tendency to overthink and pad; and anchoring with real files counters the observed habit of ignoring existing code in favor of a fresh rewrite. None of these are exotic prompt-engineering tricks. They are just the specific accommodations this specific model rewards, and they are the difference between the "extremely impressive" reviews and the "it got stuck" ones.
4. Where it actually shines: inside a real agent harness
Here is the most important practical fact in this entire guide, and it is one that a casual chat test will hide from you: users who run Ox Alpha inside a real agent harness with tools rate it far higher than chat-only users do - The Cherry Creek News. The gap is not subtle. In plain chat, direct code generation looks middling. Wired into an agent loop with a terminal, a filesystem, and the ability to run its own tests, the same model produced one documented run of 69 tool calls with a single error and no retry loop - OrcaRouter. That is the difference between judging a model by how it talks and judging it by what it can do.
The first-principles reason for the split is that a coding benchmark in chat mode asks the model to emit correct code in one pass, with no feedback, while an agent harness lets the model do what human engineers do: write, run, read the error, and fix. A model with strong long-context reasoning but imperfect one-shot generation, which is exactly what the numbers below describe, will look dramatically better in the second setting because the harness converts its reasoning into an iterative loop. This is why the model's adoption was concentrated in agent frameworks. Coding agents pushed billions of tokens through it almost immediately, with Nous Research's Hermes Agent (~2.74T tokens) and Claude Code (~1.37T tokens) among the top consumers - The Next Web.
Because it is OpenAI-compatible, wiring it into a harness is a configuration change, not an integration project. The documented use cases are agentic software engineering inside tools like Claude Code, OpenCode, Zed, and Hermes Agent, each of which lets you point at a custom base URL and model - Kingy AI. In OpenCode specifically, the model debuted alongside OpenRouter and you select it by running /models and choosing the free Ox Alpha entry rather than hardcoding the alias, since the route id can change during the preview. The practical takeaway is simple and it should reshape your evaluation: if your first test is a chat window and you come away unimpressed, you have tested the weaker half of the model. Test it where it lives, inside a loop with tools, before you form a verdict.
5. Set expectations before you build: the honest coding numbers
You cannot decide how to use a model responsibly without an honest read on how good it actually is, and Ox Alpha's coding story is the most misreported thing about it. The brand of this site is that we always show you all three numbers, because each one is true and each one is incomplete on its own. The viral headline was roughly 80% pass@1 on a hand-picked 10-task subset of DeepSWE, where it beat Claude Fable 5, GLM-5.3, Grok 4.6, and GPT-5.6 Sol - oxalpha.com. That is the number that launched the "topped the coding benchmarks" headlines, and the same tester explicitly warned that 10 tasks carry high variance and should be read as directional only.
When that same tester ran the full 113-task DeepSWE set, the score corrected down to about 63% pass@1, described as on par with GPT-5.6 Sol mid-tier and near Claude Opus 4.8, which is a genuinely respectable result but not a record - The Cherry Creek News. Then an independent, reproducible run on LiveCodeBench v6 (greedy, temperature 0, single attempt, zero generation failures) put it at just 28% pass@1, well below the frontier and a direct contradiction of the "frontier coder" framing - xnasarx on GitHub. All three are real. The honest synthesis is that Ox Alpha is a capable mid-tier coder that punches above its measured weight inside an agent harness and below it in one-shot generation.
Two facts complete the picture and both should temper any plan to bet a product on the raw model. First, Ox Alpha has no entry on Artificial Analysis or LMArena, and its own listing carries no capability benchmark, so there is no authoritative aggregator score to point to - Local AI Zone. Second, on the LM Market Cap composite it ranked #392 of 394 on coding, a number that reflects the absence of verified capability data more than a measured floor, but which underscores how thin the independent evidence still is - LM Market Cap. For the fully charted, side-by-side version of every score, see the benchmarks page; for the model-by-model matrix against GLM-5.3 and the rest of the frontier, see the comparison and the guide Ox Alpha vs the field.
6. The limits to plan around, and the fallback-model rule
Every model has failure modes; the reason Ox Alpha's matter more than usual is that you cannot call anyone when you hit them. Planning around the limits is therefore not optional polish, it is the core of using this model safely. The single most consistent real-user complaint is that it gets stuck in loops or stalls on long tasks until manually nudged, with one user watching it apologize for a mistake and immediately repeat it - Elixir Forum. This is precisely why the agent-harness pattern helps: a harness with a step limit and a human-in-the-loop nudge turns a silent stall into a recoverable event, whereas a fire-and-forget batch job just hangs.
The other documented weaknesses cluster into a short, specific list, and knowing them lets you route around them rather than discover them in production. Keep each in mind as a boundary, not a dealbreaker:
- Weak frontend and CSS. Visual reasoning is a known soft spot; a Hacker News reviewer who called it impressive still flagged poor visual reasoning - Hacker News.
- Verbosity and overthinking. It produces verbose output and overthinks at high reasoning effort; one user called its Elixir "the most verbose you will ever see."
- Middling latency. A median multi-step agent turn runs around 11.6 seconds, so it is better for background work than snappy interactive UX - ox-alpha.net.
- Tool-call errors. A ~4.45% tool-call error rate means roughly one in twenty tool calls needs a retry path in your harness.
- 503s under load. During the viral free window, traffic produced 503 errors, dropped streams, and throughput dips to ~20-30 tok/s within the first day.
Read those together and one rule falls out that is non-negotiable for anything beyond experimentation: wire in a fallback model. These failure modes were observed during the stealth preview, and even now that the model is revealed as GLM-5.3 Flash, a cheap high-throughput model is exactly the kind you want a safety net behind: any system that depends on a response must be able to fail over on a transient error, a timeout, or a detected loop. The clean way to do this is at the router layer: try your primary model first, and on failure retry the request against a second model with a contract behind it. This is also the honest reason a raw model API is an evaluation tool and not a production foundation on its own, and it is exactly the gap that a managed build platform closes for you, which is the next section.
7. The safety rules that are not optional
The safety story for Ox Alpha is unusual because the risk is not really about the model's outputs, it is about who is holding your inputs. During the stealth preview the operator was anonymous and retained every prompt and completion, and users could not name the company holding their data - TechTimes. That operator is now identified as Z.ai (Zhipu AI), a China-based lab, but the caution does not evaporate: anything you sent during the preview was retained by a provider you could not audit at the time, and Z.ai's own data terms are what govern the model going forward. The preview data terms were also ambiguous and conflicting: the canonical listing said prompts were retained but not used for training, another host marketed "zero data retention," and the broader Stealth Program agreement reportedly permitted collection for training, so the only defensible posture is to follow the strictest boundary and assume anything you send is kept - OpenRouter.
The practical consequences are simple to state and hard to walk back if you ignore them. One Hacker News user memorably likened free use of an anonymous provider to getting free steak smuggled out of a grocery store, and the sober advice underneath the joke is exactly right: never paste secrets into a host whose data handling you have not vetted - Hacker News. Now that the operator is confirmed as a China-based lab, the documented supply-chain and geopolitical-data concerns and the observed language-dependent censorship are concrete rather than hypothetical, and they matter for regulated or sensitive workloads - SiliconANGLE. Here are the boundaries to treat as hard lines:
- Never send secrets or credentials. No API keys, tokens, passwords, or connection strings, ever.
- Never send PII or customer data to a model without a data agreement you can rely on.
- Never send private source or regulated data. Use sanitized, non-sensitive test data only.
- Assume prompts are retained. Route anything you would not want logged through a separate model with clear terms.
- Know the upstream free window has closed. It ran roughly one week from the August 20 launch; when it closed the model was revealed as GLM-5.3 Flash, now a cheap paid model. - CellCog
Those five rules are the price of admission for using any model whose data terms you cannot fully control, and they are not burdensome once you internalize the frame: Ox Alpha is a brilliant place to test ideas with throwaway data and a poor place to put anything you cannot afford to hand to a third party under China-based data terms. The FAQ answers the specific data-and-safety questions in more depth, and the model page carries the full, sourced data-policy row with its caveats. If a workload cannot tolerate any of the five lines above, that workload belongs on a model you can contract with under terms you have vetted.
8. Build something real with Ox Alpha
Everything so far has been about evaluation: a first run to get a feel, API to test under tools, benchmarks to set expectations, limits and safety rules to know the boundaries. But evaluation is not the goal. The goal is to build something, and this is where a raw model API runs out of road, because a real product needs the things a single model endpoint cannot give you: a fallback when the model 503s, auth and a database for your users, billing if you charge, and a deploy target that stays up. Wiring all of that yourself is weeks of undifferentiated plumbing before you write a single line of the thing you actually care about.
That is exactly the gap founden.ai/app/build is built to close, and since Founden is run by the founder of this site, read what follows as our own description rather than a neutral review. You describe what you want, and it stands up the whole application (the website, the signed-in product, the back-office, the database, the auth, the billing, the deploy) as real code you own, with the model wired in behind a router that can fall back when it needs to. Testing it through the API is how you decide Ox Alpha is worth using; build a company with Ox Alpha is how you turn that decision into a live product without becoming an infrastructure engineer first. And because it generates a real codebase rather than a locked template, you are not choosing between "easy" and "yours," which is usually the trade you are forced to make.
The honest caveat, in keeping with the rest of this guide, is that a build platform does not repeal the safety rules from the last section. Now that Ox Alpha is confirmed as Z.ai's GLM-5.3 Flash, you inherit its China-based data terms and its high-throughput-but-not-frontier reality, so keep sensitive data out and keep a fallback wired regardless of how you deploy. What the platform changes is the effort, not the physics: the plumbing is handled, but the judgment about what to send any model is still yours. Used that way, you can build a 3D world with Ox Alpha or a straightforward SaaS with equal ease, evaluating the model on the merits and shipping only what its limits genuinely support.
9. The bottom line
Using Ox Alpha well comes down to a sequence, not a single action. Pick a route first: Z.ai's own API, OpenRouter (z-ai/glm-5.3-flash), or the self-hosted open weights, all listed at where to run GLM-5.3 Flash now. Test it with a real task right away, because it is OpenAI-compatible and drops into existing SDKs; if your code used this site's retired ox-alpha endpoint, the API page is the migration guide. Prompt the 1M context by loading whole and retrieving narrow, while respecting the 128K output cap that includes reasoning tokens. And above all, test it inside an agent harness, because that is where the same model that looks middling in chat produces the runs that made people call it impressive.
Then build honestly. Set expectations against the real numbers (80% viral, ~63% on the full run, 28% independent), wire a fallback because it stalls and 503s under load, and never send a model anything you cannot afford to have retained under terms you have not vetted. If those constraints fit your use case, Ox Alpha (now known to be Z.ai's GLM-5.3 Flash) is a genuinely capable mid-tier model and worth the time you spend on it. For the fuller context, the honest guide covers whether to trust it, who built Ox Alpha covers how the case was cracked, and Ox Alpha vs the field places it against the current frontier. When you are ready to turn a promising test into something people can actually use, create something with Ox Alpha and let the platform own the plumbing while you own the idea.
This guide, like the rest of this site, is written to be useful rather than promotional, in the same spirit that drives founder Yuma Heymans (@yumahey), co-founder of HeroHunt.ai, who works at the intersection of AI and the tools developers actually build with.
This guide reflects the picture as of August 29, 2026. When the stealth preview closed, OpenRouter (the host that ran it) retired the "stealth/ox-alpha" slug and identified the model as Z.ai's GLM-5.3 Flash (z-ai/glm-5.3-flash), confirming the community's leading theory. Z.ai has not issued a separate formal statement of its own. Prices, access routes, and terms can still change, so treat the live listing as the only reliable signal and verify before you depend on anything here.