Ox Alpha vs GLM, Claude, and GPT (2026)
August 28, 2026 · 28 min read · Ox Alpha
The buyer's scorecard for the stealth model the community cracked as Z.ai's GLM-5.3 Flash, placed honestly against the 2026 frontier.
On a scorecard weighted for what a real buyer needs, Ox Alpha finishes last of nine, not because it is bad, but because two of the five things that mattered during its stealth week (verified capability and a nameable, accountable vendor) were exactly the two it could not supply at the time. That single sentence is the whole comparison, and the rest of this guide is the sourced proof behind it. Ox Alpha was the most-downloaded free coder of the week and, by any auditable measure, an unranked mid-tier model that wore a mystery costume until the preview closed and revealed it as Z.ai's GLM-5.3 Flash: the community called it, and this scorecard's own logic pointed at exactly that maker.
The problem with almost every "Ox Alpha vs" post is that it grades on the wrong curve. It cites the viral 80% coding number, notes the 1M context, waves at the free access, and calls it a frontier killer - oxalpha.com. None of those three facts is wrong. All three are also the least decision-relevant facts about a model you would actually put into production, because two of them were temporary (the stealth preview has since closed) and one of them is a hand-picked sample. This guide fixes the curve, and its conclusion held up: when the preview ended, OpenRouter (the host that ran it) retired the "stealth/ox-alpha" slug and identified the model as Z.ai's GLM-5.3 Flash, exactly where the forensics pointed.
What this covers: a weighted scorecard across the eight models the community actually benchmarks Ox Alpha against, an at-a-glance spec matrix, honest head-to-heads (including the GLM-vs-GLM problem that quietly invalidates half the comparisons online), a like-for-like DeepSWE chart with its caveat printed on it, and a decision tree for picking the right model for a real job. Every figure comes from the community and reported data collected on our benchmarks page; nothing here is invented.
Contents
- The weighted scorecard
- The comparison at a glance
- Ox Alpha vs GLM-5.3: the GLM-vs-GLM problem
- Ox Alpha vs Claude Opus 4.8 and Fable 5
- Ox Alpha vs GPT-5.6 Sol
- Ox Alpha vs Grok 4.6 and Gemini 3.7 Flash: the speed axis
- Ox Alpha vs DeepSeek V4 Pro and Qwen 3.8 Max: the value axis
- The DeepSWE field, charted (with the caveat)
- Which should you pick?
- The bottom line
1. The weighted scorecard
A comparison table is only honest if it says out loud what it is optimizing for, so start with the weights. A model buyer is not buying a benchmark, they are buying an outcome delivered repeatedly: code that ships, a long context that holds a whole repository, latency they can tolerate, a price they can forecast, and a vendor they can trust to still exist next quarter and to not walk off with their prompts. Those five things are not equally important, and they are certainly not equally supplied across the field. Coding capability carries the most weight because it is the job; trust and long-horizon context come next because they decide whether the capability is usable at all; speed and price matter but are the easiest to design around.
The scores below are 0 to 10, each cell carrying the real datapoint that earned it, and the whole table is sorted by the weighted final descending. It is the scorecard as the field looked during Ox Alpha's stealth week, which is exactly the window a buyer had to decide in. The result is uncomfortable for the hype and clarifying for everyone else: the models with named makers, published benchmarks, and real SLAs cluster at the top, and the anonymous free preview sits at the bottom of a table it would top if the only column were price. That inversion is the point of scoring instead of cheerleading. It is also why our full, continuously-sourced matrix lives on the compare page rather than in a single viral screenshot.
| # | Model | What It Does | Coding (30%) | Context (20%) | Speed (15%) | Price (15%) | Trust (20%) | Final |
|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.7 Flash | Fast, cheap, 1M-context coder from a named lab | 7 - 65.3% DeepSWE v1.1 | 9 - 1M context | 10 - ~340 tok/s, far fastest | 8 - $0.75 / $3.75 per 1M | 10 - Google, SLA, model card | 8.6 |
| 2 | DeepSeek V4 Pro | Open-weight (MIT) frontier coder, auditable | 9 - 93.5% LiveCodeBench v6, 80.6% SWE-bench | 9 - 1M context | 6 - no published tok/s | 9 - ~$0.66 / $1.98 per 1M | 8 - open weights, self-hostable | 8.4 |
| 3 | Claude Fable 5 | Highest measured coder, premium price | 10 - 95% SWE-bench, 69.7% official DeepSWE | 9 - 1M context | 6 - no published tok/s | 3 - $10 / $50 per 1M | 10 - Anthropic, SLA | 8.2 |
| 4 | GPT-5.6 Sol | Largest context, top-tier verified coding | 9 - 82.2% SWE-bench Verified | 10 - 1.05M context | 6 - no published tok/s | 4 - $4 / $20 per 1M | 10 - OpenAI, SLA | 8.2 |
| 5 | Claude Opus 4.8 | Prior Anthropic flagship, strong all-round | 9 - 88.6% SWE-bench, ~59% DeepSWE | 9 - 1M context | 6 - no published tok/s | 4 - $5 / $25 per 1M | 10 - Anthropic, SLA | 8.0 |
| 6 | GLM-5.3 | The prime fingerprint suspect, priced and named | 8 - 91.25% Kingbench, 66.9% DeepSWE | 9 - 1M context | 6 - no published tok/s | 7 - $1.40 / $4.40 per 1M | 7 - Zhipu, named, geopolitics | 7.6 |
| 7 | Qwen 3.8 Max | Alibaba flagship, strong Kingbench | 7 - 81.25% Kingbench | - not in dataset | - not in dataset | - not in dataset | 7 - Alibaba, named | 7.0 |
| 8 | Grok 4.6 | Fast, expensive, smaller context | 7 - 65.9% DeepSWE v1.1 | 6 - 500K context | 8 - ~67-86 tok/s | 4 - ~14.9x DeepSeek's cost | 9 - xAI, named, SLA | 6.9 |
| 9 | Ox Alpha | Anonymous free 1M-context preview | 6 - ~63% full DeepSWE, 28% LiveCodeBench | 9 - 1M context | 4 - ~25-50 tok/s, 11.6s agent turn | 10 - $0 / $0 (preview) | 2 - anonymous, no SLA, prompts retained | 6.1 |
The criteria: Coding capability (30%) is each model's best-defensible coding result, favouring reproducible, verified benchmarks over hand-picked subsets. Context (20%) is the usable working-memory window. Speed (15%) is real-world throughput and turn latency. Price (15%) is the standing per-token cost a buyer can forecast, not a temporary subsidy. Trust (20%) folds together whether the maker is named, whether there is an SLA and a model card, and what happens to your data. Where a value is genuinely absent from our dataset (Qwen's context, speed, and price), the cell is marked and excluded from that model's weighted average rather than guessed.
Read the table as a warning against single-number thinking. Ox Alpha's 10 on price is the highest score in the entire matrix, and it still finishes ninth, because a subsidy you cannot forecast and a vendor you could not name (at the time) cannot carry a workload by themselves. Meanwhile Gemini 3.7 Flash wins not by leading any single column but by having no weak one - Quesma. That is what "best buy" actually looks like when you stop grading on virality, and it is the frame the rest of this guide reasons from.
One update the reveal forces on this table, in the name of honesty: the "Ox Alpha" row is now known to be Z.ai's GLM-5.3 Flash, revealed when the preview closed. Two cells move slightly and neither changes the verdict. Trust rises above a 2 now that the maker is named (Z.ai, with a model card and a real endpoint), and Price eases off a perfect 10, since the standing cost is no longer $0: Z.ai now lists it at $0.15 / $0.50 per 1M after a half-price launch at $0.075 / $0.25, and some hosts charge less ( every host's current price). Cheap, not free. A named, cheap, mid-tier coder is a better product than an anonymous free one, but it is still a mid-tier coder: the coding, context, and speed columns are unchanged, so Ox Alpha stays near the bottom of a frontier scorecard while becoming a perfectly reasonable budget pick under its real name.
2. The comparison at a glance
Before the head-to-heads, it helps to see the raw axes side by side, because the scorecard compresses them and some readers will disagree with our weights. The honest way to place Ox Alpha is on the dimensions that are objective and provider-published (context, price, modality) next to each model's best-known coding number for rough orientation, with a loud reminder that those coding numbers come from different suites and are not strictly comparable. A DeepSWE percentage and a SWE-bench Verified percentage are different exams; lining them up in one column is orientation, not a leaderboard.
The matrix below does exactly that. Notice what Ox Alpha wins and what it does not. It ties for the largest context at 1M and won price outright at zero during the stealth preview, and on every other axis it either matches the pack or trails it. Nothing in this table supports the "beats the frontier" framing; everything in it supports "remarkable value while the subsidy lasted." The price column reflects the stealth window; since it closed, the "Ox Alpha" row resolves to Z.ai's GLM-5.3 Flash at $0.15 / $0.50 per 1M on Z.ai's list price, still cheap, no longer free. For the fully sourced, model-by-model version with every underlying link, the compare page is the canonical source, and the spec sheet carries each Ox Alpha figure with its provider citation.
| Model | Context | Price /1M (in / out) | Best-known coding | Named maker |
|---|---|---|---|---|
| Ox Alpha | 1M | $0 / $0 (preview); now $0.15 / $0.50 as GLM-5.3 Flash (Z.ai list) | ~63% DeepSWE (community) | Z.ai (revealed) |
| GLM-5.3 | 1M | $1.40 / $4.40 | 66.9% DeepSWE | Zhipu / Z.ai |
| Gemini 3.7 Flash | 1M | $0.75 / $3.75 | 65.3% DeepSWE | |
| GPT-5.6 Sol | 1.05M | $4 / $20 | 82.2% SWE-bench Verified | OpenAI |
| Claude Opus 4.8 | 1M | $5 / $25 | 88.6% SWE-bench Verified | Anthropic |
| Claude Fable 5 | 1M | $10 / $50 | 95% SWE-bench Verified | Anthropic |
| DeepSeek V4 Pro | 1M | ~$0.66 / $1.98 | 80.6% SWE-bench Verified | DeepSeek (MIT) |
| Grok 4.6 | 500K | expensive | 65.9% DeepSWE | xAI |
The single most important row-to-row lesson is the price column read against the coding column. DeepSeek V4 Pro delivers 80.6% SWE-bench Verified at roughly $0.66 input - Speedway Media, which is the number that should haunt anyone tempted to build on Ox Alpha for cost reasons: there is already a cheap, auditable, open-weight coder that scores far higher on a reproducible benchmark. Ox Alpha's price advantage was real but narrow, it was zero versus cheap during the preview, and it was temporary, whereas DeepSeek's is structural. Now that the preview has closed and Ox Alpha's real identity (GLM-5.3 Flash, $0.15 / $0.50 at Z.ai's list price) is priced like everything else, that "zero versus cheap" edge collapses to "cheap versus cheap," and DeepSeek's reproducible-benchmark lead is what remains. That is the kind of comparison the viral posts never run, and it is exactly the comparison a buyer needs.
3. Ox Alpha vs GLM-5.3: the GLM-vs-GLM problem
This is the head-to-head that quietly broke most of the internet's Ox Alpha coverage, and it deserves to go first because it changes how you read everything else, and because it is the one the reveal has now settled. Independent fingerprinting (tokenizer counts matching GLM-5.3 exactly, video-token accounting matching GLM-5V-Turbo, shared error codes, and a leaked Java stack trace naming Zhipu's own API route) pointed overwhelmingly at Zhipu's GLM family as the operator behind Ox Alpha - OrcaRouter. When the preview closed, that call was confirmed: OpenRouter retired the stealth slug and identified the model as Z.ai's GLM-5.3 Flash. We ranked that whole body of evidence, with every hole in it, in who built Ox Alpha. Because the attribution is now confirmed, a large share of "Ox Alpha vs GLM-5.3" benchmarks were effectively GLM comparing against a variant of itself, which is not a comparison at all.
That is not a rhetorical flourish, it has concrete consequences for the numbers. On Kingbench, GLM-5.3 scored 91.25% and Ox Alpha 87.5%, a near-match that the tester cited at the time as circumstantial evidence the two are the same family - daily.dev. When two models land four points apart on an unaudited benchmark and also share a tokenizer, the tight score is not independent corroboration of quality, it is a tell about lineage, and it turned out to be pointing at the truth. Conversely, on the 10-task DeepSWE subset Ox Alpha (80%) beat GLM-5.3 (62%) - oxalpha.com, which, given they are relatives (Ox Alpha being the GLM-5.3 Flash variant), most likely reflects different reasoning settings or a distinct checkpoint beating the larger shipped model on a tiny slice, not a genuinely new competitor dethroning an old one.
The practical takeaways from a now-confirmed same-family comparison come down to a few concrete choices.
- Price is the clean differentiator. During the stealth preview Ox Alpha was $0; that preview has closed, and the revealed model (GLM-5.3 Flash) now lists at $0.15 / $0.50 per 1M on Z.ai (it launched at $0.075 / $0.25, and some OpenRouter hosts charge less). GLM-5.3, the larger sibling, is $1.40 / $4.40 per 1M - OpenRouter.
- Modality is shared. Ox Alpha accepted image and video input, matching GLM-5.3 Flash's multimodal listing, which was always one of the honest tells it was a distinct, newer variant of the family rather than the base text route.
- Auditability now applies to both. GLM-5.3 and GLM-5.3 Flash each have a named maker (Z.ai), a model card, and a price you can plan around, which the anonymous "ox-alpha" slug did not offer while the preview ran.
What this means in practice is that "Ox Alpha or GLM-5.3?" was always a strange question, and the reveal makes it clearer: you were choosing between two members of the same GLM-5.3 family at two price points, the cheap fast Flash (which was masked as Ox Alpha) and the larger standard model. With the mask off, the sensible framing is simply "GLM-5.3 Flash or GLM-5.3," a normal size-versus-cost trade-off, not a mystery model versus a known one.
4. Ox Alpha vs Claude Opus 4.8 and Fable 5
Anthropic is where the comparison stops being about lineage and starts being about a real capability gap, so it is worth being precise about which Claude the community actually tested. The head-to-heads used Claude Opus 4.8, the prior Anthropic flagship, not the current one; Anthropic's present flagship is Claude Opus 5, which leads the Artificial Analysis Intelligence Index - OrcaRouter. That matters because even the older Opus 4.8 already sits comfortably above Ox Alpha on verified coding, and Fable 5 sits far above both. Comparing Ox Alpha to Claude is therefore a comparison to Anthropic's second-string and specialist models, and it still does not win.
On the numbers, Ox Alpha's full 113-task DeepSWE run of roughly 63% lands near Opus 4.8's ~59% on the same suite, and on the non-standard Kingbench it edged Opus 4.8 (87.5% vs 80%) - OrcaRouter. Read quickly, that looks like parity. Read carefully, it is not: Opus 4.8's headline is 88.6% SWE-bench Verified, a reproducible, widely-trusted benchmark, while Ox Alpha's most reproducible independent result is 28% on LiveCodeBench v6 - GitHub. When one model's strong number comes from an audited suite and the other's comes from a hand-picked subset, "parity on DeepSWE" is not parity, it is a coincidence of one exam.
Fable 5 removes any ambiguity. It posts 95% SWE-bench Verified and 69.7% on the official full DeepSWE leaderboard, where Ox Alpha was never officially evaluated - oxalpha.com. On the 10-task subset Ox Alpha (80%) beat Fable 5 (65%), which tells you how much a 10-task sample can mislead, because on every larger, audited measurement Fable 5 is in a different tier. The honest framing: against Claude, Ox Alpha traded verified frontier coding for zero cost during a preview and, at the time, an unknown maker (since resolved to Z.ai's GLM-5.3 Flash). The price gap is enormous (Fable 5 is $10 / $50 per 1M, Opus 4.8 is $5 / $25, and even the revealed GLM-5.3 Flash is only $0.15 / $0.50 at Z.ai's list price), and for throwaway experimentation that gap can justify the cheap model. For anything you would put your name on, it cannot, because the one thing Claude sells that GLM-5.3 Flash does not (verified frontier capability) is precisely the thing that fails silently until it matters.
5. Ox Alpha vs GPT-5.6 Sol
The GPT-5.6 Sol comparison is where Ox Alpha's single best marketing line comes from, so it is worth dismantling carefully. On the 10-task DeepSWE subset, Ox Alpha scored 80% against GPT-5.6 Sol's 52%, winning 7 of the 10 tasks, and that 28-point gap is the source of most "beats GPT" headlines - oxalpha.com. It is also the single least representative number in the entire comparison, because the same tester who produced it warned that 10 tasks carry high variance and should be treated as directional, and because the GPT baseline was measured on a larger run, so the two sides were never strictly like-for-like even within that one test.
Run the full suite and the story collapses toward the mean. Ox Alpha's full DeepSWE result is described as "more or less on par with GPT-5.6 Sol mid," and on the official SWE-bench Verified framework GPT-5.6 Sol scores about 82.2% - a different, non-comparable exam, but a reproducible one - oxalpha.com. The pattern is identical to the Claude case: Ox Alpha's win exists only inside a tiny hand-picked slice, and every broader, more auditable measurement pulls it back to mid-tier. A buyer who picks Ox Alpha over GPT-5.6 Sol on the strength of the 80% number is optimizing for a sample size of ten.
Where Ox Alpha does hold a real, defensible edge over GPT-5.6 Sol is cost and context economics, not capability. GPT-5.6 Sol is $4 / $20 per 1M with a 1.05M context; Ox Alpha ran free during the preview and its revealed identity (GLM-5.3 Flash) is still just $0.15 / $0.50 at Z.ai's list price with a 1M context, so for high-volume, long-context work the running cost difference is not marginal: roughly 27 times cheaper on input and 40 times cheaper on output. That is a genuine reason to reach for the cheap model for a bulk, low-stakes job. It is not a reason to believe it is a better coder, and conflating the two is exactly the mistake the viral subset invites. If you want the full picture of what Ox Alpha actually is beneath the benchmark noise, what is Ox Alpha lays out the specs and the story without the hype.
6. Ox Alpha vs Grok 4.6 and Gemini 3.7 Flash: the speed axis
These two comparisons matter because they isolate the axis Ox Alpha is quietly worst on, which almost no benchmark post mentions: interactive speed. A coding agent is a latency product. You are waiting on it, turn after turn, and a model that is 60% as fast feels far worse than 60% as smart, because the slowness compounds across every tool call in a long agentic loop. Ox Alpha's measured throughput is roughly 25 to 50 tok/s, dropping to 20-30 under load, with an ~11.6s median multi-step agent turn - ox-alpha.net. That is the number that turns a strong model into a frustrating one in a live session.
Grok 4.6 and Gemini 3.7 Flash both beat it, in different ways. On raw coding they are close to Ox Alpha's full-run figure: Grok 4.6 at 65.9% and Gemini 3.7 Flash at 65.3% on DeepSWE - daily.dev, both of which actually sit slightly above Ox Alpha's ~63%, undercutting any "next-gen" framing. But the separation is speed. Grok 4.6 runs at roughly 67-86 tok/s and Gemini 3.7 Flash at a startling ~340 tok/s, versus Ox Alpha's ~50 - Speedway Media. A model that is both slightly smarter and roughly seven times faster is not a close call for interactive work.
The trade-offs sort cleanly once speed is on the table.
- Gemini 3.7 Flash wins the practical trifecta: comparable coding, 1M context, and by far the best throughput, at $0.75 / $3.75 per 1M.
- Grok 4.6 is fast and named but carries a smaller 500K context and a high price (about 14.9x DeepSeek's cost), which narrows its niche.
- Ox Alpha won on price during the preview (and, revealed as GLM-5.3 Flash, remains very cheap), but pays for it in latency either way.
The lesson for an interactive coding workflow is that Ox Alpha's slowness is not a footnote, it is a first-order cost that the cheap price partly disguises. When you are iterating live in an editor, Gemini 3.7 Flash's throughput will save more of your day than a near-zero token price saves your budget, unless your budget is the only constraint that exists. This is the clearest case in the whole field where the cheapest model is genuinely more expensive, because it is spending your time instead of your money.
7. Ox Alpha vs DeepSeek V4 Pro and Qwen 3.8 Max: the value axis
The final cluster is the one that should most trouble the "Ox Alpha is unbeatable value" thesis, because it contains the models that already occupy the cheap-and-capable corner Ox Alpha claims. On the full DeepSWE run, Ox Alpha (~63%) is described as tied with DeepSeek V4 Pro - OpenRouter, and a casual reader stops there and concludes rough equality. The reproducible benchmarks say otherwise, emphatically. On LiveCodeBench v6, DeepSeek V4 Pro scores 93.5% and even the smaller V4-Flash Non-Think scores 55.2%, against Ox Alpha's 28% - GitHub. A 65-point gap on an independent, reproducible suite is not a tie by any honest reading.
DeepSeek also wins the structural argument that Ox Alpha could not answer during the preview. It ships under an MIT open-weights license, so you can download it, self-host it, audit it, and run it on private data with no third-party operator retaining your prompts. Its price is ~$0.66 / $1.98 per 1M, permanently cheap. During the stealth window Ox Alpha's one clear lead was raw sticker price ($0 versus cheap); with the preview closed and the model revealed as GLM-5.3 Flash ($0.15 / $0.50 at Z.ai's list price), that lead narrows from free-versus-cheap to cheap-versus-cheap, at about a quarter of DeepSeek's price, and since Z.ai published GLM-5.3 Flash's weights under MIT too, the open-weight gap has closed as well. User preference is split: a weather-app test judged Ox Alpha "more accurate and less buggy" than DeepSeek V4 Flash, while others found DeepSeek better for writing. That is the texture of two mid-tier models, not a challenger beating an incumbent.
Qwen 3.8 Max rounds out the value tier, though the public data on it is thinner. On Kingbench it scored 81.25%, behind Ox Alpha's 87.5%, and in fingerprint tests its video-token behaviour was "distinctly different" from Ox Alpha, which is one of the cleaner pieces of evidence that they are not the same model - Studio Global AI. Qwen is a named Alibaba flagship with a real maker behind it, which is why it outranked Ox Alpha on the stealth-week scorecard despite a lower Kingbench number: the untracked axes (a maker you could name, terms you could read) broke the tie in favour of accountability. The value tier's verdict is blunt: the cheap-capable corner is already occupied by models you can audit, and Ox Alpha's one distinguishing feature there was a temporary zero that its neighbours nearly match. With the reveal, that too is gone: it is now GLM-5.3 Flash, a named, cheap-but-not-free member of the same crowded value tier, distinguished by nothing its neighbours lack.
8. The DeepSWE field, charted (with the caveat)
Numbers scattered across seven sections are hard to hold in your head, so here they are on one axis, with the caveat printed directly in the caption because it is the whole point. The chart below places Ox Alpha's community full-run DeepSWE score next to each peer's own reported DeepSWE figure. These are not all from one leaderboard: Ox Alpha's 63% is a community full run and the peers are each maker's reported DeepSWE (v1.1) number, so this is orientation, not an official ranking. Even read generously, it shows Ox Alpha at the bottom of a tight band, not the top.
The chart does something the prose cannot: it shows how narrow the band is. The entire spread from GPT-5.6 Sol's 72.7 to Ox Alpha's 63 is under ten points, which means that on this particular exam the whole field is roughly comparable and Ox Alpha is a competent member of it - glm5.app. That is a fair, even flattering, way to read Ox Alpha, and it is worlds away from "tops the coding benchmarks." A ten-point band with your subject at the low end is the visual definition of mid-tier, not frontier. Set this against the LiveCodeBench chart on our benchmarks page, where Ox Alpha's 28% sits far below the same peers, and you have the full honest range: competitive on one soft suite, well behind on a hard reproducible one.
The reason both readings can be true is the thing every serious comparison eventually lands on: Ox Alpha has no authoritative aggregate score at all. It has no entry on Artificial Analysis and none on LMArena, and its own public listing shows no capability benchmark - Local AI Zone. Every peer in this chart has a tracked, contested, third-party-audited position in the market. Ox Alpha has a scatter of community runs and a #392-of-394 coding rank on one composite - LM Market Cap. When you cannot place a model on the shared board, you are not comparing capability, you are comparing anecdotes, and that uncertainty is itself a reason it scores low on trust.
9. Which should you pick?
All nine models can be sorted by a handful of real questions, and the honest answer for most production work is not Ox Alpha. The decision is not "which is smartest" but "which failure can you least afford": paying too much, waiting too long, leaking data to a party you cannot name, or building on something that vanishes without warning. During the stealth week those last two disqualified Ox Alpha for anything load-bearing, which is why it sits at the "throwaway experiment only" leaf of the tree below rather than the trunk. The reveal answers one of them (the operator is now named: Z.ai) and closes the other (the preview is over, the model is GLM-5.3 Flash), so the tree reads as the stealth-era decision, with the trust gate now relaxed for the named model. The full reasoning behind wiring in a fallback and sanitizing data is in how to use Ox Alpha.
The tree encodes the guide's core finding as a set of gates. During the stealth week, if your data was sensitive or the work was production, Ox Alpha was off the table before capability even entered, because the operator was anonymous and retained every prompt, and you could not contract with or audit a party you could not name - The Next Web. The reveal (Z.ai's GLM-5.3 Flash) removes the "cannot name the operator" objection, so the choice becomes a normal engineering trade-off: Fable 5 or Opus 4.8 for maximum verified capability, DeepSeek V4 Pro for cheap-and-auditable-and-open-weight, Gemini 3.7 Flash for fast-and-balanced. GLM-5.3 Flash survives to the bottom-right leaf, the cost-first job where cheap tokens outweigh the mid-tier and latency costs, and even there the guidance is to keep a fallback wired and feed any stealth or third-party endpoint only sanitized data.
This is also where an evaluation background pays off. Yuma Heymans (@yumahey), who runs Ox Alpha and co-founded the AI recruitment platform HeroHunt.ai, has spent years sourcing and vetting candidates with models like these, and the discipline is the same: you never hire on one glowing reference, and you never deploy a model on one viral benchmark. The scorecard above is that vetting discipline applied to nine candidates, and Ox Alpha was the brilliant intern who showed up with no references (until the background check came back and named the employer: Z.ai's GLM-5.3 Flash). Checking references is exactly what closed this case.
10. The bottom line
Now that the mystery is solved, the comparison resolves into one clean structural truth: Ox Alpha competed on the two axes that are cheapest to supply and lost on the three that are hardest. Price and raw context window are commodity features in 2026, matched or nearly matched across the field, and Ox Alpha maxed them by being free (during the preview) and 1M-token. Verified capability, real speed, and vendor trust are the expensive, durable features, and those are exactly where a named lab with a benchmark, an SLA, and a data policy pulls ahead. A scorecard weighted for what a buyer actually needs will always favour the durable features, which is why the stealth model finished last of nine despite topping the price column, and why knowing it is Z.ai's GLM-5.3 Flash does not move it up the coding, speed, or trust-of-capability columns: those numbers were always the model, not the mask.
That does not make Ox Alpha uninteresting, it makes it correctly placed. It was the most compelling free experiment of the moment, a genuinely capable mid-tier coder worth an afternoon inside an agent harness with throwaway data while the preview lasted - Hacker News, and its real identity (GLM-5.3 Flash, $0.15 / $0.50 per 1M at Z.ai's list price) is a cheap, fast, mid-tier model that fits that profile exactly. It is not a frontier model by any reproducible measure, it is not something to build a critical product on without a fallback, and it is not cheaper than the alternatives once you count latency and risk instead of just the sticker. Every one of those claims is sourced above. The best part of the story is that the community's forensics called the maker correctly, and the reveal proved them right. For the fastest gut-check, see where to run GLM-5.3 Flash now (this site's own free chat was retired on September 30, 2026); for the receipts, the spec sheet, the benchmarks, the API details, and the FAQ carry every figure with its citation, and Ox Alpha, honestly is the companion overview to this comparison.
When you have picked your model and want to build with it, start at founden.ai/app/build (Founden and Ox Alpha have the same founder).
This guide reflects the picture as of late August 2026, updated when the stealth preview closed and OpenRouter (the host that ran it) retired the "stealth/ox-alpha" slug and identified the model as Z.ai's GLM-5.3 Flash. Benchmarks here remain community-sourced and mixed, and the model's price and limits can still change, so verify every live figure before you rely on it.