Benchmarks
Ox Alpha's reputation was built on one number and complicated by the next. Here is every score we could find, each with its source and a confidence level, so you can judge the model rather than the headline.
The viral number
The 80% that launched a thousand headlines came from a hand-picked 10-task slice of DeepSWE. When the same tester ran the full 113-task set, the score came down to roughly 63%: mid-tier, near Claude Opus 4.8 (~59%) and DeepSeek V4 Pro, and behind Grok 4.6 and Gemini 3.7 Flash.
DeepSWE, pass@1 (Ox Alpha)
Reproducible coding
An independent, reproducible run on LiveCodeBench v6 (greedy, temperature 0, single attempt, zero generation failures) put Ox Alpha at 28% overall, well below frontier coders. It handled easy problems and fell off sharply on hard ones.
LiveCodeBench v6, pass@1 by difficulty
Community standings
On Kingbench, a non-standard community benchmark, Ox Alpha placed second at 87.5%, just behind GLM-5.3. That near-match was part of why so many people concluded the two share a lineage, now confirmed: Ox Alpha was GLM-5.3 Flash. There is no published, auditable methodology for this benchmark, so weigh it accordingly.
Kingbench score
In context
On DeepSWE, the one agentic coding suite where Ox Alpha has a number, its community full run sits at the back of this pack. The peers are each maker's own reported DeepSWE, run on different testers and versions, so read the gap as directional, not like-for-like.
It has no SWE-bench Verified score at all, the benchmark where the current leaders sit: Claude Fable 5 at 95%, Claude Opus 4.8 at 88.6% (the prior Anthropic flagship, now behind Opus 5), GPT-5.6 Sol at 82.2%, and DeepSeek V4 at 80.6%. Founden.ai, the AI builder that shares an owner with Ox Alpha, has a dedicated look at Claude Opus 4.8's own benchmarks and pricing.
DeepSWE pass@1: Ox Alpha (community full run) vs peers' reported DeepSWE
Ox Alpha figure: community full run. Peer figures are each maker's reported DeepSWE (v1.1); full sources on Compare.
Every result
All 9 benchmarks we found, with score, rank, confidence, and source. Last verified September 4, 2026.
| Benchmark | Category | Score | Rank | Confidence | Source |
|---|---|---|---|---|---|
DeepSWE (10-task subset) Beat Claude Fable 5 (65%), GLM-5.3 (62%), Grok 4.6 (62%), and GPT-5.6 Sol (52%) on the same 10 tasks. The single most-cited number and the basis for 'topped coding benchmarks' headlines. Tester warned '10 tasks can have high variance' and to treat it as directional; baselines were measured on larger runs, so the comparison is not strictly like-for-like. | coding | ~80% pass@1 (8/10) | 1st on this subset | Reported | oxalpha.com (Ben Davis run) |
DeepSWE (full 113-task set) The same tester's full run corrected the headline 80% down to ~63%, described as on par with GPT-5.6 Sol mid-tier and near Claude Opus 4.8 (~59%), tied with DeepSeek V4 Pro, and behind Grok 4.6 and Gemini 3.7 Flash. The exact full-run figure is reported inconsistently (58 vs 63). This is the more representative result. | coding | ~63% pass@1 (also cited 58% / 58.4%) | mid-tier | Reported | The Cherry Creek News / OrcaRouter |
LiveCodeBench release_v6 Independent reproducible run (greedy, temp=0, single attempt), zero generation failures. By difficulty: Easy 51.2%, Medium 30.8%, Hard 13.8%. 27 points below even DeepSeek V4-Flash Non-Think (55.2%); DeepSeek V4-Pro scored 93.5%. Directly contradicts the 'frontier coder' framing. | coding | 28.0% pass@1 (49/175) | well below frontier | Reported | xnasarx/ox-alpha-benchmarks (GitHub) |
Kingbench Placed 2nd behind GLM-5.3 (91.25%), ahead of Qwen 3.8 Max (81.25%) and Claude Opus 4.8 (80%). Non-standard benchmark with no published, auditable methodology. The near-match with GLM-5.3 is cited as circumstantial evidence the two are the same model family. | coding | 87.5% | 2nd | Reported | daily.dev (AICodeKing) |
aicodingdaily OpenCode coding leaderboard Mid-pack. Sits between Gemini-3.1-Pro (10.1 pts, #25) and DeepSeek-V4-Flash (8.2 pts, #27), avg 12:36 per prompt. Not a frontier result. | coding | 8.9 / 20 points | #26 | Reported | AI Coding Daily |
LM Market Cap composite Coding, Quality, and Adoption all ranked #392 of 394. Benchmark signal 17/100 drags the composite down, though Pricing (100), Recency (100), and Context Window (96) score high. Reflects lack of verified capability data more than a measured capability floor. | general | 20 / 100 | Coding #392 of 394 | Reported | LM Market Cap |
OpenCode adoption leaderboard Its strongest 'leaderboard' placement, driven by being free. Usage by Aug 23: ~12-16 trillion tokens, ~180,000-221,000 unique users, ~3.56-5M+ sessions. This is an adoption/usage ranking, distinct from capability benchmarks. | agentic | No. 2 within 3 days | #2 | Reported | The Cherry Creek News |
benchable.ai domain scores TREAT WITH HEAVY SKEPTICISM. These near-perfect numbers contradict LiveCodeBench (28%), the LM Market Cap #392 coding rank, and the #26 OpenCode rank; the page's own results table showed 'No results match the current filters.' Likely auto-generated and unreliable. | general | 95-100% (Math 97%, Reasoning 96%, Coding 95%, perfect Hallucinations/Ethics/General Knowledge) | outlier | Speculative | benchable.ai |
Artificial Analysis Intelligence Index (v4.1.1) After the reveal and open-weight release, Artificial Analysis lists GLM-5.3 Flash at an Intelligence Index of 57 (v4.1.1), just below the current frontier band (60-63) and level with Claude Opus 4.8, at a fraction of the cost (about $0.045 per task). During the stealth preview it had no aggregator entry; that is no longer true. | general | 57 | mid-tier | Confirmed | Artificial Analysis |
Head-to-head against GLM-5.3, Opus 4.8, GPT-5.6 Sol, and the rest.
You have seen the data and the live model. When you want a working product rather than an API call, Founden builds it: a company, a product, a 3D world, whatever you have in mind, without wiring up any of the infrastructure yourself.