Benchmarks
Ox Alpha's reputation was built on one number and complicated by the next. Here is every score we could find, each with its source and a confidence level, so you can judge the model rather than the headline.
The viral number
The 80% that launched a thousand headlines came from a hand-picked 10-task slice of DeepSWE. When the same tester ran the full 113-task set, the score came down to roughly 63%, mid-tier and near GPT-5.6 Sol and Claude Opus 4.8.
DeepSWE, pass@1 (Ox Alpha)
Reproducible coding
An independent, reproducible run on LiveCodeBench v6 (greedy, temperature 0, single attempt, zero generation failures) put Ox Alpha at 28% overall, well below frontier coders. It handled easy problems and fell off sharply on hard ones.
LiveCodeBench v6, pass@1 by difficulty
Community standings
On Kingbench, a non-standard community benchmark, Ox Alpha placed second at 87.5%, just behind GLM-5.3. That near-match is part of why so many people suspect the two share a lineage. There is no published, auditable methodology for this benchmark, so weigh it accordingly.
Kingbench score
In context
On DeepSWE, the one agentic coding suite where Ox Alpha has a number, it sits at the back of the current pack. And it has no SWE-bench Verified score at all, the benchmark where the leaders live: Claude Fable 5 at 95%, Opus 4.8 at 88.6%, GPT-5.6 Sol at 82.2%, DeepSeek V4 at 80.6%.
DeepSWE family (v1.1 and community runs; versions and testers differ, so read as directional).
DeepSWE, pass@1 (higher is better)
Every result
All 9 benchmarks we found, with score, rank, confidence, and source. Last verified August 24, 2026.
| Benchmark | Category | Score | Rank | Confidence | Source |
|---|---|---|---|---|---|
DeepSWE (10-task subset) Beat Claude Fable 5 (65%), GLM-5.3 (62%), Grok 4.6 (62%), and GPT-5.6 Sol (52%) on the same 10 tasks. The single most-cited number and the basis for 'topped coding benchmarks' headlines. Tester warned '10 tasks can have high variance' and to treat it as directional; baselines were measured on larger runs, so the comparison is not strictly like-for-like. | coding | ~80% pass@1 (8/10) | 1st on this subset | Reported | oxalpha.com (Ben Davis run) |
DeepSWE (full 113-task set) The same tester's full run corrected the headline 80% down to ~63%, described as on par with GPT-5.6 Sol mid-tier and near Claude Opus 4.8 (~59%), tied with DeepSeek V4 Pro, and behind Grok 4.6 and Gemini 3.7 Flash. The exact full-run figure is reported inconsistently (58 vs 63). This is the more representative result. | coding | ~63% pass@1 (also cited 58% / 58.4%) | mid-tier | Reported | The Cherry Creek News / OrcaRouter |
LiveCodeBench release_v6 Independent reproducible run (greedy, temp=0, single attempt), zero generation failures. By difficulty: Easy 51.2%, Medium 30.8%, Hard 13.8%. 27 points below even DeepSeek V4-Flash Non-Think (55.2%); DeepSeek V4-Pro scored 93.5%. Directly contradicts the 'frontier coder' framing. | coding | 28.0% pass@1 (49/175) | well below frontier | Reported | xnasarx/ox-alpha-benchmarks (GitHub) |
Kingbench Placed 2nd behind GLM-5.3 (91.25%), ahead of Qwen 3.8 Max (81.25%) and Claude Opus 4.8 (80%). Non-standard benchmark with no published, auditable methodology. The near-match with GLM-5.3 is cited as circumstantial evidence the two are the same model family. | coding | 87.5% | 2nd | Reported | daily.dev (AICodeKing) |
aicodingdaily OpenCode coding leaderboard Mid-pack. Sits between Gemini-3.1-Pro (10.1 pts, #25) and DeepSeek-V4-Flash (8.2 pts, #27), avg 12:36 per prompt. Not a frontier result. | coding | 8.9 / 20 points | #26 | Reported | AI Coding Daily |
LM Market Cap composite Coding, Quality, and Adoption all ranked #392 of 394. Benchmark signal 17/100 drags the composite down, though Pricing (100), Recency (100), and Context Window (96) score high. Reflects lack of verified capability data more than a measured capability floor. | general | 20 / 100 | Coding #392 of 394 | Reported | LM Market Cap |
OpenCode adoption leaderboard Its strongest 'leaderboard' placement, driven by being free. Usage by Aug 23: ~12-16 trillion tokens, ~180,000-221,000 unique users, ~3.56-5M+ sessions. This is an adoption/usage ranking, distinct from capability benchmarks. | agentic | No. 2 within 3 days | #2 | Reported | The Cherry Creek News |
benchable.ai domain scores TREAT WITH HEAVY SKEPTICISM. These near-perfect numbers contradict LiveCodeBench (28%), the LM Market Cap #392 coding rank, and the #26 OpenCode rank; the page's own results table showed 'No results match the current filters.' Likely auto-generated and unreliable. | general | 95-100% (Math 97%, Reasoning 96%, Coding 95%, perfect Hallucinations/Ethics/General Knowledge) | outlier | Speculative | benchable.ai |
Artificial Analysis Intelligence Index / LMArena As of Aug 22, 2026 Ox Alpha had NO entry on Artificial Analysis or LMArena/Chatbot Arena, and its own public listing shows no intelligence/coding/agentic benchmark. There is no authoritative aggregator score. For reference, the AA Intelligence Index (Aug 2026) was led by Claude Opus 5 at 63.0%. | general | Not listed (no entry) | unranked | Confirmed | Local AI Zone |
Head-to-head against GLM-5.3, Opus 4.8, GPT-5.6 Sol, and the rest.