Compare
A single, sourced master table. Ox Alpha next to nine models it gets measured against, on the axes that actually decide a build: context, price, coding ability, speed, and whether you can run the weights yourself.
| Model | Released | Context | Input /1M | Output /1M | Coding benchmark | Speed | Reasoning | Open weights |
|---|---|---|---|---|---|---|---|---|
Ox Alpha Undisclosed (fingerprints to Zhipu GLM) | Aug 20, 2026 | 1M | $0 | $0 | DeepSWE full 113-task ~63% (community) | ~25-50 tok/s | Yes | No |
GLM-5.3 Z.ai (Zhipu AI) | Aug 14, 2026 | 1M | $1.40 | $4.40 | DeepSWE v1.1 66.9% (vendor-reported) | ~102 tok/s | Yes | No |
Claude Opus 4.8 Anthropic | May 28, 2026 | 1M | $5.00 | $25.00 | SWE-bench Verified 88.6% | ~61 tok/s | Yes | No |
GPT-5.6 Sol OpenAI | Jul 9, 2026 | 1.05M | $4.00 | $20.00 | SWE-bench Verified 82.2% | ~73 tok/s | Yes | No |
Claude Fable 5 Anthropic | June 9, 2026 | 1M | $10.00 | $50.00 | SWE-bench Verified 95% | ~50-75 tok/s | Yes | No |
Qwen 3.8 Max Alibaba (Qwen team / Alibaba Cloud) | Aug 3, 2026 | 1M | $2.00 | $6.00 | SWE-bench Pro 67.7% (vendor-reported) | ~68.9 tok/s | Yes | Yes |
Grok 4.6 xAI | Aug 12, 2026 | 500K | $2.00 | $6.00 | Terminal-Bench v2.1 88.4% | ~62 tok/s | Yes | No |
Gemini 3.7 Flash Google (Google DeepMind) | Aug 13, 2026 | 1M | $0.75 | $3.75 | DeepSWE v1.1 65.3% | ~348 tok/s | Yes | No |
DeepSeek V4 Pro DeepSeek | V4-Pro-0813, Aug 13, 2026 | 1M | $0.66 | $1.98 | SWE-bench Verified 80.6% (vendor-reported) | ~70 tok/s | Yes | Yes |
Context and price are shortened for scanning (full figures and exact terms are in the notes). Ox Alpha is highlighted.
What the table says
$0 during the preview against $1.40 to $10 per 1M input elsewhere, with a full 1M-token window matched only by the largest models here. On those two axes nothing beats it right now.
Its community DeepSWE run (~63%) sits near GLM-5.3 and Gemini 3.7 Flash, but well behind the SWE-bench Verified leaders (Fable 5 at 95%, Opus 4.8 at 88.6%, GPT-5.6 Sol at 82.2%).
At ~25 to 50 tok/s it lags Gemini 3.7 Flash (~348) and GLM-5.3 (~102), and unlike the open-weight models here (Qwen, DeepSeek) you cannot run or audit it. Great for experiments, not for production.
Notes and sources
The detail behind each row: what is vendor-reported vs independent, pricing nuances, and where the numbers come from.
Cloaked/undisclosed model on OpenRouter; behavioral fingerprints point to a Zhipu GLM model. Free during preview. Multimodal input (text, image, video), 1M context, 131K max output, adjustable reasoning effort. Benchmark figures are community-run and independent, not vendor-published, so treat as reported.
Released Aug 14, 2026 (Aug 13 PT). Same base model as GLM-5.2; gains from extended post-training. Pricing ($1.40 in / $4.40 out, $0.26 cached read per 1M) and 1M context confirmed across OpenRouter and Artificial Analysis. Coding/agent benchmarks are ALL vendor-reported by Z.ai and not independently re-run: Terminal-Bench 3.0 28.3, DeepSWE v1.1 66.9, SWE-Marathon v1.1 42.5, Toolathlon Verified 73.0. No standard SWE-bench Verified or LiveCodeBench figure published. Standout result: CyberGym 84.5% (found 2,436 vulns across 269 projects). Marketed as strongest open-weights coding model, but weights NOT yet released as of 2026-08-24; currently API/Coding-Plan only. Output speed varies by provider.
Model ID claude-opus-4-8 (pinned snapshot). Now a legacy model, superseded by Claude Opus 5 (Jun 9, 2026). Knowledge and training cutoff Jan 2026. Fast mode at 2.5x speed for $10 in / $50 out per 1M. Core specs (context, max output, price, modalities, reasoning, release) CONFIRMED from Anthropic official docs. SWE-bench Verified 88.6% and SWE-bench Pro ~69% are REPORTED (appear only in the announcement image/System Card; figure taken from LLM Stats and Vellum). Online-Mind2Web 84% is an official Anthropic figure. AA Intelligence Index 57.
Flagship of OpenAI's July 2026 GPT-5.6 family (Sol > Terra > Luna). GA Jul 9, 2026 (preview Jun 26); knowledge cutoff Feb 16, 2026. Context 1.05M total = up to 922K input + 128K output. PRICING CONFLICT: OpenAI docs and OpenRouter first-party give $4 in / $20 out per 1M (reported here); Artificial Analysis and some aggregators list $5/$30; OpenRouter also showed a promo $2/$10. Long-context surcharge: 2x input and 1.5x output when input exceeds 272K tokens. Other reported benchmarks: AA Coding Agent Index 80, AA Intelligence Index 59, Terminal-Bench 2.1 88.8, DeepSWE 72.7, FrontierMath v2 89.1; GPQA Diamond also cited as high as 94.6% by some sources. Postdates assistant training cutoff; all figures from live web, treat as reported.
Core specs (pricing $10/$50 per 1M, June 9 2026 release, vision/PDF input) confirmed on Anthropic pages. Context (1M) and max output (128K) confirmed via OpenRouter; Anthropic pages did not state token limits directly. Prompt caching gives 90% input discount (cache read ~$1.00/M; write ~$12.50/M, ~$20/M for 1h). US-only inference 1.1x. Coding figures are third-party aggregated (BenchLM: SWE-bench Verified 95%, SWE-bench Pro ~80%; SWE-bench Pro 80.3% cited as Anthropic's preferred agentic metric vs Opus 4.8 at 69.2%). Anthropic does not publish exact SWE-bench percentages, so coding numbers are reported, not confirmed. Public Mythos-class tier; a gated Mythos 5 tier exists. #3 of 224 on BenchAlign (82.7/100), #2 in coding.
Architecture: 2.4T-param sparse MoE with ~95B active params; multimodal from the ground up. API released Aug 3, 2026 via Alibaba Cloud Model Studio; open weights followed Aug 8, 2026 on Hugging Face (Qwen/Qwen3.8-2.4T-A95B plus FP8), ungated, under a bespoke license. PRICING VARIES BY REGION: OpenRouter and Alibaba Cloud Singapore list $2.00 in / $6.00 out per 1M (headline); China/Germany/US/Japan/HK standard is $1.65 in / $4.951 out. Prompt caching up to ~88% discount. Vendor-reported benchmarks: Terminal-Bench 2.1 86.6, OSWorld-Verified 86.1, SWE-bench Pro 67.7. AA Intelligence Index 58 (#13 of 186); output 68.9 tok/s (slow, verbose), TTFT ~2.28s. Much secondary coverage appears AI-generated; core specs corroborated by official docs, OpenRouter and AA, but benchmark percentages are vendor-reported and not independently verified.
Released Aug 12, 2026 (35 days after Grok 4.5); knowledge cutoff Feb 1, 2026. Proprietary, hosted-only via xAI/SpaceXAI API. Standard price $2/$6 per 1M; blended ~$1.35/1M (AA). Long-context tier: over 200K tokens billed $4/M input, $1/M cached input, $12/M output. A fast low-latency variant runs at 2x price. AA Intelligence Index 61 (~#6). Coding scores are version-sensitive: AA reports Terminal-Bench v2.1 88.4%, while xAI's launch post lists harder Terminal-Bench v3.0 26%, DeepSWE v1.1 65.9%, APEX-SWE 56.4%, CursorBench v3.2 69.9%, FrontierCode v1.1 61.3%. CAUTION: a circulating 95.60% SWE-bench Verified figure is not corroborated by xAI or AA and is treated as unverified, NOT reported. Max output tokens not published by authoritative sources.
Released Aug 13, 2026 as Google's most intelligent workhorse model for coding and agents, three weeks after Gemini 3.6 Flash. PRICING NUANCE: $0.75/$3.75 per 1M is an introductory rate through Dec 31, 2026; standard rate rises to $1.50/$7.50 per 1M on Jan 1, 2027. Google does not publish SWE-bench Verified or LiveCodeBench; official card uses newer benchmarks. Best coding scores from the card: DeepSWE v1.1 65.3% (up from 49.0% on 3.6 Flash), Terminal-bench 2.1 85.8%, FrontierCode 1.1 43.6%. Other results: WebDev Arena 1588 Elo, GDM-MRCR v2 (128k) 97.0%, CharXiv 84.5%, OSWorld-2.0 47.9%. Reasoning via thinking_level (older thinking_budget deprecated). Output ~348 tok/s (AA), TTFT ~12.2s. Closed weights, API-only (Gemini API, AI Studio, Vertex, Antigravity).
Main columns use DeepSeek V4 Pro (GA 0813). V4 previewed Apr 24, 2026 (0423); GA in two waves: V4-Flash-0731 (Jul 31) and V4-Pro-0813 (Aug 13). New pricing took effect Aug 16, 2026 16:00 UTC. V4 Pro ~1.6T-param MoE; V4 Flash 284B MoE (~13B active). Both 1M context, 384K max output, text-only (separate DeepSeek Vision Exp handles images). Pricing (official, per 1M, cache-miss input): Pro input $0.66 off-peak / $1.32 peak, output $1.98 / $3.96, cache-hit input ~$0.022; Flash input $0.22 / $0.44, output $0.66 / $1.32. Peak hours 01:00-04:00 and 06:00-10:00 UTC Mon-Fri. AA displays Pro at peak rate. Speed: Pro reasoning/max ~69.8-71 tok/s, Pro non-reasoning ~32 tok/s, Flash-Max ~131 tok/s (fastest). Vendor benchmarks (not independently reproduced): SWE-bench Verified 80.6% (Pro-Max), LiveCodeBench 93.5 (Pro-Thinking), AA Intelligence Index 53. Both Pro and Flash open-weights under MIT.
Comparison compiled 9-model wide from 42 primary sources on Aug 24, 2026.