- Benchmarks
- Leaderboard
46 systems · 46 metrics · updated October 7, 2026
The leaderboard with no composite score.
Models, coding CLIs, routers, harnesses and inference providers, each with its best-supported values for quality, speed and cost. Every value keeps its study, sample size, interval or range and configuration.
Where models differ, with uncertainty
No composite score. Pick a metric. The systems sort on that one value. The shaded band holds the rows the data cannot separate from the first.
Pass rate on eight hard tasks (Strict pass)
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks · higher is better · sorted by the point value, not a ranking
| System | n | Interval or range | Configuration | Against the first row | |
|---|---|---|---|---|---|
| Claude Sonnet 5.5ModelClaude CodeCoding CLI | 100% (24/24) | 24 | 95% CI 86%–100% | Claude Code · eight hard validated tasks | first row |
| Claude Opus 5.5ModelClaude CodeCoding CLI | 100% (24/24) | 24 | 95% CI 86%–100% | Claude Code · eight hard validated tasks | overlaps: a tie · compare |
| Claude Opus 5.5ModelClaude CodeCoding CLI | 100% (24/24) | 24 | 95% CI 86%–100% | Claude Code · effort high · eight hard validated tasks | overlaps: a tie · compare |
| Claude Fable 5.1ModelClaude CodeCoding CLI | 100% (24/24) | 24 | 95% CI 86%–100% | Claude Code · eight hard validated tasks | overlaps: a tie · compare |
| GPT-6.1 Sol (Codex CLI)ModelCodex CLICoding CLI | 100% (16/16) | 16 | 95% CI 81%–100% | Codex CLI · effort medium · eight hard validated tasks | overlaps: a tie · compare |
| GPT-6.1 Sol (Codex CLI)ModelCodex CLICoding CLI | 100% (16/16) | 16 | 95% CI 81%–100% | Codex CLI · effort high · eight hard validated tasks | overlaps: a tie · compare |
| Claude Haiku 4.5ModelClaude CodeCoding CLI | 46% (11/24) | 24 | 95% CI 28%–65% | Claude Code · eight hard validated tasks | does not overlap · compare |
Whiskers and lines: 95% Wilson interval7 measured points; n under each value6 of 7 at 100%: this task set cannot separate them.
7 measured points of pass rate on eight hard tasks (strict pass) in Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks, sorted by value. First row: Claude Sonnet 5.5 + Claude Code, 100% (24/24) (n = 24).
One row is one measured point; a run that counts for both a model and its CLI is one row that names both. The band and the shaded column follow the comparison rules: overlapping 95% intervals or run ranges cannot be separated, and neither can separate run ranges with fewer than 5 runs per side. A comparison page names a winner only under those rules, on the same chart.
Why there is no overall score
Per entity, the best-supported metrics across studies: a 95% interval first, then a run range, then the larger sample. Each keeps its study, n, interval or range and configuration. There is no composite score: the studies differ in task, route, effort and sample size, so one weighted number would hide what each value measured.
See how sample size changes pass-rate intervals. The planner is a calculation on fixed rates; it does not predict a winner.
A weighted sum would need weights nobody measured. It would mix measured values with list-price calculations, and many values here are ties at a ceiling (every call passed). So you rank within one metric of one study, where the comparison rules already say when a gap is real. How to read AI benchmarks honestly.
Every system, with its evidence
Up to 3 values per category, one per study, best supported first: a 95% interval, then a run range, then the larger n. The cells under each name show its evidence depth: one cell per study.
46 systems
Best-supported values per system
Quality: pass rates, accuracy and scores. Speed: time per call or decision. Cost: US dollars per call, pass or decision.
Quality
97% (189/194)Per-question accuracy97% (189/194)Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)92% (115/125)Per-question accuracy on unseen decisionsQuality
97% (189/194)Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)92% (115/125)Per-question accuracy on unseen decisions100% (24/24)Pass rate on eight hard tasks (Strict pass)Quality
94% (183/194)Per-question accuracy91% (177/194)Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)82% (102/125)Per-question accuracy on unseen decisionsQuality
100% (16/16)Pass rate on eight hard tasks (Strict pass)100% (16/16)Strict pass rate by effort on eight hard tasks63% (10/16)Strict pass rate: single call vs agent loop on eight hard tasksQuality
100% (16/16)Pass rate on eight hard tasks (Strict pass)100% (16/16)Strict pass rate by effort on eight hard tasks69% (11/16)Pass rate on 4 harder tasks (Strict pass)Quality
100% (24/24)Pass rate on eight hard tasks (Strict pass)100% (16/16)Strict pass rate by effort on eight hard tasks100% (15/15)Pass rate on five validated tasksQuality
77% (805/1048)Who decides right? Accuracy on 1,000+ checkable decisions100% (246/246)Routing calls that returned a decision95% (184/194)Per-question accuracyQuality
100% (24/24)Pass rate on eight hard tasks (Strict pass)100% (15/15)Pass rate on five validated tasks64.2%Reasoning share of output tokens per call on hard tasks (calculation)Speed
Not measuredSpeed
Not measured
| Quality | Speed | Cost | ||||
|---|---|---|---|---|---|---|
| Claude Sonnet 5.5Anthropic | Model | 16 of 19 | 229 |
|
|
|
| Claude CodeAnthropic | Coding CLI | 14 of 16 | 412 |
|
|
|
| Claude Haiku 4.5Anthropic | Model | 14 of 16 | 198 |
|
|
|
| Codex CLIOpenAI | Coding CLI | 11 of 12 | 167 |
|
|
|
| GPT-6.1 Sol (Codex CLI)OpenAI | Model | 9 of 10 | 128 |
|
|
|
| Claude Opus 5.5Anthropic | Model | 8 of 10 | 127 |
|
|
|
| Jev 1.13TypeSafe | Router | 4 | 57 |
|
|
|
| AgentAgent | Agent harness | 3 of 4 | 23 |
|
|
|
| Claude Fable 5.1Anthropic | Model | 3 of 4 | 23 |
|
|
|
| GPT-6 Luna (Codex CLI)OpenAI | Model | 3 | 34 |
|
|
|
| Claude Opus 4.5Anthropic | Model | 2 | 3 |
| Not measured |
|
| Claude Opus 4.6Anthropic | Model | 2 | 3 |
| Not measured |
|
| Claude Sonnet 4.5Anthropic | Model | 2 | 3 |
| Not measured |
|
| DeepSeek V3.2DeepSeek | Model | 2 | 3 |
| Not measured |
|
| Gemini 3 FlashGoogle | Model | 2 | 3 |
| Not measured |
|
| GLM 5Z.ai | Model | 2 | 3 |
| Not measured |
|
| GPT 5 miniOpenAI | Model | 2 | 3 |
| Not measured |
|
| GPT 5.2OpenAI | Model | 2 | 3 |
| Not measured |
|
| Kimi K2.5Moonshot AI | Model | 2 | 3 |
| Not measured |
|
| MiniMax M2.5MiniMax | Model | 2 | 3 |
| Not measured |
|
| Amazon BedrockAmazon | Inference provider | 1 | 23 | Not measured | Not measured |
|
| AnthropicAnthropic | Inference provider | 1 | 35 | Not measured | Not measured |
|
| AzureMicrosoft | Inference provider | 1 | 33 | Not measured | Not measured |
|
| BasetenBaseten | Inference provider | 1 | 9 | Not measured | Not measured |
|
| CerebrasCerebras | Inference provider | 1 | 3 | Not measured | Not measured |
|
| Claude Fable 5Anthropic | Model | 1 | 1 |
| Not measured | Not measured |
| Claude Opus 5Anthropic | Model | 1 | 1 |
| Not measured | Not measured |
| Claude Platform on AWSAnthropic | Inference provider | 1 | 15 | Not measured | Not measured |
|
| Cloudflare Workers AICloudflare | Inference provider | 1 | 11 | Not measured | Not measured |
|
| DeepInfraDeepInfra | Inference provider | 1 | 16 | Not measured | Not measured |
|
| Deterministic routing policyAgent | Router | 1 | 7 |
|
|
|
| Fireworks AIFireworks AI | Inference provider | 1 | 6 | Not measured | Not measured |
|
| Google AI StudioGoogle | Inference provider | 1 | 16 | Not measured | Not measured |
|
| Google Vertex AIGoogle | Inference provider | 1 | 37 | Not measured | Not measured |
|
| GPT 5.5OpenAI | Model | 1 | 1 |
| Not measured | Not measured |
| GPT-6 Luna (OpenAI API)OpenAI | Model | 1 | 5 | Not measured |
| Not measured |
| GPT-6.1 Sol (OpenAI API)OpenAI | Model | 1 | 13 | Not measured |
| Not measured |
| GroqGroq | Inference provider | 1 | 6 | Not measured | Not measured |
|
| NebiusNebius | Inference provider | 1 | 4 | Not measured | Not measured |
|
| Novita AINovita AI | Inference provider | 1 | 15 | Not measured | Not measured |
|
| OpenAIOpenAI | Inference provider | 1 | 14 | Not measured | Not measured |
|
| OpenRouterOpenRouter | Inference provider | 1 | 20 | Not measured | Not measured |
|
| ParasailParasail | Inference provider | 1 | 20 | Not measured | Not measured |
|
| SambaNovaSambaNova | Inference provider | 1 | 4 | Not measured | Not measured |
|
| SiliconFlowSiliconFlow | Inference provider | 1 | 12 | Not measured | Not measured |
|
| Together AITogether AI | Inference provider | 1 | 10 | Not measured | Not measured |
|
Rates on a 0–100% track; times and costs on the span of every value of the same metric in the same studyHollow marks and hatched cells: list-price calculationsWhiskers: 95% interval; thin lines: run range or median to p95 (not intervals)
46 systems. Each row lists the system's best-supported quality, speed and cost values with study, n and interval or range. The rows are sorted by the reader's choice, not by a score.