- Benchmarks
- Models
26 systems · 28 studies · updated October 7, 2026
Every model we measured.
Models, coding CLIs, routers and harnesses, each with every value the studies recorded for it: the configuration it ran under, the sample size and the interval.
All models, CLIs, routers and harnesses
26 systems
Where each system was measured
One row per system, one column per study. A darker cell holds more values: it counts evidence, not quality.
| System | SWE-bench | SWE-bench probe | Blind review | Head-to-head | Hard tasks | Coding agents | Effort ladder | Caching | Agent memory | Decision arena | Routing accuracy | Routing overhead | Provider index | Cost what-ifs | CLI vs API | Calibration | Agent loop | Haiku thinking | JSON schema | Session cache | Cache break-even | Routing holdout | Thinking bill | LLM speed anatomy | Haiku retry cost | Harder tasks | Voice budget | Routing at scale |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 | — | 3 (1 calc.) | — | 8 (2 calc.) | 6 (1 calc.) | 4 (1 calc.) | 16 (4 calc.) | 13 (2 calc.) | 40 (8 calc.) | — | 9 (1 calc.) | 9 (5 calc.) | — | 3 (3 calc.) | 3 | — | 27 (2 calc.) | 7 (1 calc.) | 16 (4 calc.) | — | 18 (18 calc.) | 12 (2 calc.) | 14 (14 calc.) | 10 (5 calc.) | 8 (8 calc.) | 10 (1 calc.) | — | — |
| Claude Opus 5.5 | — | 3 (1 calc.) | — | 24 (6 calc.) | 12 (2 calc.) | 4 (1 calc.) | 16 (4 calc.) | 4 (2 calc.) | — | — | — | — | — | 3 (3 calc.) | — | — | — | — | — | — | 30 (30 calc.) | — | 20 (20 calc.) | 10 (5 calc.) | — | 10 (1 calc.) | — | — |
| Claude Haiku 4.5 | 2 | — | — | 8 (2 calc.) | 6 (1 calc.) | — | — | 9 | 40 (8 calc.) | — | 9 (1 calc.) | 13 (5 calc.) | — | 4 (3 calc.) | — | — | 27 (2 calc.) | 23 (4 calc.) | 16 (4 calc.) | — | 3 (3 calc.) | 12 (2 calc.) | 6 (6 calc.) | 10 (5 calc.) | 8 (8 calc.) | 9 | — | — |
| Claude Fable 5.1 | — | — | — | 8 (2 calc.) | 6 (1 calc.) | — | — | — | — | — | — | — | — | 3 (3 calc.) | — | — | — | — | — | — | 3 (3 calc.) | — | 6 (6 calc.) | 3 (2 calc.) | — | — | — | — |
| GPT-6.1 Sol (Codex CLI) | — | — | — | 24 (6 calc.) | 12 (2 calc.) | 4 (1 calc.) | 12 (3 calc.) | 9 | — | — | — | — | — | — | 13 | — | — | — | 16 (4 calc.) | — | — | — | 18 (18 calc.) | 10 (5 calc.) | — | 10 (1 calc.) | — | — |
| Claude Code | — | — | — | 48 (12 calc.) | 30 (5 calc.) | 8 (2 calc.) | 32 (8 calc.) | 26 (4 calc.) | 1 | — | — | 5 | — | — | 3 | — | 54 (4 calc.) | 27 (3 calc.) | 32 (8 calc.) | — | — | 22 (2 calc.) | 46 (46 calc.) | 33 (17 calc.) | 16 (16 calc.) | 29 (2 calc.) | — | — |
| Codex CLI | — | — | — | 24 (6 calc.) | 12 (2 calc.) | 4 (1 calc.) | 12 (3 calc.) | 9 | — | — | — | 4 | — | — | 19 | — | 26 (2 calc.) | — | 16 (4 calc.) | — | — | — | 18 (18 calc.) | 13 (7 calc.) | — | 10 (1 calc.) | — | — |
| Jev 1.13 | — | — | — | — | — | — | — | — | — | 30 (1 calc.) | 8 (1 calc.) | 8 (5 calc.) | — | 1 (1 calc.) | — | — | — | — | — | — | — | 12 (2 calc.) | — | — | — | — | — | — |
| Agent | 13 (2 calc.) | — | 5 | — | — | — | — | — | — | — | — | — | — | 1 (1 calc.) | — | 4 (2 calc.) | — | — | — | — | — | — | — | — | — | — | — | — |
| GPT-6.1 Sol (OpenAI API) | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 13 | — | — | — | — | — | — | — | — | — | — | — | — | — |
| GPT-6 Luna (Codex CLI) | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 5 | — | 26 (2 calc.) | — | — | — | — | — | — | 3 (2 calc.) | — | — | — | — |
| GPT-6 Luna (OpenAI API) | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 5 | — | — | — | — | — | — | — | — | — | — | — | — | — |
| GPT 5.2 | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Gemini 3 Flash | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| GLM 5 | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Claude Sonnet 4.5 | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Claude Opus 4.5 | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Claude Opus 4.6 | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| DeepSeek V3.2 | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| MiniMax M2.5 | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Kimi K2.5 | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| GPT 5 mini | 2 | — | — | — | — | — | — | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Claude Opus 5 | — | — | 1 | — | — | — | — | — | — | — | — | — | — | 1 (1 calc.) | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Claude Fable 5 | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| GPT 5.5 | — | — | 1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Deterministic routing policy | — | — | — | — | — | — | — | — | — | — | — | 7 (4 calc.) | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
1–2, 3–5, 6–11, 12–23, 24+ valuesHatched: list-price calculations only131 system-study pairs with values
26 systems across 28 studies; 131 system-study pairs hold at least one value. Open a cell to see the values.
There is no cross-study plot of models: different studies use different tasks, so their values do not share an axis. Rank within one metric on the leaderboard.
Models
Model · Anthropic
Claude Sonnet 5.5
Anthropic’s mid-tier Claude model. Agent’s default coding model; measured here through Claude Code at low, medium, high and default effort, and as an LLM router.
Best-supported quality value
97% (189/194)
Per-question accuracy · n = 194 · 95% CI 94%–99%
Model · Anthropic
Claude Opus 5.5
Anthropic’s large Claude model, measured through Claude Code at low, medium, high and default effort.
Best-supported quality value
100% (24/24)
Pass rate on eight hard tasks (Strict pass) · n = 24 · 95% CI 86%–100%
Model · Anthropic
Claude Haiku 4.5
Anthropic’s small, low-price Claude model, measured through Claude Code, as an LLM router and in the public SWE-bench panel.
Best-supported quality value
94% (183/194)
Per-question accuracy · n = 194 · 95% CI 90%–97%
Model · Anthropic
Claude Fable 5.1
Anthropic’s highest-priced Claude model in these studies, measured through Claude Code.
Best-supported quality value
100% (24/24)
Pass rate on eight hard tasks (Strict pass) · n = 24 · 95% CI 86%–100%
Model · OpenAI
GPT-6.1 Sol (Codex CLI)
OpenAI’s GPT-6.1 Sol model run through the Codex CLI at low, medium and high effort.
Best-supported quality value
100% (16/16)
Pass rate on eight hard tasks (Strict pass) · n = 16 · 95% CI 81%–100%
Model · OpenAI
GPT-6.1 Sol (OpenAI API)
OpenAI’s GPT-6.1 Sol model called directly through the OpenAI API, without a CLI.
Model · OpenAI
GPT-6 Luna (Codex CLI)
OpenAI’s GPT-6 Luna model run through the Codex CLI.
Best-supported quality value
63% (10/16)
Strict pass rate: single call vs agent loop on eight hard tasks · n = 16 · 95% CI 39%–82%
Model · OpenAI
GPT-6 Luna (OpenAI API)
OpenAI’s GPT-6 Luna model called directly through the OpenAI API, without a CLI.
Model · OpenAI
GPT 5.2
OpenAI model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).
Best-supported quality value
85% (28/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 69%–93%
Model · Google
Gemini 3 Flash
Google model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).
Best-supported quality value
82% (27/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 66%–91%
Model · Z.ai
GLM 5
Z.ai model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).
Best-supported quality value
79% (26/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 62%–89%
Model · Anthropic
Claude Sonnet 4.5
Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).
Best-supported quality value
76% (25/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 59%–87%
Model · Anthropic
Claude Opus 4.5
Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).
Best-supported quality value
73% (24/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 56%–85%
Model · Anthropic
Claude Opus 4.6
Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2).
Best-supported quality value
70% (23/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 53%–83%
Model · DeepSeek
DeepSeek V3.2
DeepSeek model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).
Best-supported quality value
73% (24/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 56%–85%
Model · MiniMax
MiniMax M2.5
MiniMax model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).
Best-supported quality value
70% (23/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 53%–83%
Model · Moonshot AI
Kimi K2.5
Moonshot AI model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).
Best-supported quality value
70% (23/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 53%–83%
Model · OpenAI
GPT 5 mini
OpenAI model in the public SWE-bench Verified panel (mini-SWE-agent v2).
Best-supported quality value
64% (21/33)
Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 47%–78%
Model · Anthropic
Claude Opus 5
Anthropic model, measured as a blind code-review critic and in a list-price calculation.
Best-supported quality value
70% (28/40)
Does the judge’s model family matter? · n = 40 · 95% CI 55%–82%
Model · Anthropic
Claude Fable 5
Anthropic model, measured as a blind code-review critic.
Best-supported quality value
65% (26/40)
Does the judge’s model family matter? · n = 40 · 95% CI 50%–78%
Model · OpenAI
GPT 5.5
OpenAI model, measured as a blind code-review critic on a small number of pairs.
Best-supported quality value
100% (2/2)
Does the judge’s model family matter? · n = 2 · 95% CI 34%–100%
Coding CLIs
Coding CLI · Anthropic
Claude Code
Anthropic’s coding CLI. Each measurement pairs it with one Claude model; the context names the model.
Best-supported quality value
97% (189/194)
Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy) · n = 194 · 95% CI 94%–99%
Coding CLI · OpenAI
Codex CLI
OpenAI’s coding CLI. Each measurement pairs it with one GPT model and effort; the context names them.
Best-supported quality value
100% (16/16)
Pass rate on eight hard tasks (Strict pass) · n = 16 · 95% CI 81%–100%
Routers
Router · TypeSafe
Jev 1.13
A small routing model from TypeSafe that answers typed routing decisions over an API. Priced on input tokens only. Measured here as a router (accuracy, live time per call and cost per 1,000 decisions) and in the System One arena.
Best-supported quality value
77% (805/1048)
Who decides right? Accuracy on 1,000+ checkable decisions · n = 1048 · 95% CI 74%–79%
Router · Agent
Deterministic routing policy
Agent’s rule-based routing: an in-process policy picks the model and effort for each call from the task stage and signals. No model call, so no token cost.
Best-supported quality value
100% (20000/20000)
Routing calls that returned a decision · n = 20000 · 95% CI 100%–100%
Agent harnesses
Agent harness · Agent
Agent
The Agent coding pipeline: onboarding, research, plan, act, verify and review, on Claude Sonnet 5.5 through a Claude subscription.
Best-supported quality value
69% (91/132)
Single critic verdicts that preferred the AI change · n = 132 · 95% CI 61%–76%