26 systems · 28 studies

Every model we measured.

Models, coding CLIs, routers and harnesses, each with every value the studies recorded for it: the configuration it ran under, the sample size and the interval.

All models, CLIs, routers and harnesses

26 systems

Where each system was measured

One row per system, one column per study. A darker cell holds more values: it counts evidence, not quality.

System
Claude Sonnet 5.5Model
Claude Opus 5.5Model
Claude Haiku 4.5Model
Claude Fable 5.1Model
GPT-6.1 Sol (Codex CLI)Model
Claude CodeCoding CLI
Codex CLICoding CLI
Jev 1.13Router
AgentAgent harness
GPT-6.1 Sol (OpenAI API)Model
GPT-6 Luna (Codex CLI)Model
GPT-6 Luna (OpenAI API)Model
GPT 5.2Model
Gemini 3 FlashModel
GLM 5Model
Claude Sonnet 4.5Model
Claude Opus 4.5Model
Claude Opus 4.6Model
DeepSeek V3.2Model
MiniMax M2.5Model
Kimi K2.5Model
GPT 5 miniModel
Claude Opus 5Model
Claude Fable 5Model
GPT 5.5Model
Deterministic routing policyRouter

1–2, 3–5, 6–11, 12–23, 24+ valuesHatched: list-price calculations only131 system-study pairs with values

26 systems across 28 studies; 131 system-study pairs hold at least one value. Open a cell to see the values.

There is no cross-study plot of models: different studies use different tasks, so their values do not share an axis. Rank within one metric on the leaderboard.

Models

  • Model · Anthropic

    Claude Sonnet 5.5

    Anthropic’s mid-tier Claude model. Agent’s default coding model; measured here through Claude Code at low, medium, high and default effort, and as an LLM router.

    Best-supported quality value

    97% (189/194)

    Per-question accuracy · n = 194 · 95% CI 94%–99%

    236 values20 studies8 comparisons

  • Model · Anthropic

    Claude Opus 5.5

    Anthropic’s large Claude model, measured through Claude Code at low, medium, high and default effort.

    Best-supported quality value

    100% (24/24)

    Pass rate on eight hard tasks (Strict pass) · n = 24 · 95% CI 86%–100%

    136 values11 studies4 comparisons

  • Model · Anthropic

    Claude Haiku 4.5

    Anthropic’s small, low-price Claude model, measured through Claude Code, as an LLM router and in the public SWE-bench panel.

    Best-supported quality value

    94% (183/194)

    Per-question accuracy · n = 194 · 95% CI 90%–97%

    205 values17 studies18 comparisons

  • Model · Anthropic

    Claude Fable 5.1

    Anthropic’s highest-priced Claude model in these studies, measured through Claude Code.

    Best-supported quality value

    100% (24/24)

    Pass rate on eight hard tasks (Strict pass) · n = 24 · 95% CI 86%–100%

    29 values6 studies4 comparisons

  • Model · OpenAI

    GPT-6.1 Sol (Codex CLI)

    OpenAI’s GPT-6.1 Sol model run through the Codex CLI at low, medium and high effort.

    Best-supported quality value

    100% (16/16)

    Pass rate on eight hard tasks (Strict pass) · n = 16 · 95% CI 81%–100%

    128 values10 studies7 comparisons

  • Model · OpenAI

    GPT-6.1 Sol (OpenAI API)

    OpenAI’s GPT-6.1 Sol model called directly through the OpenAI API, without a CLI.

    13 values1 study4 comparisons

  • Model · OpenAI

    GPT-6 Luna (Codex CLI)

    OpenAI’s GPT-6 Luna model run through the Codex CLI.

    Best-supported quality value

    63% (10/16)

    Strict pass rate: single call vs agent loop on eight hard tasks · n = 16 · 95% CI 39%–82%

    34 values3 studies5 comparisons

  • Model · OpenAI

    GPT-6 Luna (OpenAI API)

    OpenAI’s GPT-6 Luna model called directly through the OpenAI API, without a CLI.

    5 values1 study3 comparisons

  • Model · OpenAI

    GPT 5.2

    OpenAI model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).

    Best-supported quality value

    85% (28/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 69%–93%

    3 values2 studies11 comparisons

  • Model · Google

    Gemini 3 Flash

    Google model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).

    Best-supported quality value

    82% (27/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 66%–91%

    3 values2 studies11 comparisons

  • Model · Z.ai

    GLM 5

    Z.ai model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).

    Best-supported quality value

    79% (26/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 62%–89%

    3 values2 studies11 comparisons

  • Model · Anthropic

    Claude Sonnet 4.5

    Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).

    Best-supported quality value

    76% (25/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 59%–87%

    3 values2 studies11 comparisons

  • Model · Anthropic

    Claude Opus 4.5

    Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).

    Best-supported quality value

    73% (24/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 56%–85%

    3 values2 studies11 comparisons

  • Model · Anthropic

    Claude Opus 4.6

    Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2).

    Best-supported quality value

    70% (23/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 53%–83%

    3 values2 studies11 comparisons

  • Model · DeepSeek

    DeepSeek V3.2

    DeepSeek model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).

    Best-supported quality value

    73% (24/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 56%–85%

    3 values2 studies11 comparisons

  • Model · MiniMax

    MiniMax M2.5

    MiniMax model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).

    Best-supported quality value

    70% (23/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 53%–83%

    3 values2 studies11 comparisons

  • Model · Moonshot AI

    Kimi K2.5

    Moonshot AI model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).

    Best-supported quality value

    70% (23/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 53%–83%

    3 values2 studies11 comparisons

  • Model · OpenAI

    GPT 5 mini

    OpenAI model in the public SWE-bench Verified panel (mini-SWE-agent v2).

    Best-supported quality value

    64% (21/33)

    Resolved rate on the same 33 SWE-bench Verified instances · n = 33 · 95% CI 47%–78%

    3 values2 studies11 comparisons

  • Model · Anthropic

    Claude Opus 5

    Anthropic model, measured as a blind code-review critic and in a list-price calculation.

    Best-supported quality value

    70% (28/40)

    Does the judge’s model family matter? · n = 40 · 95% CI 55%–82%

    2 values2 studies

  • Model · Anthropic

    Claude Fable 5

    Anthropic model, measured as a blind code-review critic.

    Best-supported quality value

    65% (26/40)

    Does the judge’s model family matter? · n = 40 · 95% CI 50%–78%

    1 values1 study

  • Model · OpenAI

    GPT 5.5

    OpenAI model, measured as a blind code-review critic on a small number of pairs.

    Best-supported quality value

    100% (2/2)

    Does the judge’s model family matter? · n = 2 · 95% CI 34%–100%

    1 values1 study

Coding CLIs

  • Coding CLI · Anthropic

    Claude Code

    Anthropic’s coding CLI. Each measurement pairs it with one Claude model; the context names the model.

    Best-supported quality value

    97% (189/194)

    Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy) · n = 194 · 95% CI 94%–99%

    412 values16 studies1 comparison

  • Coding CLI · OpenAI

    Codex CLI

    OpenAI’s coding CLI. Each measurement pairs it with one GPT model and effort; the context names them.

    Best-supported quality value

    100% (16/16)

    Pass rate on eight hard tasks (Strict pass) · n = 16 · 95% CI 81%–100%

    167 values12 studies1 comparison

Routers

  • Router · TypeSafe

    Jev 1.13

    A small routing model from TypeSafe that answers typed routing decisions over an API. Priced on input tokens only. Measured here as a router (accuracy, live time per call and cost per 1,000 decisions) and in the System One arena.

    Best-supported quality value

    77% (805/1048)

    Who decides right? Accuracy on 1,000+ checkable decisions · n = 1048 · 95% CI 74%–79%

    59 values5 studies3 comparisons

  • Router · Agent

    Deterministic routing policy

    Agent’s rule-based routing: an in-process policy picks the model and effort for each call from the task stage and signals. No model call, so no token cost.

    Best-supported quality value

    100% (20000/20000)

    Routing calls that returned a decision · n = 20000 · 95% CI 100%–100%

    7 values1 study3 comparisons

Agent harnesses

  • Agent harness · Agent

    Agent

    The Agent coding pipeline: onboarding, research, plan, act, verify and review, on Claude Sonnet 5.5 through a Claude subscription.

    Best-supported quality value

    69% (91/132)

    Single critic verdicts that preferred the AI change · n = 132 · 95% CI 61%–76%

    23 values4 studies11 comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.