46 systems · 46 metrics

The leaderboard with no composite score.

Models, coding CLIs, routers, harnesses and inference providers, each with its best-supported values for quality, speed and cost. Every value keeps its study, sample size, interval or range and configuration.

Where models differ, with uncertainty

Which model should I use?

No composite score. Pick a metric. The systems sort on that one value. The shaded band holds the rows the data cannot separate from the first.

Claude Sonnet 5.5 + Claude CodeModel + Coding CLI · Claude Code · eight hard validated tasks
100% (24/24)n = 24
Claude Opus 5.5 + Claude CodeModel + Coding CLI · Claude Code · eight hard validated tasks
100% (24/24)n = 24
Claude Opus 5.5 + Claude CodeModel + Coding CLI · Claude Code · effort high · eight hard validated tasks
100% (24/24)n = 24
Claude Fable 5.1 + Claude CodeModel + Coding CLI · Claude Code · eight hard validated tasks
100% (24/24)n = 24
GPT-6.1 Sol (Codex CLI)Model + Coding CLI · Codex CLI · effort medium · eight hard validated tasks
100% (16/16)n = 16
GPT-6.1 Sol (Codex CLI)Model + Coding CLI · Codex CLI · effort high · eight hard validated tasks
100% (16/16)n = 16
Claude Haiku 4.5 + Claude CodeModel + Coding CLI · Claude Code · eight hard validated tasks
46% (11/24)n = 24

Whiskers and lines: 95% Wilson interval7 measured points; n under each value6 of 7 at 100%: this task set cannot separate them.

7 measured points of pass rate on eight hard tasks (strict pass) in Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks, sorted by value. First row: Claude Sonnet 5.5 + Claude Code, 100% (24/24) (n = 24).

One row is one measured point; a run that counts for both a model and its CLI is one row that names both. The band and the shaded column follow the comparison rules: overlapping 95% intervals or run ranges cannot be separated, and neither can separate run ranges with fewer than 5 runs per side. A comparison page names a winner only under those rules, on the same chart.

Why there is no overall score

Per entity, the best-supported metrics across studies: a 95% interval first, then a run range, then the larger sample. Each keeps its study, n, interval or range and configuration. There is no composite score: the studies differ in task, route, effort and sample size, so one weighted number would hide what each value measured.

See how sample size changes pass-rate intervals. The planner is a calculation on fixed rates; it does not predict a winner.

A weighted sum would need weights nobody measured. It would mix measured values with list-price calculations, and many values here are ties at a ceiling (every call passed). So you rank within one metric of one study, where the comparison rules already say when a gap is real. How to read AI benchmarks honestly.

Every system, with its evidence

Up to 3 values per category, one per study, best supported first: a 95% interval, then a run range, then the larger n. The cells under each name show its evidence depth: one cell per study.

46 systems

Best-supported values per system

Quality: pass rates, accuracy and scores. Speed: time per call or decision. Cost: US dollars per call, pass or decision.

Sorted by the number of studies that measured it, not by a score.

Rates on a 0–100% track; times and costs on the span of every value of the same metric in the same studyHollow marks and hatched cells: list-price calculationsWhiskers: 95% interval; thin lines: run range or median to p95 (not intervals)

46 systems. Each row lists the system's best-supported quality, speed and cost values with study, n and interval or range. The rows are sorted by the reader's choice, not by a score.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.