What is the best AI model for coding? Our data says four models tie
4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.
TL;DR
- On our hard coding set, four models share the top tier: Claude Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol in the Codex CLI. Six configurations passed every call. Sonnet, Opus (default and high effort) and Fable passed 24/24 (95% interval 86% to 100%). GPT-6.1 Sol at medium and high effort passed 16/16 (81% to 100%).
- Claude Haiku 4.5 is the only model whose interval separates. It passed 11/24 strictly (46%, 95% interval 28% to 65%).
- The set is at its ceiling. A tie means our tasks cannot find a gap, not that the models are equal.
- On SWE-bench Verified, the panel is one tier too. On the same 33 instances, 11 public runs resolved 21 to 28, and Agent resolved 25 (59% to 87%). Every interval overlaps.
- Inside the top tier, choose on cost, speed and route. Sonnet had the lowest list-price cost per strict pass at $0.0143. Fable cost $0.0933, 6.5x as much (calculation). The speed ranges overlap.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks
139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.
Transcript
- Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
- 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
- Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
- Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
- Open benchmarks: intervals, sources and every failure kept.
The short answer
A list that ranks models one to ten can hide ties. Our rule: models that our data cannot separate share one tier. On eight hard tasks with strict validators, that gives two tiers: four models in Tier 1 and Haiku 4.5 in Tier 2. Study: /benchmarks/hard-model-head-to-head.
The tier list, built from intervals
Tier 1 holds the option with the highest rate and every option whose 95% interval overlaps it. The next tier starts at the first option whose interval sits fully below, and the rule repeats.
| Tier | Configuration | Strict passes | 95% interval | Cost per strict pass (calculation) | Median time per call (fastest to slowest) |
|---|---|---|---|---|---|
| 1 | Claude Sonnet 5.5 · Claude Code | 24/24 | 86% to 100% | $0.0143 | 7.7 s (2.3 to 34.8 s) |
| 1 | Claude Opus 5.5 · Claude Code | 24/24 | 86% to 100% | $0.0282 | 9.2 s (4.2 to 27.2 s) |
| 1 | Claude Opus 5.5 (high) · Claude Code | 24/24 | 86% to 100% | $0.0334 | 11.0 s (3.6 to 63.0 s) |
| 1 | GPT-6.1 Sol (medium) · Codex CLI | 16/16 | 81% to 100% | $0.0256 | 13.1 s (8.5 to 61.6 s) |
| 1 | Claude Fable 5.1 · Claude Code | 24/24 | 86% to 100% | $0.0933 | 16.1 s (4.5 to 90.0 s) |
| 1 | GPT-6.1 Sol (high) · Codex CLI | 16/16 | 81% to 100% | $0.0151 | 18.1 s (11.7 to 92.2 s) |
| 2 | Claude Haiku 4.5 · Claude Code | 11/24 | 28% to 65% | $0.0672 | 39.0 s (15.3 to 75.1 s) |
Haiku's interval stops at 65%; the lowest bound in Tier 1 is 81%. They do not overlap, so Tier 1 is ahead of Haiku on this set. Haiku's 13 non-passes were 5 format misses (right answer, wrong format) and 8 wrong answers.
Repeated prompts agree. Sonnet and GPT-6.1 Sol (medium) passed 10/10 on each of three prompts (72% to 100%). Haiku passed 0/10 on an exact-number prompt (0% to 28%) (study).
Effort does not change the tier. On the effort ladder, all 11 cells of Sonnet, Opus and GPT-6.1 Sol passed 16/16 (81% to 100% each).
SWE-bench Verified: one tier for twelve systems
On the same 33 instances, the 11 public mini-SWE-agent v2 runs resolved 21 to 28. Agent, our pipeline on Sonnet 5.5, resolved 25 (76%, 95% interval 59% to 87%).
The highest lower bound is 69% (GPT 5.2 at high effort, 28/33). The lowest upper bound is 78% (GPT 5 mini, 21/33). Because 69% is below 78%, every pair of intervals overlaps. All twelve systems share Tier 1.
Note: the panel is third-party, with older models in a bash-only harness. Agent is a full pipeline, not a bare model. Study: /benchmarks/swe-bench-verified.
How to choose inside Tier 1
1. Cost per pass (a calculation)
- Sonnet 5.5: $0.0143 per strict pass, the lowest.
- GPT-6.1 Sol (high): $0.0151, about 5.5% more.
- Opus 5.5: $0.0282, 1.97x Sonnet.
- Fable 5.1: $0.0933, 6.5x Sonnet.
Haiku lists at half of Sonnet's price, but it cost $0.0672 per strict pass, 4.7x Sonnet.
Only Sonnet 5.5 in Claude Code sits on the frontier. No other configuration passes as often for the same cost per pass or less.
The cheapest option depends on your workload. We repriced our agent's SWE-bench tokens (94.0% cache reads) at each list price. Per resolved instance, GPT-6.1 Sol $2.10, Sonnet $3.49, Opus $5.75, Fable $12.85 (calculation). GPT-6.1 Sol charges $0.10 per million cache-read tokens; Sonnet charges $0.20.
On the effort ladder, low effort was cheapest per strict pass for each model (calculation): Sonnet $0.0122, Opus $0.0212, GPT-6.1 Sol $0.0128.
2. Speed (the ranges overlap)
Tier 1 medians ran from 7.7 s (Sonnet) to 18.1 s (GPT-6.1 Sol, high), but the per-call ranges overlap. The medians describe one batch. They do not rank the models.
3. Route: CLI or API
On five short tasks, the Codex CLI sent a median 12,124 input tokens per call; Claude Code sent 2,130. On a scheduler repair, the same GPT-6.1 Sol took a median 61.2 s through the Codex CLI and 17.3 s through the OpenAI API. We ran each route 3 times, and the ranges do not overlap. More: the hidden context tax.
Compare pages: Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI) and Sonnet 5.5 vs Opus 5.5.
Why we do not crown one model
- Ceiling. 6 of 7 configurations passed every call. When nearly every option scores 100%, the set cannot rank them.
- Small samples. A 24/24 has a 95% interval of 86% to 100%; a 16/16 has 81% to 100%. A tie at these sizes can hide a gap of about 14 to 19 points (calculation).
- Route and model together. A Claude-vs-GPT row compares Claude Code and the Codex CLI too, not the models alone.
- The tier depends on the tasks. On five easier tasks, all nine configurations shared one tier: 127 of 130 calls passed (93% to 99%), and Haiku passed 15/15 (80% to 100%).
What would change the answer
- Harder tasks. We plan a harder set with strict validators. Until then, Tier 1 stays a tie.
- More calls. A 100/100 would have an interval of about 96% to 100% (calculation).
- Your own validator. Run the Tier 1 models on your tasks with a strict check. Then compare the intervals.
How we measured
- Hard set: 8 tasks, each with a deterministic validator in a sandbox without network. Claude Code: 3 repetitions per task (n = 24). Codex CLI: 2 (n = 16). One turn, tools off, a fresh folder.
- Strict pass: the whole reply passes as given. A right answer in the wrong format is a format miss. Intervals are 95% Wilson intervals.
- Costs: reported tokens × list price for every call, divided by strict passes. These are calculations; the calls ran on subscriptions.
- SWE-bench: one Agent attempt per instance; the panel is public mini-SWE-agent v2 runs.
Caveats
- Lenient reading. If format misses count, Haiku's interval (47% to 82%) overlaps GPT-6.1 Sol's (81% to 100%). We report the strict result.
- Different days. The Claude and Codex batches ran on different days on the same host.
- SWE-bench limits. n = 33, older panel models, and public issues, so training-data contamination is not controlled.
- Repricing is not a run. It uses Sonnet's tokens; another model would use different ones. List prices change.
What to read next
- The leaderboard and every comparison
- An AI model leaderboard without a composite score
- How to read AI benchmarks honestly
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable
Find the best model for your own work
Agent records the model, tokens and result of every step, so you can see where a different model changes the outcome. Try Agent.