• Tier List
  • Head to head
  • Claude Sonnet
  • Claude Opus

What is the best AI model for coding? Our data says four models tie

4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.

TL;DR

  • On our hard coding set, four models share the top tier: Claude Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol in the Codex CLI. Six configurations passed every call. Sonnet, Opus (default and high effort) and Fable passed 24/24 (95% interval 86% to 100%). GPT-6.1 Sol at medium and high effort passed 16/16 (81% to 100%).
  • Claude Haiku 4.5 is the only model whose interval separates. It passed 11/24 strictly (46%, 95% interval 28% to 65%).
  • The set is at its ceiling. A tie means our tasks cannot find a gap, not that the models are equal.
  • On SWE-bench Verified, the panel is one tier too. On the same 33 instances, 11 public runs resolved 21 to 28, and Agent resolved 25 (59% to 87%). Every interval overlaps.
  • Inside the top tier, choose on cost, speed and route. Sonnet had the lowest list-price cost per strict pass at $0.0143. Fable cost $0.0933, 6.5x as much (calculation). The speed ranges overlap.
Live story · 44 sHaiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.

Transcript
  1. Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
  2. 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  3. Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  4. Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  5. Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  6. Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  7. Open benchmarks: intervals, sources and every failure kept.

The short answer

A list that ranks models one to ten can hide ties. Our rule: models that our data cannot separate share one tier. On eight hard tasks with strict validators, that gives two tiers: four models in Tier 1 and Haiku 4.5 in Tier 2. Study: /benchmarks/hard-model-head-to-head.

The tier list, built from intervals

Tier 1 holds the option with the highest rate and every option whose 95% interval overlaps it. The next tier starts at the first option whose interval sits fully below, and the rule repeats.

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

TierConfigurationStrict passes95% intervalCost per strict pass (calculation)Median time per call (fastest to slowest)
1Claude Sonnet 5.5 · Claude Code24/2486% to 100%$0.01437.7 s (2.3 to 34.8 s)
1Claude Opus 5.5 · Claude Code24/2486% to 100%$0.02829.2 s (4.2 to 27.2 s)
1Claude Opus 5.5 (high) · Claude Code24/2486% to 100%$0.033411.0 s (3.6 to 63.0 s)
1GPT-6.1 Sol (medium) · Codex CLI16/1681% to 100%$0.025613.1 s (8.5 to 61.6 s)
1Claude Fable 5.1 · Claude Code24/2486% to 100%$0.093316.1 s (4.5 to 90.0 s)
1GPT-6.1 Sol (high) · Codex CLI16/1681% to 100%$0.015118.1 s (11.7 to 92.2 s)
2Claude Haiku 4.5 · Claude Code11/2428% to 65%$0.067239.0 s (15.3 to 75.1 s)

Haiku's interval stops at 65%; the lowest bound in Tier 1 is 81%. They do not overlap, so Tier 1 is ahead of Haiku on this set. Haiku's 13 non-passes were 5 format misses (right answer, wrong format) and 8 wrong answers.

Repeated prompts agree. Sonnet and GPT-6.1 Sol (medium) passed 10/10 on each of three prompts (72% to 100%). Haiku passed 0/10 on an exact-number prompt (0% to 28%) (study).

Effort does not change the tier. On the effort ladder, all 11 cells of Sonnet, Opus and GPT-6.1 Sol passed 16/16 (81% to 100% each).

SWE-bench Verified: one tier for twelve systems

GPT 5.2 (high)
Gemini 3 Flash (high)
GLM 5 (high)
Agent (Sonnet 5.5, full pipeline)
Claude 4.5 Sonnet (high)
Claude 4.5 Haiku (high)
Claude 4.5 Opus (high)
DeepSeek V3.2 (high)
MiniMax M2.5 (high)
Claude 4.6 Opus
Kimi K2.5 (high)
GPT 5 mini

Every interval overlaps every other: this chart does not order these rows.

12 rows. Highest GPT 5.2 (high) 85% (95% interval 69%–93%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 33 per row

Agent vs 11 public mini-SWE-agent v2 runs, one attempt each

Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

On the same 33 instances, the 11 public mini-SWE-agent v2 runs resolved 21 to 28. Agent, our pipeline on Sonnet 5.5, resolved 25 (76%, 95% interval 59% to 87%).

The highest lower bound is 69% (GPT 5.2 at high effort, 28/33). The lowest upper bound is 78% (GPT 5 mini, 21/33). Because 69% is below 78%, every pair of intervals overlaps. All twelve systems share Tier 1.

Note: the panel is third-party, with older models in a bash-only harness. Agent is a full pipeline, not a bare model. Study: /benchmarks/swe-bench-verified.

How to choose inside Tier 1

1. Cost per pass (a calculation)

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

  • Sonnet 5.5: $0.0143 per strict pass, the lowest.
  • GPT-6.1 Sol (high): $0.0151, about 5.5% more.
  • Opus 5.5: $0.0282, 1.97x Sonnet.
  • Fable 5.1: $0.0933, 6.5x Sonnet.

Haiku lists at half of Sonnet's price, but it cost $0.0672 per strict pass, 4.7x Sonnet.

Calculation
  • Claude Code
  • Codex CLI
Better: upper left

Haloed: on the frontier (1 of 7). A point in the shaded area is no better on either axis than a haloed point.

List-price calculation, not a run. 7 points: Strict pass rate against USD per strict pass (list-price calculation). USD per strict pass (list-price calculation) runs from $0.014 to $0.093; Strict pass rate from 46% to 100%. Highlighted: Claude Sonnet 5.5 · Claude Code.

Notesn 16–24 per point

Strict pass rate against list-price cost per strict pass

Upper-left is better. Highlighted points are on the frontier: no other configuration passes at least as often for at most the same cost per pass. Frontier: Claude Sonnet 5.5 · Claude Code. Costs are calculations from tokens. Pass rates with their 95% intervals are in the pass-rate chart.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Only Sonnet 5.5 in Claude Code sits on the frontier. No other configuration passes as often for the same cost per pass or less.

The cheapest option depends on your workload. We repriced our agent's SWE-bench tokens (94.0% cache reads) at each list price. Per resolved instance, GPT-6.1 Sol $2.10, Sonnet $3.49, Opus $5.75, Fable $12.85 (calculation). GPT-6.1 Sol charges $0.10 per million cache-read tokens; Sonnet charges $0.20.

On the effort ladder, low effort was cheapest per strict pass for each model (calculation): Sonnet $0.0122, Opus $0.0212, GPT-6.1 Sol $0.0128.

2. Speed (the ranges overlap)

Tier 1 medians ran from 7.7 s (Sonnet) to 18.1 s (GPT-6.1 Sol, high), but the per-call ranges overlap. The medians describe one batch. They do not rank the models.

3. Route: CLI or API

On five short tasks, the Codex CLI sent a median 12,124 input tokens per call; Claude Code sent 2,130. On a scheduler repair, the same GPT-6.1 Sol took a median 61.2 s through the Codex CLI and 17.3 s through the OpenAI API. We ran each route 3 times, and the ranges do not overlap. More: the hidden context tax.

Compare pages: Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI) and Sonnet 5.5 vs Opus 5.5.

Why we do not crown one model

  1. Ceiling. 6 of 7 configurations passed every call. When nearly every option scores 100%, the set cannot rank them.
  2. Small samples. A 24/24 has a 95% interval of 86% to 100%; a 16/16 has 81% to 100%. A tie at these sizes can hide a gap of about 14 to 19 points (calculation).
  3. Route and model together. A Claude-vs-GPT row compares Claude Code and the Codex CLI too, not the models alone.
  4. The tier depends on the tasks. On five easier tasks, all nine configurations shared one tier: 127 of 130 calls passed (93% to 99%), and Haiku passed 15/15 (80% to 100%).

What would change the answer

  1. Harder tasks. We plan a harder set with strict validators. Until then, Tier 1 stays a tie.
  2. More calls. A 100/100 would have an interval of about 96% to 100% (calculation).
  3. Your own validator. Run the Tier 1 models on your tasks with a strict check. Then compare the intervals.

How we measured

  • Hard set: 8 tasks, each with a deterministic validator in a sandbox without network. Claude Code: 3 repetitions per task (n = 24). Codex CLI: 2 (n = 16). One turn, tools off, a fresh folder.
  • Strict pass: the whole reply passes as given. A right answer in the wrong format is a format miss. Intervals are 95% Wilson intervals.
  • Costs: reported tokens × list price for every call, divided by strict passes. These are calculations; the calls ran on subscriptions.
  • SWE-bench: one Agent attempt per instance; the panel is public mini-SWE-agent v2 runs.

Caveats

  • Lenient reading. If format misses count, Haiku's interval (47% to 82%) overlaps GPT-6.1 Sol's (81% to 100%). We report the strict result.
  • Different days. The Claude and Codex batches ran on different days on the same host.
  • SWE-bench limits. n = 33, older panel models, and public issues, so training-data contamination is not controlled.
  • Repricing is not a run. It uses Sonnet's tokens; another model would use different ones. List prices change.

Find the best model for your own work

Agent records the model, tokens and result of every step, so you can see where a different model changes the outcome. Try Agent.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.