• Head to head
  • Hard Tasks
  • Claude Haiku
  • Claude Sonnet
  • Claude Opus
  • Claude Fable
  • GPT-6.1 Sol
  • Codex CLI
  • Format Misses
  • Latency

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

On 8 hard tasks with deterministic validators, does pass rate separate the Claude Code models and GPT-6.1 Sol through the Codex CLI, and what do speed, tokens and cost per pass add?

Published · 7 charts · Download the data or a carousel

91%

95% CI 86%–95% · n = 152

139/152 · Calls that passed strictly (hard set)

The answer

139 of 152 calls that reached a model passed strictly (91%). 6 of 7 configurations passed every call: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Fable 5.1 · Claude Code (24/24 each, 95% interval 86% to 100%) and GPT-6.1 Sol (medium) · Codex CLI, GPT-6.1 Sol (high) · Codex CLI (16/16 each, 95% interval 81% to 100%), so the hard set still has a ceiling for these models and pass rate does not separate them. Claude Haiku 4.5 · Claude Code passed 11/24 strictly (46%, 95% interval 28% to 65%). 5 more replies had the right answer in the wrong format (for example inside a code fence), so 16/24 on a lenient reading (47% to 82%); 8 replies were wrong. It passed none of the tasks “Predict JavaScript event-loop output order”, “Solve a multi-constraint room schedule” and “Write a SQLite reporting query (fan-out, ties, boundaries)”. Median total time per call was Sonnet 7.7 s, Opus 9.2 s, Opus (high) 11.0 s, GPT-6.1 Sol (medium) 13.1 s, Fable 16.1 s, GPT-6.1 Sol (high) 18.1 s, Haiku 39.0 s. The fastest and slowest single calls of every configuration overlap with every other, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on a subscription), the lowest cost per strict pass was Claude Sonnet 5.5 · Claude Code at $0.0143; the quality-vs-cost frontier is Claude Sonnet 5.5 · Claude Code. 30 earlier attempts were blocked before any model call (Codex CLI: the CLI reported no signed-in account); they are reported, not scored, and that route ran in a later batch.

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 44 sHaiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.

Transcript
  1. Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
  2. 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  3. Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  4. Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  5. Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  6. Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  7. Open benchmarks: intervals, sources and every failure kept.

Key numbers

95% (144/152)

Calls with a correct answer, format misses included (lenient reading)

95% CI 90%–97% · n = 152

5

Non-passes that were format misses, not wrong answers

of 13 non-passes (8 wrong answers) · n = 13

6

Configurations that passed every call

of 7 (4 at 24/24, 2 at 16/16) · n = 7

7.7s

Claude Sonnet 5.5 · Claude Code

Lowest observed median time among configurations that passed every call (separate batches)

n = 24

$0.0143

Claude Sonnet 5.5 · Claude Code

Lowest list-price cost per strict pass (calculation)

n = 24

30

Attempts blocked before any model call (not scored)

(Codex CLI; 0 model calls) · n = 182

Estimate what serial waiting costs your team

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
  • Strict pass
  • Format miss (correct answer, wrong format)
  • Wrong answer
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

One square per call; counts at the right are exact and in legend order.

7 rows, 3 series: Strict pass, Format miss (correct answer, wrong format), Wrong answer. Strict pass: highest Claude Sonnet 5.5 · Claude Code 24 (n 24). Lowest Claude Haiku 4.5 · Claude Code 11 (n 24). Format miss (correct answer, wrong format): highest Claude Haiku 4.5 · Claude Code 5 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI 0 (n 16).

Notesn 16–24 per row

Counts per configuration: strict passes, format misses and wrong answers

A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
Entrance: medians race at 28× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

7 rows. Slowest Claude Haiku 4.5 · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 16–24 per row

Median per configuration; whiskers = fastest and slowest call

One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
Entrance: medians race at 25× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

7 rows. Slowest Claude Haiku 4.5 · Claude Code 35.5 s (range 12.9 s–70.3 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 6 s (range 0.9 s–30.6 s, n 24). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 16–24 per row

Median per configuration; whiskers = fastest and slowest call

One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).

Notesn 16–24 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
  • Claude Code
  • Codex CLI
Better: upper left

Haloed: on the frontier (1 of 7). A point in the shaded area is no better on either axis than a haloed point.

List-price calculation, not a run. 7 points: Strict pass rate against USD per strict pass (list-price calculation). USD per strict pass (list-price calculation) runs from $0.014 to $0.093; Strict pass rate from 46% to 100%. Highlighted: Claude Sonnet 5.5 · Claude Code.

Notesn 16–24 per point

Strict pass rate against list-price cost per strict pass

Upper-left is better. Highlighted points are on the frontier: no other configuration passes at least as often for at most the same cost per pass. Frontier: Claude Sonnet 5.5 · Claude Code. Costs are calculations from tokens. Pass rates with their 95% intervals are in the pass-rate chart.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)

Tables

Pass matrix on hard tasks: configuration × task (strict passes)

Configuration by column: each cell is k of n from the table; a cell opens its model

ConfigurationFix an interval-merge function (off-by-one and edge cases)1Fix a time-zone day-length function (DST)2Write a CSV parser (quoted newlines, strict errors)3Predict JavaScript event-loop output order4Solve a multi-constraint room schedule5Write a strict SemVer 2.0.0 regex6Refactor to remove duplication, keep 20 tests green7Write a SQLite reporting query (fan-out, ties, boundaries)8
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code
  1. Fix an interval-merge function (off-by-one and edge cases)
  2. Fix a time-zone day-length function (DST)
  3. Write a CSV parser (quoted newlines, strict errors)
  4. Predict JavaScript event-loop output order
  5. Solve a multi-constraint room schedule
  6. Write a strict SemVer 2.0.0 regex
  7. Refactor to remove duplication, keep 20 tests green
  8. Write a SQLite reporting query (fan-out, ties, boundaries)

Shade is the share k of n in each cell. * The cell has a note (hover or focus it). Totals and other columns are in the Table view.

Counts from the table, not an intervaln = 2–3 per cell

7 rows by 8 columns. 50 of 56 k/n cells are full (n of n) and 3 are zero.

Validator controls run before the first model call

TaskReference answerChecksPlausible wrong answers rejectedWrapped reference flagged as format miss
Fix an interval-merge function (off-by-one and edge cases)passes185/5yes
Fix a time-zone day-length function (DST)passes363/3yes
Write a CSV parser (quoted newlines, strict errors)passes223/3yes
Predict JavaScript event-loop output orderpasses13/3yes
Solve a multi-constraint room schedulepasses143/3yes
Write a strict SemVer 2.0.0 regexpasses303/3yes
Refactor to remove duplication, keep 20 tests greenpasses252/2yes
Write a SQLite reporting query (fan-out, ties, boundaries)passes84/4yes

Method

  1. Protocol declared before the first call. A follow-up to the five-task head-to-head (/benchmarks/model-head-to-head), where pass rate hit a ceiling.
  2. 8 hard tasks, each with a deterministic validator that runs in a sandbox without network: Fix an interval-merge function (off-by-one and edge cases); Fix a time-zone day-length function (DST); Write a CSV parser (quoted newlines, strict errors); Predict JavaScript event-loop output order; Solve a multi-constraint room schedule; Write a strict SemVer 2.0.0 regex; Refactor to remove duplication, keep 20 tests green; Write a SQLite reporting query (fan-out, ties, boundaries).
  3. Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 references wrapped in a fence or prose are flagged as format misses.
  4. Configurations that reached a model: Sonnet, Opus, Opus (high), GPT-6.1 Sol (medium), Fable, GPT-6.1 Sol (high), Haiku. Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task).
  5. Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no code fence and no other text. This is stricter than the five-task study, which removed one wrapping fence.
  6. Format miss: the strict check failed, but a lenient extractor (fenced block, outer JSON, first code line, one output line) finds an answer that passes the same validator. Reported apart from wrong answers, never as a pass.
  7. Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call. One call at a time per account.
  8. Attempts: Claude Code: 120 attempts, 120 reached a model, 0 blocked; Codex CLI: 62 attempts, 32 reached a model, 30 blocked. Nothing was trimmed or retried, and no run hit a usage or rate limit.
  9. One Codex batch was resumed after its process ended (declared in the protocol before the resume): the resume skipped every task, repetition and effort the batch file already held, so no recorded call was repeated or replaced.
  10. Cost per strict pass: list price × reported tokens for every call in the configuration (cache reads and writes priced as in the five-task study), divided by its strict passes. A calculation.

Caveats

  • 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  • Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  • Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  • The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.
  • Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  • CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.
  • Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  • List-price costs are calculations; the calls used a flat subscription.

Sources

  • Provider head-to-head, hard set: eight hard tasks with strict validators

    Our recorded runs ·

    Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.

    Raw data: provider-h2h-hard/receipts.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/hard-model-head-to-head.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

More comparisons based on this study (5)

These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.

More write-ups that cite this study (1)

Models and comparisons in this study

More studies

All benchmarks
Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.