Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
On 8 hard tasks with deterministic validators, does pass rate separate the Claude Code models and GPT-6.1 Sol through the Codex CLI, and what do speed, tokens and cost per pass add?
Published · 7 charts · Download the data or a carousel
91%
The answer
139 of 152 calls that reached a model passed strictly (91%). 6 of 7 configurations passed every call: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Fable 5.1 · Claude Code (24/24 each, 95% interval 86% to 100%) and GPT-6.1 Sol (medium) · Codex CLI, GPT-6.1 Sol (high) · Codex CLI (16/16 each, 95% interval 81% to 100%), so the hard set still has a ceiling for these models and pass rate does not separate them. Claude Haiku 4.5 · Claude Code passed 11/24 strictly (46%, 95% interval 28% to 65%). 5 more replies had the right answer in the wrong format (for example inside a code fence), so 16/24 on a lenient reading (47% to 82%); 8 replies were wrong. It passed none of the tasks “Predict JavaScript event-loop output order”, “Solve a multi-constraint room schedule” and “Write a SQLite reporting query (fan-out, ties, boundaries)”. Median total time per call was Sonnet 7.7 s, Opus 9.2 s, Opus (high) 11.0 s, GPT-6.1 Sol (medium) 13.1 s, Fable 16.1 s, GPT-6.1 Sol (high) 18.1 s, Haiku 39.0 s. The fastest and slowest single calls of every configuration overlap with every other, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on a subscription), the lowest cost per strict pass was Claude Sonnet 5.5 · Claude Code at $0.0143; the quality-vs-cost frontier is Claude Sonnet 5.5 · Claude Code. 30 earlier attempts were blocked before any model call (Codex CLI: the CLI reported no signed-in account); they are reported, not scored, and that route ran in a later batch.
Live story
Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks
139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.
Transcript
- Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
- 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
- Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
- Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
- Open benchmarks: intervals, sources and every failure kept.
Key numbers
95% (144/152)
Calls with a correct answer, format misses included (lenient reading)
95% CI 90%–97% · n = 152
5
Non-passes that were format misses, not wrong answers
of 13 non-passes (8 wrong answers) · n = 13
6
Configurations that passed every call
of 7 (4 at 24/24, 2 at 16/16) · n = 7
7.7s
Claude Sonnet 5.5 · Claude Code
Lowest observed median time among configurations that passed every call (separate batches)
n = 24
$0.0143
Claude Sonnet 5.5 · Claude Code
Lowest list-price cost per strict pass (calculation)
n = 24
30
Attempts blocked before any model call (not scored)
(Codex CLI; 0 model calls) · n = 182
Estimate what serial waiting costs your team
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
| Claude Opus 5.5 · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
| Claude Opus 5.5 (high) · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 100% | Strict pass: 81%–100%; Lenient (format misses counted): 81%–100% | 16 |
| Claude Fable 5.1 · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
| GPT-6.1 Sol (high) · Codex CLI | 100% | 100% | Strict pass: 81%–100%; Lenient (format misses counted): 81%–100% | 16 |
| Claude Haiku 4.5 · Claude Code | 46% | 67% | Strict pass: 28%–65%; Lenient (format misses counted): 47%–82% | 24 |
7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
- Strict pass
- Format miss (correct answer, wrong format)
- Wrong answer
One square per call; counts at the right are exact and in legend order.
| Item | Strict pass | Format miss (correct answer, wrong format) | Wrong answer | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 24 | 0 | 0 | 24 |
| Claude Opus 5.5 · Claude Code | 24 | 0 | 0 | 24 |
| Claude Opus 5.5 (high) · Claude Code | 24 | 0 | 0 | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 16 | 0 | 0 | 16 |
| Claude Fable 5.1 · Claude Code | 24 | 0 | 0 | 24 |
| GPT-6.1 Sol (high) · Codex CLI | 16 | 0 | 0 | 16 |
| Claude Haiku 4.5 · Claude Code | 11 | 5 | 8 | 24 |
7 rows, 3 series: Strict pass, Format miss (correct answer, wrong format), Wrong answer. Strict pass: highest Claude Sonnet 5.5 · Claude Code 24 (n 24). Lowest Claude Haiku 4.5 · Claude Code 11 (n 24). Format miss (correct answer, wrong format): highest Claude Haiku 4.5 · Claude Code 5 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI 0 (n 16).
Notesn 16–24 per row
Counts per configuration: strict passes, format misses and wrong answers
A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call on hard tasks (separate batches) | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 7.8 s | 2.3 s–34.8 s | 24 |
| Claude Opus 5.5 · Claude Code | 9.2 s | 4.2 s–27.2 s | 24 |
| Claude Opus 5.5 (high) · Claude Code | 11 s | 3.6 s–63 s | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 8.5 s–61.6 s | 16 |
| Claude Fable 5.1 · Claude Code | 16.1 s | 4.5 s–90 s | 24 |
| GPT-6.1 Sol (high) · Codex CLI | 18.1 s | 11.7 s–92.2 s | 16 |
| Claude Haiku 4.5 · Claude Code | 39 s | 15.3 s–75.1 s | 24 |
7 rows. Slowest Claude Haiku 4.5 · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 6 s | 0.9 s–30.6 s | 24 |
| Claude Opus 5.5 · Claude Code | 6.8 s | 2.4 s–21.8 s | 24 |
| Claude Opus 5.5 (high) · Claude Code | 7.1 s | 2.2 s–56.2 s | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 10.2 s | 6.1 s–40.4 s | 16 |
| Claude Fable 5.1 · Claude Code | 11.6 s | 2 s–85.3 s | 24 |
| GPT-6.1 Sol (high) · Codex CLI | 12.7 s | 8.9 s–75.9 s | 16 |
| Claude Haiku 4.5 · Claude Code | 35.5 s | 12.9 s–70.3 s | 24 |
7 rows. Slowest Claude Haiku 4.5 · Claude Code 35.5 s (range 12.9 s–70.3 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 6 s (range 0.9 s–30.6 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1,050 | 585 | 24 |
| Claude Opus 5.5 · Claude Code | 945 | 529 | 24 |
| Claude Opus 5.5 (high) · Claude Code | 1,052 | 614 | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 335 | 150 | 16 |
| Claude Fable 5.1 · Claude Code | 1,366 | 889 | 24 |
| GPT-6.1 Sol (high) · Codex CLI | 436 | 225 | 16 |
| Claude Haiku 4.5 · Claude Code | 5,064 | 4,556 | 24 |
7 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).
Notesn 16–24 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.014 | 24 |
| GPT-6.1 Sol (high) · Codex CLI | $0.015 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | 16 |
| Claude Opus 5.5 · Claude Code | $0.028 | 24 |
| Claude Opus 5.5 (high) · Claude Code | $0.033 | 24 |
| Claude Haiku 4.5 · Claude Code | $0.067 | 24 |
| Claude Fable 5.1 · Claude Code | $0.093 | 24 |
List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).
Notesn 16–24 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Claude Code
- Codex CLI
Haloed: on the frontier (1 of 7). A point in the shaded area is no better on either axis than a haloed point.
| Point | Series | USD per strict pass (list-price calculation) | Strict pass rate | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | Claude Code | $0.014 | 100% | 24 |
| Claude Opus 5.5 · Claude Code | Claude Code | $0.028 | 100% | 24 |
| Claude Opus 5.5 (high) · Claude Code | Claude Code | $0.033 | 100% | 24 |
| Claude Fable 5.1 · Claude Code | Claude Code | $0.093 | 100% | 24 |
| Claude Haiku 4.5 · Claude Code | Claude Code | $0.067 | 46% | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | Codex CLI | $0.026 | 100% | 16 |
| GPT-6.1 Sol (high) · Codex CLI | Codex CLI | $0.015 | 100% | 16 |
List-price calculation, not a run. 7 points: Strict pass rate against USD per strict pass (list-price calculation). USD per strict pass (list-price calculation) runs from $0.014 to $0.093; Strict pass rate from 46% to 100%. Highlighted: Claude Sonnet 5.5 · Claude Code.
Notesn 16–24 per point
Strict pass rate against list-price cost per strict pass
Upper-left is better. Highlighted points are on the frontier: no other configuration passes at least as often for at most the same cost per pass. Frontier: Claude Sonnet 5.5 · Claude Code. Costs are calculations from tokens. Pass rates with their 95% intervals are in the pass-rate chart.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
Tables
Pass matrix on hard tasks: configuration × task (strict passes)
Configuration by column: each cell is k of n from the table; a cell opens its model
- Fix an interval-merge function (off-by-one and edge cases)
- Fix a time-zone day-length function (DST)
- Write a CSV parser (quoted newlines, strict errors)
- Predict JavaScript event-loop output order
- Solve a multi-constraint room schedule
- Write a strict SemVer 2.0.0 regex
- Refactor to remove duplication, keep 20 tests green
- Write a SQLite reporting query (fan-out, ties, boundaries)
Shade is the share k of n in each cell. * The cell has a note (hover or focus it). Totals and other columns are in the Table view.
| Configuration | Fix an interval-merge function (off-by-one and edge cases) | Fix a time-zone day-length function (DST) | Write a CSV parser (quoted newlines, strict errors) | Predict JavaScript event-loop output order | Solve a multi-constraint room schedule | Write a strict SemVer 2.0.0 regex | Refactor to remove duplication, keep 20 tests green | Write a SQLite reporting query (fan-out, ties, boundaries) | Strict total | Format misses | Wrong answers | Median total (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 24/24 | 0 | 0 | 7.8 s |
| Claude Opus 5.5 · Claude Code | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 24/24 | 0 | 0 | 9.2 s |
| Claude Opus 5.5 (high) · Claude Code | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 24/24 | 0 | 0 | 11 s |
| GPT-6.1 Sol (medium) · Codex CLI | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 16/16 | 0 | 0 | 13.1 s |
| Claude Fable 5.1 · Claude Code | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 24/24 | 0 | 0 | 16.1 s |
| GPT-6.1 Sol (high) · Codex CLI | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 16/16 | 0 | 0 | 18.1 s |
| Claude Haiku 4.5 · Claude Code | 3/3 | 1/3 | 2/3 (+1 format miss) | 0/3 | 0/3 (+2 format misses) | 3/3 | 2/3 | 0/3 (+2 format misses) | 11/24 | 5 | 8 | 39 s |
Counts from the table, not an intervaln = 2–3 per cell
7 rows by 8 columns. 50 of 56 k/n cells are full (n of n) and 3 are zero.
Validator controls run before the first model call
| Task | Reference answer | Checks | Plausible wrong answers rejected | Wrapped reference flagged as format miss |
|---|---|---|---|---|
| Fix an interval-merge function (off-by-one and edge cases) | passes | 18 | 5/5 | yes |
| Fix a time-zone day-length function (DST) | passes | 36 | 3/3 | yes |
| Write a CSV parser (quoted newlines, strict errors) | passes | 22 | 3/3 | yes |
| Predict JavaScript event-loop output order | passes | 1 | 3/3 | yes |
| Solve a multi-constraint room schedule | passes | 14 | 3/3 | yes |
| Write a strict SemVer 2.0.0 regex | passes | 30 | 3/3 | yes |
| Refactor to remove duplication, keep 20 tests green | passes | 25 | 2/2 | yes |
| Write a SQLite reporting query (fan-out, ties, boundaries) | passes | 8 | 4/4 | yes |
Method
- Protocol declared before the first call. A follow-up to the five-task head-to-head (/benchmarks/model-head-to-head), where pass rate hit a ceiling.
- 8 hard tasks, each with a deterministic validator that runs in a sandbox without network: Fix an interval-merge function (off-by-one and edge cases); Fix a time-zone day-length function (DST); Write a CSV parser (quoted newlines, strict errors); Predict JavaScript event-loop output order; Solve a multi-constraint room schedule; Write a strict SemVer 2.0.0 regex; Refactor to remove duplication, keep 20 tests green; Write a SQLite reporting query (fan-out, ties, boundaries).
- Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 references wrapped in a fence or prose are flagged as format misses.
- Configurations that reached a model: Sonnet, Opus, Opus (high), GPT-6.1 Sol (medium), Fable, GPT-6.1 Sol (high), Haiku. Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task).
- Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no code fence and no other text. This is stricter than the five-task study, which removed one wrapping fence.
- Format miss: the strict check failed, but a lenient extractor (fenced block, outer JSON, first code line, one output line) finds an answer that passes the same validator. Reported apart from wrong answers, never as a pass.
- Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call. One call at a time per account.
- Attempts: Claude Code: 120 attempts, 120 reached a model, 0 blocked; Codex CLI: 62 attempts, 32 reached a model, 30 blocked. Nothing was trimmed or retried, and no run hit a usage or rate limit.
- One Codex batch was resumed after its process ended (declared in the protocol before the resume): the resume skipped every task, repetition and effort the batch file already held, so no recorded call was repeated or replaced.
- Cost per strict pass: list price × reported tokens for every call in the configuration (cache reads and writes priced as in the five-task study), divided by its strict passes. A calculation.
Caveats
- 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
- Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.
- Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
- CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.
- Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
- List-price costs are calculations; the calls used a flat subscription.
Sources
Provider head-to-head, hard set: eight hard tasks with strict validators
Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.
Repricing calculation
Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Token prices as listed by the vendor on 2026-10-03.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/hard-model-head-to-head.
Explainers that cite this study
Read the methods and terms in the context of these recorded results.
- Benchmark saturation: when every model scores 100%
- Cost per correct answer: the LLM price that counts failures
- How many runs does an LLM eval need? Sample size, with real intervals
- How to read AI benchmarks honestly
- The Pareto frontier: reading LLM cost against quality
- pass@k explained: pass@1, pass@k and pass^k with real runs
- Reasoning tokens, explained: the hidden output you pay for
- Strict grading and format misses: when a right answer fails
- Time to first token (TTFT), explained with CLI and API timings
- What is reasoning effort?
- Wilson confidence intervals for AI benchmarks
More comparisons based on this study (5)
These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.
More write-ups that cite this study (1)
Models and comparisons in this study
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Haiku 4.5
- Claude Fable 5.1
- GPT-6.1 Sol (Codex CLI)
- Claude Code
- Codex CLI
- Claude Sonnet 5.5 vs Claude Opus 5.5
- Claude Haiku 4.5 vs Claude Sonnet 5.5
- Claude Opus 5.5 vs Claude Fable 5.1
- Claude Sonnet 5.5 vs Claude Fable 5.1
- Claude Haiku 4.5 vs Claude Opus 5.5
- Claude Code vs Codex CLI
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)
- Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)
- All comparisons
Write-ups on this study
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
AI coding cost per developer: a formula built on recorded work
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
Best LLM for JSON output? Haiku, Sonnet and GPT-6.1 Sol, 10 runs each
Best LLM for JSON output? On one prompt, Sonnet and GPT-6.1 Sol passed 10/10 (95%: 72–100%); Haiku passed 1/10 (2–40%). CLI results, not JSON mode.
Claude Code cost per task: a price ladder from one decision to one agent run
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.
Claude Fable 5.1 vs Opus 5.5 vs Sonnet 5.5: speed, tokens and price tested
24 of 24: Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed every hard task. Fable cost 3.3x Opus and 6.5x Sonnet per pass (list-price calculation).
Claude Haiku 4.5 vs Sonnet 5.5: all 80 comparison rows, and where the small model loses
Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.
Does LLM routing save money? The saving, the router and the net
Routing would save 2.8% ($3.01) on 2,362 recorded calls (a calculation). A Sonnet router on every call costs about $11.80, so the net is a loss.
Format misses vs wrong answers: your LLM eval may be failing right answers
5 of 13 failed calls on our hard set were right answers in the wrong format. How to grade LLM output strictly, test validators first and report both numbers.
GPT-6.1 Sol vs Claude Opus 5.5 and Sonnet 5.5 on harder tasks: only one strict pair separates
GPT-6.1 Sol vs Claude Opus and Sonnet on four selected harder tasks. Pass intervals overlap for these models; only Sol vs Haiku separates on strict passes.
GPT-6.1 Sol vs Claude Sonnet 5.5 vs Opus 5.5: every row we measured
Of 70 comparison rows for GPT-6.1 Sol, Claude Sonnet 5.5 and Opus 5.5, only 2 have a winner (speed). All 15 pass-rate rows tie. Tokens, price and route differ.
How long does an AI coding agent take per task? Minutes, calls and where the time goes
Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.
How many runs do you need to compare two AI models? A sample-size table
Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.
How much of your AI bill is thinking tokens? Claude and GPT-6.1 Sol, measured
Thinking tokens: 46% to 92% of output per median call, 9% to 80% of pooled list-price cost. Calculation over recorded Claude and GPT-6.1 Sol calls.
How to turn off extended thinking in Claude Code, and what we measured for Haiku 4.5
Set MAX_THINKING_TOKENS=0, then check both counters. Our Haiku 4.5 test covers routing and hard tasks, with timings, costs and uncertainty.
Is Claude Haiku cheaper if you retry or escalate to Sonnet? A calculation on real receipts
Haiku 4.5 first, one retry, then Sonnet 5.5? On 8 hard tasks, the calculation gives 4.3x the cost and 8.7x the time per correct answer versus Sonnet alone.
Is Claude Haiku cheaper than Sonnet? Cost per correct answer, with retries
Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.
LLM API pricing comparison, October 2026: Claude vs GPT vs Gemini per million tokens
21 LLM API prices per million tokens, October 2026: Claude, GPT, Gemini and more. Blended prices span 100x. Price per token is not price per task.
Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
Opus at low effort or Sonnet at high effort? A bigger model that thinks less, tested
Opus 5.5 at low effort and Sonnet 5.5 at high effort both passed 16/16. Opus low cost 1.3x as much per pass (calculation). Sonnet at low effort cost least.
Plan for p95, not the median: LLM tail latency in our runs
4.30 s p95 against a 2.60 s median for Claude Sonnet 5.5; 34.5 s against 12.5 s for Haiku 4.5. Measured LLM tail latency and what to do about it.
The cheapest way to run an AI coding agent: 7 levers from measured runs
7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.
What does an agent loop cost? 1.3 to 8.7 times the median tokens (calculation)
Agent loop cost on 8 hard tasks: 1.3–8.7 times the median tokens and up to 2.1 times the price per strict pass (calculations), with no clear pass-rate gain.
What is the best AI model for coding? Our data says four models tie
4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.
What thinking costs: reasoning tokens per call for Claude and GPT-6.1 Sol
Haiku 4.5 spent a median 4,556 reasoning tokens per hard call, Sonnet 5.5 585, GPT-6.1 Sol 150. List price per 1,000 calls: $22.78, $5.85, $1.50 (calculation).
Which Claude model is fastest? It depends on the task, and on thinking
1.94 s was the lowest median on short calls (Fable 5.1). 7.75 s on hard calls (Sonnet 5.5). Haiku 4.5 took 4.43 s and 39.01 s with default thinking.
Which Claude model should you use? A task-by-task guide from our measurements
Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.
Why AI coding agents fail on real pull requests: every failure from 12 attempts
32 AI coding agent misses from our own runs, sorted into 8 classes: wrong answers, gates, caps, lost context and more. Counts, not rates.
Why is Claude Code slow? Where the seconds go in a coding CLI call, and what to change
A one-word Claude Code call took a median 2.5 s, with 1.7 s outside the model (n = 5). Where the rest goes: thinking, effort, route. Measured splits and fixes.
Would majority voting fix it? A self-consistency thought experiment on 90 real calls
Calculation on 90 recorded calls: a vote over 10 Haiku replies returns 289, not 282. Repeated errors and format misses limit self-consistency voting.
Your AI agent says it is done. Is it? Claimed vs verified in our runs
6 of 6 Haiku 4.5 sessions invented a late-fee rate and reported done. Sonnet 5.5 flagged the gap in 9 of 9. Claimed vs verified, with small n.
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
An AI model leaderboard without a composite score: why, and how to read ours
Our leaderboard lists 46 models, CLIs, routers and providers with their best-supported facts, each with n and an interval. No single score. Here is why.
Claude Code vs Codex CLI on hidden tests: when every agent passes, what differs?
36 hidden-test sessions: Sonnet 5.5 and Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI. All passed. Time, tool calls, diffs and one confound differ.
Claude Code vs Codex CLI: the hidden context tax and what it does to speed
Codex CLI sent a median 12,124 input tokens per call; Claude Code sent 2,130. What a coding CLI's hidden context costs, and how much time the route adds.
Claude Sonnet 5.5 vs Opus 5.5: when is Opus worth the price?
Sonnet 5.5 and Opus 5.5 tied on every quality test we ran, easy, hard and agentic. Opus cost 1.6x to 2.6x per unit of work. Where the gap comes from.
Claude vs Codex on hard tasks: GPT-6.1 Sol joins the hard set
GPT-6.1 Sol in the Codex CLI passed 16/16 hard tasks at medium and high effort. Sonnet, Opus and Fable passed 24/24. What separates them: time and cost.
Does extended thinking pay for Claude Haiku 4.5? We turned it off and measured
Claude Haiku 4.5, thinking off: 2.7x faster (a calculation from two medians), no routing accuracy gap shown. Hard tasks: 4/24 passed, thinking on 11/24.
Does letting an AI run code help? A single call vs an agent loop on 8 hard tasks
Does tool use improve LLM accuracy? 8 hard tasks, 120 attempts (118 scored): single call vs agent loop for Haiku, Sonnet and GPT-6 Luna. No clear gain.
Does reasoning effort buy quality? Claude and Codex on hard tasks, low to high
176 calls on 8 hard tasks: Sonnet, Opus and GPT-6.1 Sol at low, medium and high effort. Every cell passed 16/16. Higher effort cost more tokens and money.
How to estimate your AI coding bill from real token mixes
Estimate AI coding costs from recorded token mixes: an agent task at $2.64, a hard call at $0.014, a routing decision at $0.005. Calculations, limits stated.
Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look, limits first
Interim: 3 of 8 paired SWE-bench Verified issues graded. Opus 5.5 resolved 2, Sonnet 5.5 1; McNemar p = 1.0. Opus cost 2.6x, a calculation. Limits first.
When tasks get hard: Claude Haiku vs Sonnet vs Opus vs Fable (and where Codex is)
120 calls on 8 hard tasks with strict validators. Sonnet, Opus and Fable passed 24/24; Haiku 11/24. Speed, tokens and cost per pass. Codex results: see update.
More studies
All benchmarksGPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.