When tasks get hard: Claude Haiku vs Sonnet vs Opus vs Fable (and where Codex is)
120 calls on 8 hard tasks with strict validators. Sonnet, Opus and Fable passed 24/24; Haiku 11/24. Speed, tokens and cost per pass. Codex results: see update.
Update, 2026-10-06: the Codex side now has data. A later batch reached the model, and GPT-6.1 Sol in the Codex CLI passed 16 of 16 at medium and at high effort. The study page now holds 152 scored calls. This post keeps its original Claude-only numbers; the Codex results are in Claude vs Codex on hard tasks.
TL;DR
- We ran 120 calls on 8 hard tasks, each with a strict, deterministic validator. 107 passed strictly: 89%, 95% interval 82% to 94%.
- Claude Sonnet 5.5, Opus 5.5, Opus 5.5 at high effort and Fable 5.1 each passed 24 of 24 (95% interval 86% to 100%). The hard set did not separate them.
- Claude Haiku 4.5 passed 11 of 24 strictly (46%, 95% interval 28% to 65%). 5 more replies had the right answer in the wrong format. 8 were wrong.
- Median total time per call: Sonnet 7.7 s, Opus 9.2 s, Opus high 11.0 s, Fable 16.1 s, Haiku 39.0 s. The per-call ranges overlap, so this is not a tested ranking.
- At list price (a calculation), the lowest cost per strict pass was Sonnet at $0.0143.
- Codex has no result. All 30 Codex CLI attempts were blocked before any model call. We count them, and we make no Claude-vs-Codex claim from this run.
Every call, control and validator result: /benchmarks/hard-model-head-to-head.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks
139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.
Transcript
- Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
- 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
- Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
- Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
- Open benchmarks: intervals, sources and every failure kept.
Why a hard set
Our first head-to-head used five short tasks. 127 of 130 calls passed, so pass rate hit a ceiling. What happens when the tasks get harder?
So we wrote eight tasks that a careless answer gets wrong: fix an interval-merge function, fix a time-zone day-length function across daylight saving time, write a strict CSV parser, predict JavaScript event-loop output order, solve a multi-constraint room schedule, write a strict SemVer 2.0.0 regex, refactor without breaking 20 tests, and write a SQLite reporting query with fan-out, ties and boundaries.
Each task has a validator that runs in a sandbox without network. The validators hold from 1 check (the event-loop order) to 36 checks (the DST function).
We tested the validators first
A hard benchmark with a weak validator measures nothing. Before the first model call, we ran three controls on every task:
- 8 of 8 reference answers passed.
- 26 of 26 plausible wrong answers failed. Each task had 2 to 5 of them, written to look right at a glance.
- 8 of 8 reference answers wrapped in a code fence or prose were flagged as format misses, not passes.
The controls table is on the study page.
Pass rate: four perfect scores and one clear gap
Four configurations passed every call. With 24 calls each, a perfect score has a 95% interval of 86% to 100%. So the honest reading is: the hard set still has a ceiling for Sonnet, Opus and Fable. We cannot tell from this run which of the four is strongest. We need harder tasks for that.
Haiku is different. Its strict interval (28% to 65%) does not overlap the 86% to 100% of the other four. That is the only pass-rate difference in this study that the intervals support. On the comparison pages, it is the row where Sonnet, Opus and Fable each beat Haiku: Haiku vs Sonnet, Haiku vs Opus and Haiku vs Fable.
Format misses vs wrong answers
Every prompt states the output format: no code fence, no other text. When a reply fails the strict check, a lenient extractor looks inside it (a fenced block, an outer JSON value, the first code line). If that extracted answer passes the same validator, we call the reply a format miss. A format miss is never a pass.
All 13 non-passes came from Haiku:
- 5 format misses. The answer was right, but wrapped in a fence or prose.
- 8 wrong answers.
On the lenient reading, Haiku reached 16 of 24 (67%, 95% interval 47% to 82%). That interval still does not overlap the other four. So Haiku's gap is not only a format habit.
Per task, Haiku passed the interval merge and the SemVer regex 3 of 3. It passed none of the event-loop order, the room schedule and the SQL query. On the room schedule and the SQL query, 2 of its 3 replies each were format misses. On the event-loop order, all 3 were wrong.
On the easy set, Sonnet's only misses were format misses. Here, where every prompt stated the format rule, it passed 24 of 24. The tasks differ too, so this data cannot say the clearer rule caused that.
Speed: Sonnet fastest, Haiku slowest
Median total time per call: Sonnet 7.7 s, Opus 9.2 s, Opus high 11.0 s, Fable 16.1 s, Haiku 39.0 s. The whiskers show the fastest and slowest single call: Sonnet ran from 2.26 s to 34.79 s, Haiku from 15.27 s to 75.13 s. A range is not a confidence interval, and every range overlaps every other. So these medians describe this run; they are not a tested ranking.
Two results reverse the easy set. There, Fable was fastest at a 1.9 s median. On hard tasks it was the slowest of the four perfect configurations. And Haiku, the "small, fast" model, was the slowest Claude configuration on both sets, here by far.
Time to first useful output tells the same story: Sonnet 5.95 s, Opus 6.78 s, Opus high 7.13 s, Fable 11.63 s and Haiku 35.54 s.
Tokens: why Haiku is slow
Median output tokens per call: Opus 945, Sonnet 1,050, Opus high 1,052, Fable 1,366 and Haiku 5,064, of which 4,556 were reported reasoning tokens. At the CLI default effort, Haiku wrote about four to five times as much as the others and still got 8 answers wrong. That explains most of its time.
Cost per strict pass (a calculation)
The calls ran on a flat subscription, so there is no bill. We priced every reported token at list price, summed over all 24 calls in each configuration, and divided by the strict passes. A failed call still costs.
- Claude Sonnet 5.5: $0.0143
- Claude Opus 5.5: $0.0282
- Claude Opus 5.5 (high): $0.0334
- Claude Haiku 4.5: $0.0672
- Claude Fable 5.1: $0.0933
Haiku has the lowest list price per token, yet it costs more per pass than Sonnet or Opus. Two effects add up: far more output tokens, and fewer than half the calls passing. The quality-vs-cost frontier holds one point: Sonnet. No other configuration passes as often for less.
These gaps have no interval; read them as a direction.
Where is Codex?
We planned 30 calls for GPT-6.1 Sol in the Codex CLI. All 30 were blocked before any model call. The CLI reported no signed-in account, and our runner has no API-key fallback by design. So there are no Codex timings, tokens or answers to score.
We count those attempts as "blocked, not scored" (30 of 150), and we did not fill the gap from another run. This study makes no Claude-vs-Codex claim. For the Codex data we do have, see Claude Code vs Codex CLI: the hidden context tax and the Claude Code vs Codex CLI comparison page.
What this means if you pick a model
- For work like this, Sonnet 5.5 is the default to beat. It passed everything, had the lowest median time and the lowest cost per pass.
- Opus and Fable showed no extra quality here, because Sonnet also passed everything. These tasks do not find the difference. Price side: Sonnet vs Opus: when is Opus worth it?.
- Do not pick Haiku 4.5 at its CLI default for hard reasoning. It was slower, less accurate and more expensive per pass.
- State the output format, and validate it. A pipeline needs to know a format miss from a wrong answer.
Each model's full record across all our studies: Sonnet 5.5, Opus 5.5, Haiku 4.5 and Fable 5.1.
How we measured
- Protocol declared before the first call. Claude Code, one turn, tools off, fresh empty folder, no MCP servers, a 300 s timeout, one call at a time. Three repetitions per task, so n = 24 per configuration.
- Attempts: 120 Claude Code calls planned and run; 30 Codex CLI calls planned, 0 reached a model. Nothing was trimmed or retried.
Caveats
- Ceiling. Four configurations scored 24/24, so pass rate cannot rank them.
- Small cells. Each task cell has 3 calls.
- CLI included. Timings include CLI start-up and its system prompt. One host, one network, one session. Default effort means the CLI chose.
- Costs are calculations, not bills. No Codex data.
What to read next
- Claude Haiku vs Sonnet vs Opus vs Fable vs Codex: the five-task head-to-head
- Sonnet vs Opus: when is Opus worth the price?
- Benchmarks roundup, October 2026
See your own pass rates
Agent validates every step it delivers and records which model did the work, how long it took and what it cost. Try Agent and see the receipts for your own tasks.