Claude Haiku vs Sonnet vs Opus vs Fable vs Codex: 130 timed calls, head to head
130 timed calls on five validated tasks. Pass rate, latency, hidden input tokens and cost per passing answer for Claude Code models and Codex CLI.
Update, 2026-10-06: a declared top-up added 10 Codex CLI calls (rep 3 at medium and high effort). All 10 passed, so the study holds 127 of 130 passes, and Codex CLI medians run from 5.6 s to 6.3 s. The figures below include the top-up and match the study page. Hard-task Codex results: Claude vs Codex on hard tasks.
TL;DR
- 127 of 130 calls passed their validator (98%, 95% interval 93% to 99%). On short tasks, pass rate barely separates the models.
- The only non-passes were Claude Sonnet 5.5 on one arithmetic task, 3 of 3. Each time it reached the right answer, 2292, but added working lines that the exact-text validator rejects. A format miss, not a wrong answer.
- Speed separates them. Claude Fable 5.1 in Claude Code had the fastest median at 1.9 s per call. GPT-6.1 Sol in the Codex CLI took 5.6 to 6.3 s.
- The Codex CLI sent a median 12,124 input tokens per call. Claude Code sent 2,130. The prompts themselves are a few hundred tokens.
- At list price (a calculation), the cheapest passing answer came from Claude Sonnet 5.5 at $0.0062.
Every call and every validator result: /benchmarks/model-head-to-head.
The setup in one paragraph
We declared a protocol before the first call. Five short cases, each with a deterministic validator: behavioral checks in a sandbox, canonical JSON, or exact text. Nine configurations. Claude Code ran Haiku 4.5, Sonnet 5.5, Opus 5.5 and Fable 5.1 at default effort, plus Opus 5.5 at low and high effort, three repetitions per task. Codex CLI ran GPT-6.1 Sol at low, medium and high effort, two repetitions per task, plus a declared third repetition at medium and high effort. Each call got a fresh empty folder, tools off, no MCP servers, no session persistence and one turn. All 130 planned calls ran: 120 in the first plan and 10 in the declared Codex top-up. Nothing was trimmed or retried, and no run hit a usage or rate limit.
Pass rate: a near tie
Eight of nine configurations passed every call. Sonnet 5.5 passed 12 of 15. Its three misses were all on the same task, multi-step shift arithmetic, and all three ended on the right number with extra working lines.
Is that a Sonnet weakness? On this evidence, it is a format habit on one prompt. The validator wants the exact text, and we keep that rule because format compliance matters in real pipelines. But we would not read it as a reasoning gap. With 15 calls per configuration, the intervals for "15 of 15" still reach down to about 80%. To find where pass rates do split, we ran a harder follow-up: When tasks get hard.
Speed: the real signal
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
127 of 130 calls passed, so speed and tokens separate the models: Fable 5.1 was fastest at 1.9 s median.
Transcript
- Head-to-head · 130 timed calls · 9 configurations. Haiku vs Sonnet vs Opus vs Fable vs Codex. Five short tasks with strict validators. Every call kept, nothing retried.
- 127 of 130 calls passed. Pass rate barely separates them; speed and tokens do. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Cheapest passing answer (list-price calculation): Sonnet 5.5: $0.0062 (n = 15). Caveat: The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.
- Fable 5.1 finishes first at 1.9 s. The Codex CLI needs 5.6–6.3 s. Chart: Median total time per call · real time (n = 10–15 each). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
- What the CLI sends: a median 12,124 input tokens per call on the Codex CLI, 2,130 on Claude Code. Chart: Input tokens per call: what the CLI sends (n = 10–15 each). Caveat: The prompt cache stayed at the provider default, so cache counters differ by route and by call order.
- Speed vs cost per passing answer. Ringed: no other setup is faster, cheaper per pass and as accurate. Chart: Speed, cost and quality frontier (n = 10–15 each). Calculation, not a run. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
- Open benchmarks: intervals, sources and every failure kept.
Median total time per call, fastest first:
- Claude Fable 5.1 · Claude Code: 1.94 s
- Claude Sonnet 5.5 · Claude Code: 2.31 s
- Claude Opus 5.5 (high) · Claude Code: 2.71 s
- Claude Opus 5.5 · Claude Code: 2.75 s
- Claude Opus 5.5 (low) · Claude Code: 2.83 s
- Claude Haiku 4.5 · Claude Code: 4.43 s
- GPT-6.1 Sol (high) · Codex CLI: 5.60 s
- GPT-6.1 Sol (medium) · Codex CLI: 5.65 s
- GPT-6.1 Sol (low) · Codex CLI: 6.26 s
Two surprises. First, Haiku, the "small fast model", was the slowest Claude configuration. It reported reasoning tokens on most calls under the CLI default, and that thinking takes time. Second, effort barely moved Opus or Sol on tasks this short. Opus low, default and high all landed between 2.71 s and 2.83 s.
Look at the whiskers before you pick a winner. Fable's slowest call took 9.83 s. Haiku's slowest took 23.57 s. A median from 10 to 15 calls is a useful guide, but neighbouring configurations overlap.
The hidden prompt: what a CLI sends
A coding CLI does not send only your prompt. It wraps it in a system prompt and tool context. That costs tokens and time.
The mean input per call, split into prompt-cache reads and other input:
- Claude Code with Sonnet or Opus: 1,401 to 1,463 cache-read tokens plus 618 to 685 other input.
- Claude Code with Fable: 2,760 cache read plus 473 other.
- Claude Code with Haiku: 0 cache read plus 3,790 other.
- Codex CLI: 5,180 to 8,064 cache read plus 4,059 to 6,943 other.
The median across all calls was 12,124 input tokens for Codex CLI and 2,130 for Claude Code. That is mostly the CLI's own context, not the task. It may add to the Codex CLI's longer times; this study does not separate the two. We look at that overhead directly, against the bare API, in Claude Code vs Codex CLI vs the API.
Cost per passing answer
The calls ran on flat subscriptions, so there is no bill. We can still ask what each configuration would cost at list price. We price every reported token, cache reads and writes included, sum over every call in the configuration, and divide by its passes. A failed call still costs, so a lower pass rate raises the cost per pass.
- Claude Sonnet 5.5: $0.00624, the lowest, even with its three misses.
- Claude Opus 5.5 (low): $0.00829.
- Claude Haiku 4.5: $0.00836. Its extra reasoning output cancels its lower list price.
- GPT-6.1 Sol (low) · Codex CLI: $0.00998.
- Claude Opus 5.5: $0.01009.
- Claude Opus 5.5 (high): $0.01049.
- GPT-6.1 Sol (high): $0.01322.
- GPT-6.1 Sol (medium): $0.01564.
- Claude Fable 5.1: $0.02054, the highest.
Fable is fastest and most expensive. Sonnet is cheapest per pass and second fastest. Haiku is not the budget pick here, because of how much it thinks under the CLI default.
The frontier
Which configurations are not beaten on all three axes at once: speed, cost per pass and pass rate?
The frontier holds five Claude Code configurations: Fable 5.1, Sonnet 5.5, and Opus 5.5 at high, default and low effort. Haiku and the three Codex CLI configurations are dominated: some other configuration is at least as fast, as cheap per pass and as accurate.
Read the frontier as a shortlist, not a podium. With 10 to 15 calls per configuration, the gaps inside it are within the spread of single calls. Pick by what you care about most: Fable for raw speed, Sonnet for cost per pass, Opus when you want headroom on harder work.
How we measured
- Protocol declared before the first call, five cases with deterministic validators.
- Repetitions. Three per task for Claude Code configurations (15 calls each). Two per task for Codex at low effort (10 calls), three at medium and high effort (15 calls each, after the declared top-up).
- Isolation. Fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, one call at a time per account.
- Timing. Total time includes CLI start-up. Time to first useful output is the first streamed text that belongs to the answer.
- Cost. List price × reported tokens per call, cache reads and writes priced separately. A calculation, not a bill.
- Kept. Every attempt. Nothing retried.
Caveats
- Easy tasks. Pass rate saturates. Latency and tokens carry the signal. Do not read this as a reasoning benchmark.
- Few repetitions. Medians with ranges, not confidence intervals.
- CLI overhead is included. These are CLI-plus-model results, not raw model results.
- One host, one network, one day. Vendor latency moves over the day.
- Cache state. The prompt cache stayed at the provider default, so cache counters differ by route and call order.
- Haiku reasoning. Haiku 4.5 reported reasoning tokens on most calls under the CLI default. That explains much of its extra time and output.
- Catalog. Haiku and Fable are not in our runner catalog. These calls used the runner's CLI functions directly, and each receipt records the model the CLI reported.
What to read next
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable on 8 hard tasks
- Claude Code vs Codex CLI vs the API: latency and hidden tokens
- What if every call ran on Opus? A thought experiment
- Jev vs Claude Haiku and Sonnet as a router
Use the right model for each step
No single model wins every axis. Agent records speed, tokens and cost for every call, so you can see which model fits which step of your work. Try Agent and compare the receipts yourself.