Which Claude model is fastest? It depends on the task, and on thinking
1.94 s was the lowest median on short calls (Fable 5.1). 7.75 s on hard calls (Sonnet 5.5). Haiku 4.5 took 4.43 s and 39.01 s with default thinking.
TL;DR
- The lowest median changes with the task. On short calls it was Claude Fable 5.1 at 1.94 s (fastest to slowest call: 1.41 to 9.83 s, n = 15). On hard calls it was Claude Sonnet 5.5 at 7.75 s (2.26 to 34.79 s, n = 24).
- Haiku 4.5 had the highest median of the Claude cells on both sets: 4.43 s and 39.01 s. On hard tasks it reported a median 4,556 reasoning tokens under the CLI default. Thinking off was not tested.
- Most speed rows stay unclear. A range is not a confidence interval, and most ranges overlap. The Haiku vs Sonnet page has 22 time rows (our count). The data decides 6, all for Sonnet. Five of those 6 come from one routing run.
- Low effort had the lowest median on Sonnet. Its 5.82 s was the lowest of 11 effort cells. The ranges overlap.
- Every timing is through the Claude Code CLI on one host, with start-up included.
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
127 of 130 calls passed, so speed and tokens separate the models: Fable 5.1 was fastest at 1.9 s median.
Transcript
- Head-to-head · 130 timed calls · 9 configurations. Haiku vs Sonnet vs Opus vs Fable vs Codex. Five short tasks with strict validators. Every call kept, nothing retried.
- 127 of 130 calls passed. Pass rate barely separates them; speed and tokens do. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Cheapest passing answer (list-price calculation): Sonnet 5.5: $0.0062 (n = 15). Caveat: The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.
- Fable 5.1 finishes first at 1.9 s. The Codex CLI needs 5.6–6.3 s. Chart: Median total time per call · real time (n = 10–15 each). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
- What the CLI sends: a median 12,124 input tokens per call on the Codex CLI, 2,130 on Claude Code. Chart: Input tokens per call: what the CLI sends (n = 10–15 each). Caveat: The prompt cache stayed at the provider default, so cache counters differ by route and by call order.
- Speed vs cost per passing answer. Ringed: no other setup is faster, cheaper per pass and as accurate. Chart: Speed, cost and quality frontier (n = 10–15 each). Calculation, not a run. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
- Open benchmarks: intervals, sources and every failure kept.
The short answer
The film shows the five short tasks. In the film and in this post, "fastest" means "lowest median in this run". It is not a tested ranking.
| Task | Lowest median | Haiku 4.5 median | n per cell |
|---|---|---|---|
| Five short tasks | Fable: 1.94 s (1.41 to 9.83) | 4.43 s (3.16 to 23.57) | 15 |
| Eight hard tasks | Sonnet: 7.75 s (2.26 to 34.79) | 39.01 s (15.27 to 75.13) | 24 |
| Routing decisions | Sonnet, low effort: 2.60 s (p50 to p95: 2.60 to 4.30) | 12.67 s (p50 to p95: 12.67 to 34.41) | 82 |
| Code fix, same prompt 10 times | Sonnet: 2.67 s (2.32 to 4.34) | 5.95 s (4.89 to 7.33) | 10 |
The data decides only the last two rows, and only for Sonnet against Haiku. In the first two rows the ranges overlap. Haiku ran with the CLI default thinking in every row.
Short calls: the medians sit close together
The five-task study ran each Claude cell 15 times. Medians, lowest first: Fable 5.1 1.94 s and Sonnet 5.5 2.31 s. Opus 5.5 ran 2.71 s (high), 2.75 s (default) and 2.83 s (low). Haiku 4.5 ran 4.43 s.
Fable and Sonnet differ by 0.37 s (calculation). Their ranges overlap, so the Sonnet vs Fable page marks the row unclear.
Time to first useful output: Fable 1.20 s (0.95 to 7.90) and Sonnet 1.56 s (0.99 to 6.39). Haiku took 3.63 s (2.78 to 22.27).
Hard calls: the order changes
The hard-task study has eight tasks with strict validators and 24 calls per Claude cell. Medians: Sonnet 7.75 s, Opus 9.18 s (4.24 to 27.21) and Opus at high effort 11.03 s (3.63 to 63). Fable took 16.13 s (4.46 to 90) and Haiku 39.01 s. Fable ranked first of six Claude cells on short calls and fourth of five here.
Here the slowest cell also passed least. Sonnet, Opus, Opus (high) and Fable passed 24 of 24 (95% interval 86% to 100%). Haiku passed 11 of 24 (46%, 28% to 65%).
Where the data does decide
The data decides a row when the ranges or p50 to p95 bands do not overlap. A range is not a confidence interval. The Haiku vs Sonnet page has 22 time rows (our count). Six meet that test, all for Sonnet. The other 16 are unclear: 6 call-timing rows, 8 agent-memory session rows and 2 calculation rows.
- Routing decision, wall time: Sonnet 2.60 s against Haiku 12.67 s, n = 82 each. Model time was 1.60 s against 10.73 s. Five of the six rows come from this one run, reported in the routing study and the routing overhead study.
- Code fix, same prompt 10 times: Sonnet 2.67 s against Haiku 5.95 s. The ranges do not overlap.
The routing run mixes settings: Sonnet ran at low effort and Haiku with the CLI default thinking.
The consistency study has two more prompts, and the data decides neither. On the JSON prompt Sonnet had 2.89 s (2.68 to 5.30) and Haiku 7.03 s (5.28 to 12.27). The ranges overlap by 0.02 s (calculation). On the exact-number prompt Haiku had the lower median, 5.06 s against 6.89 s, with overlapping ranges. Haiku also gave the same wrong number in 10 of 10 calls.
The other five Claude pairs have 23 time rows (our count). The data decides none. Sonnet vs Opus has 7, Sonnet vs Fable has 4.
What slows Haiku? Thinking tokens fit the data
Under the CLI default, Haiku reported a median 4,556 reasoning tokens per hard call. Sonnet reported 585, Opus 529 and Fable 889. Median output was 5,064 tokens for Haiku and 1,050 for Sonnet, 4.8 times (calculation). Median time was 5.0 times Sonnet's (calculation).
Haiku's median time to first useful output was 35.54 s (12.88 to 70.31). Sonnet's was 5.95 s (0.86 to 30.57). The gap to the median total is 3.47 s for Haiku and 1.80 s for Sonnet (calculations). Most of the wait comes before the answer starts.
That fits the idea that thinking costs time. It is not a test: these studies did not run Haiku with thinking off. For the terms, read what reasoning effort is.
Effort: Sonnet's lowest median came at low effort
The effort ladder reran the same eight hard tasks, 16 calls per cell. Sonnet medians: low 5.82 s (2.78 to 19.96), medium 7.63 s (2.71 to 24.01), high 8.81 s (2.93 to 35.81). Median output rose with effort: 667, 770 and 1,192 tokens.
Opus ran in the same order on hard tasks: low 7.50 s, medium 9.72 s, high 10.11 s. On short calls it did not: low 2.83 s, high 2.71 s.
All 11 cells passed 16 of 16, so lower effort lost no passes on this set. The set has a ceiling and cannot rule out a gap of about 19 points. The ranges overlap.
How to get speed (our reading, not tested)
- Time your own tasks and pick by your own p95, not the median. Fable's slowest hard call took 90 s; Sonnet's took 34.79 s.
- Try low effort for latency-critical steps. On our set, Sonnet at low effort lost no passes and had the lowest median. The ranges overlap.
- Skip default thinking on Haiku for those steps. Measure thinking off first. These studies did not. Speed is not work done: Haiku passed 11 of 24 hard calls.
How we measured
- Calls. Short: 3 repetitions of 5 tasks. Hard: 3 of 8. Ladder: 2 of 8. Consistency: 10 per prompt. Routing: 82 decisions.
- Statistics. The median, with the fastest and slowest call. Routing shows p50 to p95.
- First useful output. The first streamed text that belongs to the answer. Total time runs from launch to exit, including CLI start-up.
- Isolation. Fresh empty folder, tools off, no MCP servers, no session persistence, one turn, one call at a time. We retried nothing.
- Ladder cells. The ladder reuses hard-study cells but keeps repetitions 1 and 2. Sonnet at default effort reads 7.97 s there and 7.75 s in the hard study (all 3). The new cells ran in a different hour.
Caveats
- One host, one network, one session. Provider load can move. These are CLI plus model times, not raw API times.
- Haiku ran with default thinking. The overhead study recomputed its routing median as 12.54 s; the routing study reports 12.67 s.
- Small samples. 10 to 24 calls per cell. A few slow calls move a median.
- Ceilings. The short tasks passed 127 of 130 calls across all 9 cells, Codex included (3 format misses, all Sonnet). Six of seven hard cells passed every call.
What to read next
- Claude Haiku vs Sonnet vs Opus vs Fable vs Codex, head to head
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable
- Does reasoning effort buy quality?
Time your own calls
Agent records the model, the route, the tokens and the time of every call. Try Agent and see which model is fastest on your work.