• Latency
  • Head to head
  • Claude Haiku
  • Claude Sonnet

Which Claude model is fastest? It depends on the task, and on thinking

1.94 s was the lowest median on short calls (Fable 5.1). 7.75 s on hard calls (Sonnet 5.5). Haiku 4.5 took 4.43 s and 39.01 s with default thinking.

TL;DR

  • The lowest median changes with the task. On short calls it was Claude Fable 5.1 at 1.94 s (fastest to slowest call: 1.41 to 9.83 s, n = 15). On hard calls it was Claude Sonnet 5.5 at 7.75 s (2.26 to 34.79 s, n = 24).
  • Haiku 4.5 had the highest median of the Claude cells on both sets: 4.43 s and 39.01 s. On hard tasks it reported a median 4,556 reasoning tokens under the CLI default. Thinking off was not tested.
  • Most speed rows stay unclear. A range is not a confidence interval, and most ranges overlap. The Haiku vs Sonnet page has 22 time rows (our count). The data decides 6, all for Sonnet. Five of those 6 come from one routing run.
  • Low effort had the lowest median on Sonnet. Its 5.82 s was the lowest of 11 effort cells. The ranges overlap.
  • Every timing is through the Claude Code CLI on one host, with start-up included.
Live story · 35 sHaiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

127 of 130 calls passed, so speed and tokens separate the models: Fable 5.1 was fastest at 1.9 s median.

Transcript
  1. Head-to-head · 130 timed calls · 9 configurations. Haiku vs Sonnet vs Opus vs Fable vs Codex. Five short tasks with strict validators. Every call kept, nothing retried.
  2. 127 of 130 calls passed. Pass rate barely separates them; speed and tokens do. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Cheapest passing answer (list-price calculation): Sonnet 5.5: $0.0062 (n = 15). Caveat: The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.
  3. Fable 5.1 finishes first at 1.9 s. The Codex CLI needs 5.6–6.3 s. Chart: Median total time per call · real time (n = 10–15 each). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
  4. What the CLI sends: a median 12,124 input tokens per call on the Codex CLI, 2,130 on Claude Code. Chart: Input tokens per call: what the CLI sends (n = 10–15 each). Caveat: The prompt cache stayed at the provider default, so cache counters differ by route and by call order.
  5. Speed vs cost per passing answer. Ringed: no other setup is faster, cheaper per pass and as accurate. Chart: Speed, cost and quality frontier (n = 10–15 each). Calculation, not a run. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  6. Open benchmarks: intervals, sources and every failure kept.

The short answer

The film shows the five short tasks. In the film and in this post, "fastest" means "lowest median in this run". It is not a tested ranking.

TaskLowest medianHaiku 4.5 mediann per cell
Five short tasksFable: 1.94 s (1.41 to 9.83)4.43 s (3.16 to 23.57)15
Eight hard tasksSonnet: 7.75 s (2.26 to 34.79)39.01 s (15.27 to 75.13)24
Routing decisionsSonnet, low effort: 2.60 s (p50 to p95: 2.60 to 4.30)12.67 s (p50 to p95: 12.67 to 34.41)82
Code fix, same prompt 10 timesSonnet: 2.67 s (2.32 to 4.34)5.95 s (4.89 to 7.33)10

The data decides only the last two rows, and only for Sonnet against Haiku. In the first two rows the ranges overlap. Haiku ran with the CLI default thinking in every row.

Short calls: the medians sit close together

Entrance: medians race at 4.5× real timeMotion reduced: press Replay to animateThe slowest median is 6.3 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

9 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 6.3 s (range 4.7 s–10.5 s, n 10). Fastest Claude Fable 5.1 · Claude Code 1.9 s (range 1.4 s–9.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

The five-task study ran each Claude cell 15 times. Medians, lowest first: Fable 5.1 1.94 s and Sonnet 5.5 2.31 s. Opus 5.5 ran 2.71 s (high), 2.75 s (default) and 2.83 s (low). Haiku 4.5 ran 4.43 s.

Fable and Sonnet differ by 0.37 s (calculation). Their ranges overlap, so the Sonnet vs Fable page marks the row unclear.

Time to first useful output: Fable 1.20 s (0.95 to 7.90) and Sonnet 1.56 s (0.99 to 6.39). Haiku took 3.63 s (2.78 to 22.27).

Hard calls: the order changes

Entrance: medians race at 28× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

7 rows. Slowest Claude Haiku 4.5 · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 16–24 per row

Median per configuration; whiskers = fastest and slowest call

One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

The hard-task study has eight tasks with strict validators and 24 calls per Claude cell. Medians: Sonnet 7.75 s, Opus 9.18 s (4.24 to 27.21) and Opus at high effort 11.03 s (3.63 to 63). Fable took 16.13 s (4.46 to 90) and Haiku 39.01 s. Fable ranked first of six Claude cells on short calls and fourth of five here.

Here the slowest cell also passed least. Sonnet, Opus, Opus (high) and Fable passed 24 of 24 (95% interval 86% to 100%). Haiku passed 11 of 24 (46%, 28% to 65%).

Where the data does decide

  • Wall time (CLI)
  • Wall time (direct API call)
  • Model time (API)
Claude Haiku 4.5
Claude Sonnet 5.5
Jev 1.13 (TypeSafe)

Time per decision · log scale: each gridline is 10 times the one before

3 rows, 3 series: Wall time (CLI), Wall time (direct API call), Model time (API). Wall time (CLI): slowest Claude Haiku 4.5 12.67 s (median to p95 12.67 s–34.41 s, n 82). Fastest Claude Sonnet 5.5 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–246 per row

Median wall time, whisker to the 95th percentile

Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

The data decides a row when the ranges or p50 to p95 bands do not overlap. A range is not a confidence interval. The Haiku vs Sonnet page has 22 time rows (our count). Six meet that test, all for Sonnet. The other 16 are unclear: 6 call-timing rows, 8 agent-memory session rows and 2 calculation rows.

  • Routing decision, wall time: Sonnet 2.60 s against Haiku 12.67 s, n = 82 each. Model time was 1.60 s against 10.73 s. Five of the six rows come from this one run, reported in the routing study and the routing overhead study.
  • Code fix, same prompt 10 times: Sonnet 2.67 s against Haiku 5.95 s. The ranges do not overlap.

The routing run mixes settings: Sonnet ran at low effort and Haiku with the CLI default thinking.

The consistency study has two more prompts, and the data decides neither. On the JSON prompt Sonnet had 2.89 s (2.68 to 5.30) and Haiku 7.03 s (5.28 to 12.27). The ranges overlap by 0.02 s (calculation). On the exact-number prompt Haiku had the lower median, 5.06 s against 6.89 s, with overlapping ranges. Haiku also gave the same wrong number in 10 of 10 calls.

The other five Claude pairs have 23 time rows (our count). The data decides none. Sonnet vs Opus has 7, Sonnet vs Fable has 4.

What slows Haiku? Thinking tokens fit the data

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).

Notesn 16–24 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Under the CLI default, Haiku reported a median 4,556 reasoning tokens per hard call. Sonnet reported 585, Opus 529 and Fable 889. Median output was 5,064 tokens for Haiku and 1,050 for Sonnet, 4.8 times (calculation). Median time was 5.0 times Sonnet's (calculation).

Haiku's median time to first useful output was 35.54 s (12.88 to 70.31). Sonnet's was 5.95 s (0.86 to 30.57). The gap to the median total is 3.47 s for Haiku and 1.80 s for Sonnet (calculations). Most of the wait comes before the answer starts.

That fits the idea that thinking costs time. It is not a test: these studies did not run Haiku with thinking off. For the terms, read what reasoning effort is.

Effort: Sonnet's lowest median came at low effort

Entrance: medians race at 13× real timeMotion reduced: press Replay to animateThe slowest median is 18.1 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

11 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 18.1 s (range 11.7 s–92.2 s, n 16). Fastest Claude Sonnet 5.5 (low) · Claude Code 5.8 s (range 2.8 s–20 s, n 16). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 16 per row

Median per configuration; whiskers = fastest and slowest call

Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.

Source: Effort ladder: the hard task set at each effort level

The effort ladder reran the same eight hard tasks, 16 calls per cell. Sonnet medians: low 5.82 s (2.78 to 19.96), medium 7.63 s (2.71 to 24.01), high 8.81 s (2.93 to 35.81). Median output rose with effort: 667, 770 and 1,192 tokens.

Opus ran in the same order on hard tasks: low 7.50 s, medium 9.72 s, high 10.11 s. On short calls it did not: low 2.83 s, high 2.71 s.

All 11 cells passed 16 of 16, so lower effort lost no passes on this set. The set has a ceiling and cannot rule out a gap of about 19 points. The ranges overlap.

How to get speed (our reading, not tested)

  1. Time your own tasks and pick by your own p95, not the median. Fable's slowest hard call took 90 s; Sonnet's took 34.79 s.
  2. Try low effort for latency-critical steps. On our set, Sonnet at low effort lost no passes and had the lowest median. The ranges overlap.
  3. Skip default thinking on Haiku for those steps. Measure thinking off first. These studies did not. Speed is not work done: Haiku passed 11 of 24 hard calls.

How we measured

  • Calls. Short: 3 repetitions of 5 tasks. Hard: 3 of 8. Ladder: 2 of 8. Consistency: 10 per prompt. Routing: 82 decisions.
  • Statistics. The median, with the fastest and slowest call. Routing shows p50 to p95.
  • First useful output. The first streamed text that belongs to the answer. Total time runs from launch to exit, including CLI start-up.
  • Isolation. Fresh empty folder, tools off, no MCP servers, no session persistence, one turn, one call at a time. We retried nothing.
  • Ladder cells. The ladder reuses hard-study cells but keeps repetitions 1 and 2. Sonnet at default effort reads 7.97 s there and 7.75 s in the hard study (all 3). The new cells ran in a different hour.

Caveats

  • One host, one network, one session. Provider load can move. These are CLI plus model times, not raw API times.
  • Haiku ran with default thinking. The overhead study recomputed its routing median as 12.54 s; the routing study reports 12.67 s.
  • Small samples. 10 to 24 calls per cell. A few slow calls move a median.
  • Ceilings. The short tasks passed 127 of 130 calls across all 9 cells, Codex included (3 format misses, all Sonnet). Six of seven hard cells passed every call.

Time your own calls

Agent records the model, the route, the tokens and the time of every call. Try Agent and see which model is fastest on your work.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.