113 measured metrics · 40 calculated · 14 studies

Claude Haiku 4.5vsClaude Sonnet 5.5

Claude Sonnet 5.5 ahead on 21; 58 ties, 74 unclear. A side is ahead only where the intervals or ranges do not overlap.

The verdict

Claude Haiku 4.5 and Claude Sonnet 5.5 share 113 measured metrics and 40 list-price calculations from 14 studies. Claude Sonnet 5.5 leads on 21 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); and 18 more. On those rows the 95% intervals, run ranges and p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 58 ties and 74 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 2 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Input tokens per call: what the CLI sends (Other input), 5.5x (Claude Haiku 4.5 larger).

Watch it build

A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.

Live story · 43 sClaude Haiku 4.5 vs Claude Sonnet 5.5: what the measurements say

Claude Haiku 4.5 vs Claude Sonnet 5.5: what the measurements say

153 comparison rows from 14 studies: 0 rows favour Haiku 4.5, 21 favour Sonnet 5.5, 132 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 153 rows · 14 studies. Haiku 4.5 vs Sonnet 5.5. A winner only where the 95% intervals or run ranges do not overlap.
  2. 153 comparison rows from 14 studies: Haiku 4.5 ahead on 0, Sonnet 5.5 ahead on 21. The rest do not separate them. Rows where Haiku 4.5 is ahead: 0 (of 153). Rows where Sonnet 5.5 is ahead: 21 (of 153). Ties or unclear: 132 (58 ties · 74 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. Eight hard tasks: pass rate 46% vs 100%, Sonnet 5.5 ahead: 95% intervals separate. 2 of 6 rows separate them. Table: Eight hard tasks · 5 of 6 rows · Claude Code · n = 24 per side. Source study: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks. Rows shown: Pass rate on eight hard tasks (Strict pass); Pass rate on eight hard tasks (Lenient (format misses counted)); Total time per call on hard tasks (separate batches); Time to first useful output on hard tasks; Output tokens per call on hard tasks (Output tokens). Recorded settings: Claude Code · eight hard validated tasks. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  4. Caching sessions: pass rate 0% vs 100%, Sonnet 5.5 ahead: 95% intervals separate. 3 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Code · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: how many different answers (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Code · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  5. Agent memory: pass rate 20% vs 60%, tie: 95% intervals overlap. 4 of 40 rows separate them. Table: Agent memory · 5 of 40 rows · n = 10 vs 15. Source study: Does memory help Claude Code? 8 kinds of agent memory, tested. Rows shown: Full pass rate by kind of memory: No memory; Full pass rate by kind of memory: /init CLAUDE.md; Full pass rate by kind of memory: Handbook, 210 lines; Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md; Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines. Caveat: The hook checks the same code rules as the grader. It shows what rules written as code can do; it cannot carry a fact such as the late-fee rate.
  6. Routing overhead: success rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 3 of 8 rows separate them. Table: Routing overhead · 5 of 8 rows · thinking on vs effort low · n = 82 per side. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: Time to make one routing decision; Where an LLM router’s time goes: model vs CLI (Model API time); Where an LLM router’s time goes: model vs CLI (CLI and harness time); Routing calls that returned a decision; Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)). Recorded settings: thinking on · via Claude Code · routing overhead per decision; effort low · via Claude Code · routing overhead per decision; thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts; effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts. Includes a calculation, not a bill or a new run. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  7. No winner where the data shows none. Showing 20 of 153 rows; every row and its reason online.

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Haiku 4.5
  • Claude Sonnet 5.5
  • 95% interval
  • fastest–slowest run (not an interval)
  • median to p95 (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

1 tie · 7 unclear
  • Pass rate on five validated tasks: Claude Haiku 4.5 100% (15/15) (n 15, 95% interval 80%–100%); Claude Sonnet 5.5 80% (12/15) (n 15, 95% interval 55%–93%). Tie.

    Calculation: at these rates, about 33 runs per side would separate them.

  • Total time per call: Claude Haiku 4.5 4.43 s (n 15, run range 3.2 s–23.6 s); Claude Sonnet 5.5 2.31 s (n 15, run range 2.2 s–7.7 s). Unclear.
  • Time to first useful output: Claude Haiku 4.5 3.63 s (n 15, run range 2.8 s–22.3 s); Claude Sonnet 5.5 1.56 s (n 15, run range 1 s–6.4 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): Claude Haiku 4.5 0 (n 15); Claude Sonnet 5.5 1,401 (n 15). Unclear.
  • Input tokens per call: what the CLI sends (Other input): Claude Haiku 4.5 3,790 (n 15); Claude Sonnet 5.5 685 (n 15). Unclear.
  • Output tokens per call (Output tokens): Claude Haiku 4.5 367 (n 15); Claude Sonnet 5.5 107 (n 15). Unclear.
  • List-price cost per call (calculation), calculation: Claude Haiku 4.5 $0.0057 (n 15, run range $0.0051–$0.018); Claude Sonnet 5.5 $0.0036 (n 15, run range $0.0034–$0.01). Unclear.
  • List-price cost per passing answer (calculation), calculation: Claude Haiku 4.5 $0.0084 (n 15); Claude Sonnet 5.5 $0.0062 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

Claude Sonnet 5.5 2 · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): Claude Haiku 4.5 46% (11/24) (n 24, 95% interval 28%–65%); Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Claude Sonnet 5.5 ahead.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): Claude Haiku 4.5 67% (16/24) (n 24, 95% interval 47%–82%); Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Claude Sonnet 5.5 ahead.
  • Total time per call on hard tasks (separate batches): Claude Haiku 4.5 39.0 s (n 24, run range 15.3 s–75.1 s); Claude Sonnet 5.5 7.75 s (n 24, run range 2.3 s–34.8 s). Unclear.
  • Time to first useful output on hard tasks: Claude Haiku 4.5 35.5 s (n 24, run range 12.9 s–70.3 s); Claude Sonnet 5.5 5.95 s (n 24, run range 0.9 s–30.6 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): Claude Haiku 4.5 5,064 (n 24); Claude Sonnet 5.5 1,050 (n 24). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: Claude Haiku 4.5 $0.067 (n 24); Claude Sonnet 5.5 $0.014 (n 24). Unclear.

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

Claude Sonnet 5.5 3 · 3 ties · 3 unclear
  • Same prompt, 10 times: strict pass rate (Exact number): Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 100% (10/10) (n 10, 95% interval 72%–100%). Claude Sonnet 5.5 ahead.
  • Same prompt, 10 times: strict pass rate (JSON object): Claude Haiku 4.5 10% (1/10) (n 10, 95% interval 1.8%–40%); Claude Sonnet 5.5 100% (10/10) (n 10, 95% interval 72%–100%). Claude Sonnet 5.5 ahead.
  • Same prompt, 10 times: strict pass rate (Code fix): Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
  • Same prompt, 10 times: how many different answers (Exact number): Claude Haiku 4.5 1 (n 10); Claude Sonnet 5.5 1 (n 10). Tie.
  • Same prompt, 10 times: how many different answers (JSON object): Claude Haiku 4.5 1 (n 10); Claude Sonnet 5.5 1 (n 10). Tie.
  • Same prompt, 10 times: how many different answers (Code fix): Claude Haiku 4.5 6 (n 10); Claude Sonnet 5.5 3 (n 10). Unclear.
  • Same prompt, 10 times: time per call (Exact number): Claude Haiku 4.5 5.06 s (n 10, run range 4.4 s–6.2 s); Claude Sonnet 5.5 6.89 s (n 10, run range 5.8 s–7.8 s). Unclear.
  • Same prompt, 10 times: time per call (JSON object): Claude Haiku 4.5 7.03 s (n 10, run range 5.3 s–12.3 s); Claude Sonnet 5.5 2.89 s (n 10, run range 2.7 s–5.3 s). Unclear.
  • Same prompt, 10 times: time per call (Code fix): Claude Haiku 4.5 5.95 s (n 10, run range 4.9 s–7.3 s); Claude Sonnet 5.5 2.67 s (n 10, run range 2.3 s–4.3 s). Claude Sonnet 5.5 ahead.

Does memory help Claude Code? 8 kinds of agent memory, tested

Claude Sonnet 5.5 4 · 20 ties · 16 unclear
  • Full pass rate by kind of memory: No memory: Claude Haiku 4.5 20% (2/10) (n 10, 95% interval 5.7%–51%); Claude Sonnet 5.5 60% (9/15) (n 15, 95% interval 36%–80%). Tie.

    Calculation: at these rates, about 21 runs per side would separate them.

  • Full pass rate by kind of memory: /init CLAUDE.md: Claude Haiku 4.5 20% (2/10) (n 10, 95% interval 5.7%–51%); Claude Sonnet 5.5 60% (9/15) (n 15, 95% interval 36%–80%). Tie.

    Calculation: at these rates, about 21 runs per side would separate them.

  • Full pass rate by kind of memory: Curated, 11 lines: Claude Haiku 4.5 70% (7/10) (n 10, 95% interval 40%–89%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.

    Calculation: at these rates, about 22 runs per side would separate them.

  • Full pass rate by kind of memory: Raw notes, 60 lines: Claude Haiku 4.5 60% (6/10) (n 10, 95% interval 31%–83%); Claude Sonnet 5.5 93% (14/15) (n 15, 95% interval 70%–99%). Tie.

    Calculation: at these rates, about 22 runs per side would separate them.

  • Full pass rate by kind of memory: Dreamed notes: Claude Haiku 4.5 70% (7/10) (n 10, 95% interval 40%–89%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.

    Calculation: at these rates, about 22 runs per side would separate them.

  • Full pass rate by kind of memory: Handbook, 210 lines: Claude Haiku 4.5 30% (3/10) (n 10, 95% interval 11%–60%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Claude Sonnet 5.5 ahead.
  • Full pass rate by kind of memory: Stop hook only: Claude Haiku 4.5 80% (8/10) (n 10, 95% interval 49%–94%); Claude Sonnet 5.5 80% (12/15) (n 15, 95% interval 55%–93%). Tie.
  • Full pass rate by kind of memory: Curated + hook: Claude Haiku 4.5 90% (9/10) (n 10, 95% interval 60%–98%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.

    Calculation: at these rates, about 76 runs per side would separate them.

  • Team knowledge followed, Sonnet vs Haiku: No memory: Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 40% (6/15) (n 15, 95% interval 20%–64%). Tie.

    Calculation: at these rates, about 17 runs per side would separate them.

  • Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md: Claude Haiku 4.5 10% (1/10) (n 10, 95% interval 1.8%–40%); Claude Sonnet 5.5 67% (10/15) (n 15, 95% interval 42%–85%). Claude Sonnet 5.5 ahead.
  • Team knowledge followed, Sonnet vs Haiku: Curated, 11 lines: Claude Haiku 4.5 80% (8/10) (n 10, 95% interval 49%–94%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.

    Calculation: at these rates, about 33 runs per side would separate them.

  • Team knowledge followed, Sonnet vs Haiku: Raw notes, 60 lines: Claude Haiku 4.5 60% (6/10) (n 10, 95% interval 31%–83%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.

    Calculation: at these rates, about 17 runs per side would separate them.

  • Team knowledge followed, Sonnet vs Haiku: Dreamed notes: Claude Haiku 4.5 80% (8/10) (n 10, 95% interval 49%–94%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.

    Calculation: at these rates, about 33 runs per side would separate them.

  • Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines: Claude Haiku 4.5 30% (3/10) (n 10, 95% interval 11%–60%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Claude Sonnet 5.5 ahead.
  • Team knowledge followed, Sonnet vs Haiku: Stop hook only: Claude Haiku 4.5 80% (8/10) (n 10, 95% interval 49%–94%); Claude Sonnet 5.5 67% (10/15) (n 15, 95% interval 42%–85%). Tie.

    Calculation: at these rates, about 167 runs per side would separate them.

  • Team knowledge followed, Sonnet vs Haiku: Curated + hook: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
  • A stale README command: who still ran it?: No memory: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 80% (12/15) (n 15, 95% interval 55%–93%). Tie.

    Calculation: at these rates, about 33 runs per side would separate them.

  • A stale README command: who still ran it?: /init CLAUDE.md: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 87% (13/15) (n 15, 95% interval 62%–96%). Tie.

    Calculation: at these rates, about 57 runs per side would separate them.

  • A stale README command: who still ran it?: Curated, 11 lines: Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Tie.
  • A stale README command: who still ran it?: Raw notes, 60 lines: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Claude Sonnet 5.5 ahead.
  • A stale README command: who still ran it?: Dreamed notes: Claude Haiku 4.5 10% (1/10) (n 10, 95% interval 1.8%–40%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Tie.

    Calculation: at these rates, about 75 runs per side would separate them.

  • A stale README command: who still ran it?: Handbook, 210 lines: Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Tie.
  • A stale README command: who still ran it?: Stop hook only: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 60% (9/15) (n 15, 95% interval 36%–80%). Tie.

    Calculation: at these rates, about 17 runs per side would separate them.

  • A stale README command: who still ran it?: Curated + hook: Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Tie.
  • List-price cost per fully correct result (calculation): No memory, calculation: Claude Haiku 4.5 $0.38 (n 2); Claude Sonnet 5.5 $0.14 (n 9). Unclear.
  • List-price cost per fully correct result (calculation): /init CLAUDE.md, calculation: Claude Haiku 4.5 $0.43 (n 2); Claude Sonnet 5.5 $0.14 (n 9). Unclear.
  • List-price cost per fully correct result (calculation): Curated, 11 lines, calculation: Claude Haiku 4.5 $0.11 (n 7); Claude Sonnet 5.5 $0.082 (n 15). Unclear.
  • List-price cost per fully correct result (calculation): Raw notes, 60 lines, calculation: Claude Haiku 4.5 $0.13 (n 6); Claude Sonnet 5.5 $0.10 (n 14). Unclear.
  • List-price cost per fully correct result (calculation): Dreamed notes, calculation: Claude Haiku 4.5 $0.12 (n 7); Claude Sonnet 5.5 $0.090 (n 15). Unclear.
  • List-price cost per fully correct result (calculation): Handbook, 210 lines, calculation: Claude Haiku 4.5 $0.26 (n 3); Claude Sonnet 5.5 $0.10 (n 15). Unclear.
  • List-price cost per fully correct result (calculation): Stop hook only, calculation: Claude Haiku 4.5 $0.14 (n 8); Claude Sonnet 5.5 $0.13 (n 12). Unclear.
  • List-price cost per fully correct result (calculation): Curated + hook, calculation: Claude Haiku 4.5 $0.096 (n 9); Claude Sonnet 5.5 $0.084 (n 15). Unclear.
  • Time per session: No memory: Claude Haiku 4.5 54.2 s (n 10); Claude Sonnet 5.5 18.0 s (n 15). Unclear.
  • Time per session: /init CLAUDE.md: Claude Haiku 4.5 52.9 s (n 10); Claude Sonnet 5.5 19.0 s (n 15). Unclear.
  • Time per session: Curated, 11 lines: Claude Haiku 4.5 51.7 s (n 10); Claude Sonnet 5.5 21.9 s (n 15). Unclear.
  • Time per session: Raw notes, 60 lines: Claude Haiku 4.5 51.0 s (n 10); Claude Sonnet 5.5 26.8 s (n 15). Unclear.
  • Time per session: Dreamed notes: Claude Haiku 4.5 51.8 s (n 10); Claude Sonnet 5.5 27.2 s (n 15). Unclear.
  • Time per session: Handbook, 210 lines: Claude Haiku 4.5 49.9 s (n 10); Claude Sonnet 5.5 23.6 s (n 15). Unclear.
  • Time per session: Stop hook only: Claude Haiku 4.5 68.5 s (n 10); Claude Sonnet 5.5 27.3 s (n 15). Unclear.
  • Time per session: Curated + hook: Claude Haiku 4.5 52.8 s (n 10); Claude Sonnet 5.5 22.0 s (n 15). Unclear.

Jev vs Claude as a router: accuracy and cost

Claude Sonnet 5.5 2 · 6 ties · 1 unclear
  • Typed routing decisions answered exactly right: Claude Haiku 4.5 89% (73/82) (n 82, 95% interval 80%–94%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.

    Calculation: at these rates, about 497 runs per side would separate them.

  • Per-question accuracy: Claude Haiku 4.5 94% (183/194) (n 194, 95% interval 90%–97%); Claude Sonnet 5.5 97% (189/194) (n 194, 95% interval 94%–99%). Tie.

    Calculation: at these rates, about 627 runs per side would separate them.

  • Exact rate by decision type: Failure class: Claude Haiku 4.5 94% (17/18) (n 18, 95% interval 74%–99%); Claude Sonnet 5.5 100% (18/18) (n 18, 95% interval 82%–100%). Tie.

    Calculation: at these rates, about 135 runs per side would separate them.

  • Exact rate by decision type: Message intent: Claude Haiku 4.5 100% (20/20) (n 20, 95% interval 84%–100%); Claude Sonnet 5.5 100% (20/20) (n 20, 95% interval 84%–100%). Tie.
  • Exact rate by decision type: Is it a rule?: Claude Haiku 4.5 100% (12/12) (n 12, 95% interval 76%–100%); Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • Exact rate by decision type: Context shape: Claude Haiku 4.5 75% (24/32) (n 32, 95% interval 58%–87%); Claude Sonnet 5.5 84% (27/32) (n 32, 95% interval 68%–93%). Tie.

    Calculation: at these rates, about 271 runs per side would separate them.

  • Cost per 1,000 routing decisions, calculation: Claude Haiku 4.5 $8.92 (n 82); Claude Sonnet 5.5 $5.00 (n 82). Unclear.
  • Time per routing decision (Wall time (CLI)): Claude Haiku 4.5 12,674 ms (n 82, median to p95 12.67 s–34.41 s); Claude Sonnet 5.5 2,598 ms (n 82, median to p95 2.6 s–4.3 s). Claude Sonnet 5.5 ahead.
  • Time per routing decision (Model time (API)): Claude Haiku 4.5 10,734 ms (n 82, median to p95 10.73 s–32.07 s); Claude Sonnet 5.5 1,599 ms (n 82, median to p95 1.6 s–2.57 s). Claude Sonnet 5.5 ahead.

Routing overhead: deterministic policy vs LLM routers vs Jev

Claude Sonnet 5.5 3 · 1 tie · 4 unclear
  • Time to make one routing decision: Claude Haiku 4.5 12,543 ms (n 82, median to p95 12.54 s–34.48 s); Claude Sonnet 5.5 2,597 ms (n 82, median to p95 2.6 s–4.3 s). Claude Sonnet 5.5 ahead.
  • Where an LLM router’s time goes: model vs CLI (Model API time): Claude Haiku 4.5 10,508 ms (n 82, median to p95 10.51 s–32.13 s); Claude Sonnet 5.5 1,596 ms (n 82, median to p95 1.6 s–2.58 s). Claude Sonnet 5.5 ahead.
  • Where an LLM router’s time goes: model vs CLI (CLI and harness time): Claude Haiku 4.5 1,698 ms (n 82, median to p95 1.7 s–2.68 s); Claude Sonnet 5.5 973 ms (n 82, median to p95 973 ms–1.28 s). Claude Sonnet 5.5 ahead.
  • Routing calls that returned a decision: Claude Haiku 4.5 100% (82/82) (n 82, 95% interval 96%–100%); Claude Sonnet 5.5 100% (82/82) (n 82, 95% interval 96%–100%). Tie.
  • Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)), calculation: Claude Haiku 4.5 $441.74; Claude Sonnet 5.5 $247.30. Unclear.
  • Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)), calculation: Claude Haiku 4.5 $62.47; Claude Sonnet 5.5 $34.97. Unclear.
  • Added routing delay per task (calculation) (Every model call routed (49.5 per task)), calculation: Claude Haiku 4.5 620.9 s; Claude Sonnet 5.5 128.6 s. Unclear.
  • Added routing delay per task (calculation) (Only System One decisions (7 per task)), calculation: Claude Haiku 4.5 87.8 s; Claude Sonnet 5.5 18.2 s. Unclear.

Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks

Claude Sonnet 5.5 1 · 8 ties · 5 unclear
  • Strict pass rate: single call vs agent loop on eight hard tasks: Claude Haiku 4.5 46% (11/24) (n 24, 95% interval 28%–65%); Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Claude Sonnet 5.5 ahead.
  • Strict passes per task: single call vs agent loop: Interval merge fix: Claude Haiku 4.5 100% (3/3) (n 3, 95% interval 44%–100%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
  • Strict passes per task: single call vs agent loop: DST day-length fix: Claude Haiku 4.5 33% (1/3) (n 3, 95% interval 6.2%–79%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.

    Calculation: at these rates, about 7 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: CSV parser: Claude Haiku 4.5 67% (2/3) (n 3, 95% interval 21%–94%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.

    Calculation: at these rates, about 20 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: Event-loop order: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: Room schedule: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: SemVer regex: Claude Haiku 4.5 100% (3/3) (n 3, 95% interval 44%–100%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
  • Strict passes per task: single call vs agent loop: Money refactor: Claude Haiku 4.5 67% (2/3) (n 3, 95% interval 21%–94%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.

    Calculation: at these rates, about 20 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: SQL report: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Total time per attempt: single call vs agent loop: Claude Haiku 4.5 39.0 s (n 24, run range 15.3 s–75.1 s); Claude Sonnet 5.5 7.75 s (n 24, run range 2.3 s–34.8 s). Unclear.
  • Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)): Claude Haiku 4.5 3,941 (n 24, run range 3,879–4,221); Claude Sonnet 5.5 2,281 (n 24, run range 2,234–2,669). Unclear.
  • Tokens per attempt: single call vs agent loop (Output tokens): Claude Haiku 4.5 5,064 (n 24, run range 1,899–9,321); Claude Sonnet 5.5 1,050 (n 24, run range 176–3,895). Unclear.
  • Tool calls per agent-loop attempt: Claude Haiku 4.5 3 (n 24, run range 2–18); Claude Sonnet 5.5 0 (n 16, run range 0–3). Unclear.
  • List-price cost per strict pass: single call vs agent loop (calculation), calculation: Claude Haiku 4.5 $0.067 (n 24); Claude Sonnet 5.5 $0.014 (n 24). Unclear.

Does thinking pay for Claude Haiku 4.5? Thinking on vs off

Claude Sonnet 5.5 2 · 2 ties · 3 unclear
  • Haiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right)): Claude Haiku 4.5 87% (71/82) (n 82, 95% interval 78%–92%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.

    Calculation: at these rates, about 235 runs per side would separate them.

  • Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy): Claude Haiku 4.5 91% (177/194) (n 194, 95% interval 86%–94%); Claude Sonnet 5.5 97% (189/194) (n 194, 95% interval 94%–99%). Tie.

    Calculation: at these rates, about 200 runs per side would separate them.

  • Haiku thinking study: time per routing decision (Wall time (CLI)): Claude Haiku 4.5 4.66 s (n 82, median to p95 4.7 s–8.2 s); Claude Sonnet 5.5 2.60 s (n 82, median to p95 2.6 s–4.3 s). Claude Sonnet 5.5 ahead.
  • Haiku thinking study: time per routing decision (Model time (API)): Claude Haiku 4.5 3.79 s (n 82, median to p95 3.8 s–7.4 s); Claude Sonnet 5.5 1.60 s (n 82, median to p95 1.6 s–2.6 s). Claude Sonnet 5.5 ahead.
  • Haiku thinking study: thinking and visible output tokens per routing decision (Thinking tokens): Claude Haiku 4.5 0 (n 82); Claude Sonnet 5.5 2 (n 82). Unclear.
  • Haiku thinking study: thinking and visible output tokens per routing decision (Visible output tokens): Claude Haiku 4.5 366 (n 82); Claude Sonnet 5.5 105 (n 82). Unclear.
  • Haiku thinking study: list-price cost per 1,000 routing decisions (calculation), calculation: Claude Haiku 4.5 $3.36 (n 82); Claude Sonnet 5.5 $7.32 (n 82). Unclear.

Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI

Claude Sonnet 5.5 2 · 2 ties · 4 unclear
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON), calculation: Claude Haiku 4.5 0% (0/24) (n 24, 95% interval 0%–14%); Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Claude Sonnet 5.5 ahead.
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss)), calculation: Claude Haiku 4.5 71% (17/24) (n 24, 95% interval 51%–85%); Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Strict pass): Claude Haiku 4.5 0 (n 24); Claude Sonnet 5.5 12 (n 12). Unclear.
  • What each call produced: strict pass, format miss, wrong values or error (Format miss): Claude Haiku 4.5 17 (n 24); Claude Sonnet 5.5 0 (n 12). Unclear.
  • What each call produced: strict pass, format miss, wrong values or error (Wrong values): Claude Haiku 4.5 7 (n 24); Claude Sonnet 5.5 0 (n 12). Unclear.
  • What each call produced: strict pass, format miss, wrong values or error (Error): Claude Haiku 4.5 0 (n 24); Claude Sonnet 5.5 0 (n 12). Tie.
  • Time per call, instructions vs schema mode: Claude Haiku 4.5 9.52 s (n 24, run range 5.7 s–17 s); Claude Sonnet 5.5 3.52 s (n 12, run range 2.7 s–4.1 s). Claude Sonnet 5.5 ahead.
  • Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call): Claude Haiku 4.5 1,128 (n 24); Claude Sonnet 5.5 368 (n 12). Unclear.

Jev vs Claude routers on unseen decisions: a blind holdout

Claude Sonnet 5.5 2 · 8 ties · 1 unclear
  • Unseen routing decisions answered exactly right: Claude Haiku 4.5 79% (44/56) (n 56, 95% interval 66%–87%); Claude Sonnet 5.5 88% (49/56) (n 56, 95% interval 76%–94%). Tie.

    Calculation: at these rates, about 259 runs per side would separate them.

  • Per-question accuracy on unseen decisions: Claude Haiku 4.5 82% (102/125) (n 125, 95% interval 74%–87%); Claude Sonnet 5.5 92% (115/125) (n 125, 95% interval 86%–96%). Tie.

    Calculation: at these rates, about 155 runs per side would separate them.

  • Exact rate on unseen decisions, by decision type: Failure class: Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 100% (14/14) (n 14, 95% interval 78%–100%). Tie.

    Calculation: at these rates, about 106 runs per side would separate them.

  • Exact rate on unseen decisions, by decision type: Message intent: Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 100% (14/14) (n 14, 95% interval 78%–100%). Tie.

    Calculation: at these rates, about 106 runs per side would separate them.

  • Exact rate on unseen decisions, by decision type: Is it a rule?: Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.
  • Exact rate on unseen decisions, by decision type: Context shape: Claude Haiku 4.5 36% (5/14) (n 14, 95% interval 16%–61%); Claude Sonnet 5.5 57% (8/14) (n 14, 95% interval 33%–79%). Tie.

    Calculation: at these rates, about 82 runs per side would separate them.

  • Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm)): Claude Haiku 4.5 89% (73/82) (n 82, 95% interval 80%–94%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.

    Calculation: at these rates, about 497 runs per side would separate them.

  • Tuned case set vs unseen holdout: exact rate per router (Unseen holdout): Claude Haiku 4.5 79% (44/56) (n 56, 95% interval 66%–87%); Claude Sonnet 5.5 88% (49/56) (n 56, 95% interval 76%–94%). Tie.

    Calculation: at these rates, about 259 runs per side would separate them.

  • Time per routing decision, by route (Wall time): Claude Haiku 4.5 9.44 s (n 56, median to p95 9.4 s–25.4 s); Claude Sonnet 5.5 2.36 s (n 56, median to p95 2.4 s–3.7 s). Claude Sonnet 5.5 ahead.
  • Time per routing decision, by route (Model time (API, CLI-reported)): Claude Haiku 4.5 7.52 s (n 56, median to p95 7.5 s–23.9 s); Claude Sonnet 5.5 1.49 s (n 56, median to p95 1.5 s–2.4 s). Claude Sonnet 5.5 ahead.
  • Cost per 1,000 unseen routing decisions, calculation: Claude Haiku 4.5 $7.13 (n 56); Claude Sonnet 5.5 $7.24 (n 56). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

6 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Haiku 4.5 91.7% (n 24, run range 76%–99%); Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Haiku 4.5 $0.024 (n 24); Claude Sonnet 5.5 $0.0067 (n 24). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Haiku 4.5 $0.0018 (n 24); Claude Sonnet 5.5 $0.0037 (n 24). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Haiku 4.5 $0.0045 (n 24); Claude Sonnet 5.5 $0.0040 (n 24). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Haiku 4.5 91.7% (n 24, run range 76%–99%); Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Haiku 4.5 90.2% (n 15, run range 73%–98%); Claude Sonnet 5.5 0% (n 15, run range 0%–73%). Unclear.

Where the seconds go: first text, output speed and prompt size for 6 LLMs

1 tie · 9 unclear
  • Time to first text: a 250-line answer, six models: Claude Haiku 4.5 4.00 s (n 4, run range 2.8 s–6.4 s); Claude Sonnet 5.5 1.96 s (n 4, run range 0.9 s–4.1 s). Unclear.
  • Output speed after the first text: visible tokens per second (calculation), calculation: Claude Haiku 4.5 153 (n 4, run range 153–216); Claude Sonnet 5.5 232 (n 4, run range 230–233). Unclear.
  • Output speed in characters per second after the first text (calculation), calculation: Claude Haiku 4.5 547 (n 3, run range 546–548); Claude Sonnet 5.5 517 (n 4, run range 513–519). Unclear.
  • Time to first text as the prompt grows: 1k, calculation: Claude Haiku 4.5 1.93 s (n 3, run range 1.9 s–2 s); Claude Sonnet 5.5 1.45 s (n 3, run range 1.2 s–1.7 s). Unclear.
  • Time to first text as the prompt grows: 16k, calculation: Claude Haiku 4.5 2.27 s (n 3, run range 2.2 s–2.5 s); Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s). Unclear.
  • Time to first text as the prompt grows: 64k, calculation: Claude Haiku 4.5 2.78 s (n 3, run range 2.5 s–2.9 s); Claude Sonnet 5.5 3.07 s (n 3, run range 1.4 s–3.6 s). Unclear.
  • Total time per call by prompt size (1k prompt): Claude Haiku 4.5 2.34 s (n 3, run range 2.2 s–2.5 s); Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s). Unclear.
  • Total time per call by prompt size (16k prompt): Claude Haiku 4.5 2.79 s (n 3, run range 2.6 s–2.8 s); Claude Sonnet 5.5 2.10 s (n 3, run range 2 s–2.5 s). Unclear.
  • Total time per call by prompt size (64k prompt): Claude Haiku 4.5 3.13 s (n 3, run range 2.8 s–3.3 s); Claude Sonnet 5.5 3.44 s (n 3, run range 1.7 s–4.4 s). Unclear.
  • Exact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Haiku 4.5 100% (9/9) (n 9, 95% interval 70%–100%); Claude Sonnet 5.5 100% (9/9) (n 9, 95% interval 70%–100%). Tie.

Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts

8 unclear
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Interval merge fix, calculation: Claude Haiku 4.5 $0.018 (n 3); Claude Sonnet 5.5 $0.0056 (n 3). Unclear.
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): DST day length, calculation: Claude Haiku 4.5 $0.029 (n 3); Claude Sonnet 5.5 $0.025 (n 3). Unclear.
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): CSV parser, calculation: Claude Haiku 4.5 $0.029 (n 3); Claude Sonnet 5.5 $0.015 (n 3). Unclear.
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Event-loop order, calculation: Claude Haiku 4.5 $0.037 (n 3); Claude Sonnet 5.5 $0.016 (n 3). Unclear.
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Room schedule, calculation: Claude Haiku 4.5 $0.036 (n 3); Claude Sonnet 5.5 $0.012 (n 3). Unclear.
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SemVer regex, calculation: Claude Haiku 4.5 $0.042 (n 3); Claude Sonnet 5.5 $0.0051 (n 3). Unclear.
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Money refactor, calculation: Claude Haiku 4.5 $0.020 (n 3); Claude Sonnet 5.5 $0.0096 (n 3). Unclear.
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SQLite report query, calculation: Claude Haiku 4.5 $0.036 (n 3); Claude Sonnet 5.5 $0.017 (n 3). Unclear.

GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

6 ties · 3 unclear
  • Pass rate on 4 harder tasks (Strict pass): Claude Haiku 4.5 0% (0/12) (n 12, 95% interval 0%–24%); Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%). Tie.

    Calculation: at these rates, about 18 runs per side would separate them.

  • Pass rate on 4 harder tasks (Lenient (format misses counted)): Claude Haiku 4.5 0% (0/12) (n 12, 95% interval 0%–24%); Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%). Tie.

    Calculation: at these rates, about 18 runs per side would separate them.

  • Calls that tried a tool although tools were off: Claude Haiku 4.5 8% (1/12) (n 12, 95% interval 1.5%–35%); Claude Sonnet 5.5 31% (5/16) (n 16, 95% interval 14%–56%). Unclear.
  • Strict pass rate by task: 10x10 nonogram: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 100% (4/4) (n 4, 95% interval 51%–100%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Strict pass rate by task: Sudoku, 22 givens: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%). Tie.
  • Strict pass rate by task: 6x6 Skyscrapers: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%). Tie.
  • Strict pass rate by task: Seeded shuffle output: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 50% (2/4) (n 4, 95% interval 15%–85%). Tie.

    Calculation: at these rates, about 11 runs per side would separate them.

  • Total time per call on harder tasks: Claude Haiku 4.5 109.0 s (n 10, run range 25.7 s–224 s); Claude Sonnet 5.5 70.4 s (n 12, run range 4.3 s–210 s). Unclear.
  • Output tokens per call on harder tasks (Output tokens): Claude Haiku 4.5 12,508 (n 10, run range 2,965–26,532); Claude Sonnet 5.5 9,287 (n 12, run range 407–27,921). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals); median to p95 bands (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

153 rows from 14 studies. Claude Sonnet 5.5 ahead on 21; 58 ties, 74 unclear. A side is ahead only where the intervals or ranges do not overlap.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Haiku 4.5

  • Haiku thinking study: list-price cost per 1,000 routing decisions (calculation): $3.36 vs $7.32. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0018 vs $0.0037. A list-price calculation, not a measured difference. Calculation

When to pick Claude Sonnet 5.5

  • List-price cost per call (calculation): $0.0036 vs $0.0057. A list-price calculation, not a measured difference. Calculation
  • Pass rate on eight hard tasks (Strict pass): 100% (24/24) vs 46% (11/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).
  • Pass rate on eight hard tasks (Lenient (format misses counted)): 100% (24/24) vs 67% (16/24). The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Sonnet 5.5 86% to 100%).
  • List-price cost per strict pass on hard tasks (calculation): $0.014 vs $0.067. A list-price calculation, not a measured difference. Calculation
  • Same prompt, 10 times: strict pass rate (Exact number): 100% (10/10) vs 0% (0/10). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%).
  • Same prompt, 10 times: strict pass rate (JSON object): 100% (10/10) vs 10% (1/10). The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 72% to 100%).
  • Same prompt, 10 times: time per call (Code fix): 2.67 s vs 5.95 s. The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; Claude Sonnet 5.5 2.32 s to 4.34 s). A range is not a confidence interval.
  • Full pass rate by kind of memory: Handbook, 210 lines: 100% (15/15) vs 30% (3/10). The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%).
  • Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md: 67% (10/15) vs 10% (1/10). The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 42% to 85%).
  • Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines: 100% (15/15) vs 30% (3/10). The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%).
  • A stale README command: who still ran it?: Raw notes, 60 lines: 0% (0/15) vs 100% (10/10). The 95% intervals do not overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 0% to 20%).
  • List-price cost per fully correct result (calculation): No memory: $0.14 vs $0.38. A list-price calculation, not a measured difference. Calculation
  • List-price cost per fully correct result (calculation): /init CLAUDE.md: $0.14 vs $0.43. A list-price calculation, not a measured difference. Calculation
  • List-price cost per fully correct result (calculation): Handbook, 210 lines: $0.10 vs $0.26. A list-price calculation, not a measured difference. Calculation
  • Cost per 1,000 routing decisions: $5.00 vs $8.92. A list-price calculation, not a measured difference. Calculation
  • Time per routing decision (Wall time (CLI)): 2,598 ms vs 12,674 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,674 ms to 34,413 ms; Claude Sonnet 5.5 2,598 ms to 4,298 ms); not a confidence interval.
  • Time per routing decision (Model time (API)): 1,599 ms vs 10,734 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,734 ms to 32,072 ms; Claude Sonnet 5.5 1,599 ms to 2,574 ms); not a confidence interval.
  • Time to make one routing decision: 2,597 ms vs 12,543 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,543 ms to 34,481 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval.
  • Where an LLM router’s time goes: model vs CLI (Model API time): 1,596 ms vs 10,508 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,508 ms to 32,132 ms; Claude Sonnet 5.5 1,596 ms to 2,583 ms); not a confidence interval.
  • Where an LLM router’s time goes: model vs CLI (CLI and harness time): 973 ms vs 1,698 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 1,698 ms to 2,677 ms; Claude Sonnet 5.5 973 ms to 1,277 ms); not a confidence interval.
  • Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)): $247.30 vs $441.74. A list-price calculation, not a measured difference. Calculation
  • Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)): $34.97 vs $62.47. A list-price calculation, not a measured difference. Calculation
  • Added routing delay per task (calculation) (Every model call routed (49.5 per task)): 128.6 s vs 620.9 s. A list-price calculation, not a measured difference. Calculation
  • Added routing delay per task (calculation) (Only System One decisions (7 per task)): 18.2 s vs 87.8 s. A list-price calculation, not a measured difference. Calculation
  • Strict pass rate: single call vs agent loop on eight hard tasks: 100% (24/24) vs 46% (11/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).
  • List-price cost per strict pass: single call vs agent loop (calculation): $0.014 vs $0.067. A list-price calculation, not a measured difference. Calculation
  • Haiku thinking study: time per routing decision (Wall time (CLI)): 2.60 s vs 4.66 s. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 4.66 s to 8.18 s; Claude Sonnet 5.5 2.60 s to 4.30 s); not a confidence interval.
  • Haiku thinking study: time per routing decision (Model time (API)): 1.60 s vs 3.79 s. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 3.79 s to 7.43 s; Claude Sonnet 5.5 1.60 s to 2.58 s); not a confidence interval.
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON): 100% (12/12) vs 0% (0/24). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; Claude Sonnet 5.5 76% to 100%).
  • Time per call, instructions vs schema mode: 3.52 s vs 9.52 s. The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 5.67 s to 17.0 s; Claude Sonnet 5.5 2.67 s to 4.12 s). A range is not a confidence interval.
  • Time per routing decision, by route (Wall time): 2.36 s vs 9.44 s. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 9.44 s to 25.4 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.
  • Time per routing decision, by route (Model time (API, CLI-reported)): 1.49 s vs 7.52 s. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 7.52 s to 23.9 s; Claude Sonnet 5.5 1.49 s to 2.38 s); not a confidence interval.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0067 vs $0.024. A list-price calculation, not a measured difference. Calculation
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Interval merge fix: $0.0056 vs $0.018. A list-price calculation, not a measured difference. Calculation
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): CSV parser: $0.015 vs $0.029. A list-price calculation, not a measured difference. Calculation
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Event-loop order: $0.016 vs $0.037. A list-price calculation, not a measured difference. Calculation
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Room schedule: $0.012 vs $0.036. A list-price calculation, not a measured difference. Calculation
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SemVer regex: $0.0051 vs $0.042. A list-price calculation, not a measured difference. Calculation
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Money refactor: $0.0096 vs $0.020. A list-price calculation, not a measured difference. Calculation
  • List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SQLite report query: $0.017 vs $0.036. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only these two. Open a study for every configuration.

Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

2 rows. Highest Claude Haiku 4.5 · Claude Code 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 15 per row

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 3.2× real timeMotion reduced: press Replay to animateThe slowest median is 4.4 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest Claude Haiku 4.5 · Claude Code 4.4 s (range 3.2 s–23.6 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 2.3 s (range 2.2 s–7.7 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 2.6× real timeMotion reduced: press Replay to animateThe slowest median is 3.6 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest Claude Haiku 4.5 · Claude Code 3.6 s (range 2.8 s–22.3 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1 s–6.4 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

Totals are the sum of the parts shown. Shares are calculated from the same values.

2 rows, 2 series: Cache read, Other input. Cache read: highest Claude Sonnet 5.5 · Claude Code 1,401 (n 15). Lowest Claude Haiku 4.5 · Claude Code 0 (n 15). Other input: highest Claude Haiku 4.5 · Claude Code 3,790 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 685 (n 15).

Notesn = 15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Haiku 4.5 or Claude Sonnet 5.5?
Claude Haiku 4.5 and Claude Sonnet 5.5 share 113 measured metrics and 40 list-price calculations from 14 studies. Claude Sonnet 5.5 leads on 21 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); and 18 more. On those rows the 95% intervals, run ranges and p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 58 ties and 74 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 2 at the smallest).
How were Claude Haiku 4.5 and Claude Sonnet 5.5 measured?
They share 113 measured metrics and 40 list-price calculations from 14 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Prompt caching and run-to-run consistency in Claude Code and Codex CLI; Does memory help Claude Code? 8 kinds of agent memory, tested; Jev vs Claude as a router: accuracy and cost; Routing overhead: deterministic policy vs LLM routers vs Jev; Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks; Does thinking pay for Claude Haiku 4.5? Thinking on vs off; Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI; Jev vs Claude routers on unseen decisions: a blind holdout; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on pass rate on eight hard tasks (Strict pass)?
Claude Haiku 4.5: 46% (11/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 28% to 65%). Claude Sonnet 5.5: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).
How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on pass rate on eight hard tasks (Lenient (format misses counted))?
Claude Haiku 4.5: 67% (16/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 47% to 82%). Claude Sonnet 5.5: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Sonnet 5.5 86% to 100%).
How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on same prompt, 10 times: strict pass rate (Exact number)?
Claude Haiku 4.5: 0% (0/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 0% to 28%). Claude Sonnet 5.5: 100% (10/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 72% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%).
How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on same prompt, 10 times: strict pass rate (JSON object)?
Claude Haiku 4.5: 10% (1/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 1.8% to 40%). Claude Sonnet 5.5: 100% (10/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 72% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 72% to 100%).
How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on same prompt, 10 times: time per call (Code fix)?
Claude Haiku 4.5: 5.95 s (Claude Code · same prompt repeated 10 times; n = 10; run range 4.9 s to 7.3 s). Claude Sonnet 5.5: 2.67 s (Claude Code · same prompt repeated 10 times; n = 10; run range 2.3 s to 4.3 s). The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; Claude Sonnet 5.5 2.32 s to 4.34 s). A range is not a confidence interval.

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Claude Haiku
  • Extended Thinking

Does thinking pay for Claude Haiku 4.5? Thinking on vs off

Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.

87% (71/82)Claude Haiku 4.5 (thinking off): exact routing decisions · n = 82

6 chartsUpdated October 6, 2026

  • Routing
  • Jev

Jev vs Claude routers on unseen decisions: a blind holdout

Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.

82% (46/56)Jev 1.13 (TypeSafe): exact on unseen decisions · n = 56

6 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Claude Haiku

Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts

Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.

46% (11/24)Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks · n = 24

6 chartsUpdated October 7, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.