25 measured metrics · 14 calculated · 5 studies

Claude Haiku 4.5vsClaude Opus 5.5

Claude Opus 5.5 ahead on 3; 7 ties, 29 unclear. A side is ahead only where the intervals or ranges do not overlap.

The verdict

Claude Haiku 4.5 and Claude Opus 5.5 share 25 measured metrics and 14 list-price calculations from 5 studies. Claude Opus 5.5 leads on 3 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Pass rate on 4 harder tasks (Lenient (format misses counted)), 50% (6/12) vs 0% (0/12). On those rows the 95% intervals do not overlap. The other rows are 7 ties and 29 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Output tokens per call (Output tokens), 5.7x (Claude Haiku 4.5 larger).

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Haiku 4.5
  • Claude Opus 5.5
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

1 tie · 7 unclear
  • Pass rate on five validated tasks: Claude Haiku 4.5 100% (15/15) (n 15, 95% interval 80%–100%); Claude Opus 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
  • Total time per call: Claude Haiku 4.5 4.43 s (n 15, run range 3.2 s–23.6 s); Claude Opus 5.5 2.75 s (n 15, run range 2.5 s–8.9 s). Unclear.
  • Time to first useful output: Claude Haiku 4.5 3.63 s (n 15, run range 2.8 s–22.3 s); Claude Opus 5.5 1.92 s (n 15, run range 1.6 s–7.2 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): Claude Haiku 4.5 0 (n 15); Claude Opus 5.5 1,401 (n 15). Unclear.
  • Input tokens per call: what the CLI sends (Other input): Claude Haiku 4.5 3,790 (n 15); Claude Opus 5.5 680 (n 15). Unclear.
  • Output tokens per call (Output tokens): Claude Haiku 4.5 367 (n 15); Claude Opus 5.5 64 (n 15). Unclear.
  • List-price cost per call (calculation), calculation: Claude Haiku 4.5 $0.0057 (n 15, run range $0.0051–$0.018); Claude Opus 5.5 $0.0069 (n 15, run range $0.0059–$0.022). Unclear.
  • List-price cost per passing answer (calculation), calculation: Claude Haiku 4.5 $0.0084 (n 15); Claude Opus 5.5 $0.010 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

Claude Opus 5.5 2 · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): Claude Haiku 4.5 46% (11/24) (n 24, 95% interval 28%–65%); Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Claude Opus 5.5 ahead.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): Claude Haiku 4.5 67% (16/24) (n 24, 95% interval 47%–82%); Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Claude Opus 5.5 ahead.
  • Total time per call on hard tasks (separate batches): Claude Haiku 4.5 39.0 s (n 24, run range 15.3 s–75.1 s); Claude Opus 5.5 9.18 s (n 24, run range 4.2 s–27.2 s). Unclear.
  • Time to first useful output on hard tasks: Claude Haiku 4.5 35.5 s (n 24, run range 12.9 s–70.3 s); Claude Opus 5.5 6.78 s (n 24, run range 2.4 s–21.8 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): Claude Haiku 4.5 5,064 (n 24); Claude Opus 5.5 945 (n 24). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: Claude Haiku 4.5 $0.067 (n 24); Claude Opus 5.5 $0.028 (n 24). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

6 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Haiku 4.5 91.7% (n 24, run range 76%–99%); Claude Opus 5.5 54.8% (n 24, run range 30%–96%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Haiku 4.5 $0.024 (n 24); Claude Opus 5.5 $0.013 (n 24). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Haiku 4.5 $0.0018 (n 24); Claude Opus 5.5 $0.0080 (n 24). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Haiku 4.5 $0.0045 (n 24); Claude Opus 5.5 $0.0077 (n 24). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Haiku 4.5 91.7% (n 24, run range 76%–99%); Claude Opus 5.5 54.8% (n 24, run range 30%–96%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Haiku 4.5 90.2% (n 15, run range 73%–98%); Claude Opus 5.5 0% (n 15, run range 0%–93%). Unclear.

Where the seconds go: first text, output speed and prompt size for 6 LLMs

1 tie · 9 unclear
  • Time to first text: a 250-line answer, six models: Claude Haiku 4.5 4.00 s (n 4, run range 2.8 s–6.4 s); Claude Opus 5.5 1.97 s (n 4, run range 1.7 s–2.4 s). Unclear.
  • Output speed after the first text: visible tokens per second (calculation), calculation: Claude Haiku 4.5 153 (n 4, run range 153–216); Claude Opus 5.5 156 (n 4, run range 155–156). Unclear.
  • Output speed in characters per second after the first text (calculation), calculation: Claude Haiku 4.5 547 (n 3, run range 546–548); Claude Opus 5.5 347 (n 4, run range 345–349). Unclear.
  • Time to first text as the prompt grows: 1k, calculation: Claude Haiku 4.5 1.93 s (n 3, run range 1.9 s–2 s); Claude Opus 5.5 1.51 s (n 3, run range 1.5 s–2 s). Unclear.
  • Time to first text as the prompt grows: 16k, calculation: Claude Haiku 4.5 2.27 s (n 3, run range 2.2 s–2.5 s); Claude Opus 5.5 1.74 s (n 3, run range 1.7 s–3 s). Unclear.
  • Time to first text as the prompt grows: 64k, calculation: Claude Haiku 4.5 2.78 s (n 3, run range 2.5 s–2.9 s); Claude Opus 5.5 1.79 s (n 3, run range 1.7 s–3.7 s). Unclear.
  • Total time per call by prompt size (1k prompt): Claude Haiku 4.5 2.34 s (n 3, run range 2.2 s–2.5 s); Claude Opus 5.5 1.83 s (n 3, run range 1.8 s–2.4 s). Unclear.
  • Total time per call by prompt size (16k prompt): Claude Haiku 4.5 2.79 s (n 3, run range 2.6 s–2.8 s); Claude Opus 5.5 2.36 s (n 3, run range 2.1 s–3.4 s). Unclear.
  • Total time per call by prompt size (64k prompt): Claude Haiku 4.5 3.13 s (n 3, run range 2.8 s–3.3 s); Claude Opus 5.5 2.35 s (n 3, run range 2.3 s–4.3 s). Unclear.
  • Exact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Haiku 4.5 100% (9/9) (n 9, 95% interval 70%–100%); Claude Opus 5.5 56% (5/9) (n 9, 95% interval 27%–81%). Tie.

    Calculation: at these rates, about 13 runs per side would separate them.

GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

Claude Opus 5.5 1 · 5 ties · 3 unclear
  • Pass rate on 4 harder tasks (Strict pass): Claude Haiku 4.5 0% (0/12) (n 12, 95% interval 0%–24%); Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%). Tie.

    Calculation: at these rates, about 16 runs per side would separate them.

  • Pass rate on 4 harder tasks (Lenient (format misses counted)): Claude Haiku 4.5 0% (0/12) (n 12, 95% interval 0%–24%); Claude Opus 5.5 50% (6/12) (n 12, 95% interval 25%–75%). Claude Opus 5.5 ahead.
  • Calls that tried a tool although tools were off: Claude Haiku 4.5 8% (1/12) (n 12, 95% interval 1.5%–35%); Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%). Unclear.
  • Strict pass rate by task: 10x10 nonogram: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Opus 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Strict pass rate by task: Sudoku, 22 givens: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Opus 5.5 0% (0/3) (n 3, 95% interval 0%–56%). Tie.
  • Strict pass rate by task: 6x6 Skyscrapers: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%). Tie.

    Calculation: at these rates, about 20 runs per side would separate them.

  • Strict pass rate by task: Seeded shuffle output: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%). Tie.

    Calculation: at these rates, about 20 runs per side would separate them.

  • Total time per call on harder tasks: Claude Haiku 4.5 109.0 s (n 10, run range 25.7 s–224 s); Claude Opus 5.5 80.3 s (n 9, run range 3.8 s–280 s). Unclear.
  • Output tokens per call on harder tasks (Output tokens): Claude Haiku 4.5 12,508 (n 10, run range 2,965–26,532); Claude Opus 5.5 8,420 (n 9, run range 279–40,044). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

39 rows from 5 studies. Claude Opus 5.5 ahead on 3; 7 ties, 29 unclear. A side is ahead only where the intervals or ranges do not overlap.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Haiku 4.5

  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0018 vs $0.0080. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)): $0.0045 vs $0.0077. A list-price calculation, not a measured difference. Calculation

When to pick Claude Opus 5.5

  • Pass rate on eight hard tasks (Strict pass): 100% (24/24) vs 46% (11/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Opus 5.5 86% to 100%).
  • Pass rate on eight hard tasks (Lenient (format misses counted)): 100% (24/24) vs 67% (16/24). The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Opus 5.5 86% to 100%).
  • List-price cost per strict pass on hard tasks (calculation): $0.028 vs $0.067. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.013 vs $0.024. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 64k: 1.79 s vs 2.78 s. A list-price calculation, not a measured difference. Calculation
  • Pass rate on 4 harder tasks (Lenient (format misses counted)): 50% (6/12) vs 0% (0/12). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; Claude Opus 5.5 25% to 75%).

Side by side

The study charts, showing only these two. Open a study for every configuration.

Every rate is 95% or more
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code

2 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 15 per row

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 3.2× real timeMotion reduced: press Replay to animateThe slowest median is 4.4 s. The clock runs at the recorded speed.
Claude Opus 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest Claude Haiku 4.5 · Claude Code 4.4 s (range 3.2 s–23.6 s, n 15). Fastest Claude Opus 5.5 · Claude Code 2.8 s (range 2.5 s–8.9 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 2.6× real timeMotion reduced: press Replay to animateThe slowest median is 3.6 s. The clock runs at the recorded speed.
Claude Opus 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest Claude Haiku 4.5 · Claude Code 3.6 s (range 2.8 s–22.3 s, n 15). Fastest Claude Opus 5.5 · Claude Code 1.9 s (range 1.6 s–7.2 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Opus 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

Totals are the sum of the parts shown. Shares are calculated from the same values.

2 rows, 2 series: Cache read, Other input. Cache read: highest Claude Opus 5.5 · Claude Code 1,401 (n 15). Lowest Claude Haiku 4.5 · Claude Code 0 (n 15). Other input: highest Claude Haiku 4.5 · Claude Code 3,790 (n 15). Lowest Claude Opus 5.5 · Claude Code 680 (n 15).

Notesn = 15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Haiku 4.5 or Claude Opus 5.5?
Claude Haiku 4.5 and Claude Opus 5.5 share 25 measured metrics and 14 list-price calculations from 5 studies. Claude Opus 5.5 leads on 3 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Pass rate on 4 harder tasks (Lenient (format misses counted)), 50% (6/12) vs 0% (0/12). On those rows the 95% intervals do not overlap. The other rows are 7 ties and 29 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).
How were Claude Haiku 4.5 and Claude Opus 5.5 measured?
They share 25 measured metrics and 14 list-price calculations from 5 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
How do Claude Haiku 4.5 and Claude Opus 5.5 compare on pass rate on eight hard tasks (Strict pass)?
Claude Haiku 4.5: 46% (11/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 28% to 65%). Claude Opus 5.5: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Opus 5.5 86% to 100%).
How do Claude Haiku 4.5 and Claude Opus 5.5 compare on pass rate on eight hard tasks (Lenient (format misses counted))?
Claude Haiku 4.5: 67% (16/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 47% to 82%). Claude Opus 5.5: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Opus 5.5 86% to 100%).
How do Claude Haiku 4.5 and Claude Opus 5.5 compare on pass rate on 4 harder tasks (Lenient (format misses counted))?
Claude Haiku 4.5: 0% (0/12) (Claude Code; n = 12; 95% interval 0% to 24%). Claude Opus 5.5: 50% (6/12) (Claude Code; n = 12; 95% interval 25% to 75%). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; Claude Opus 5.5 25% to 75%).
How do Claude Haiku 4.5 and Claude Opus 5.5 compare on pass rate on five validated tasks?
Claude Haiku 4.5: 100% (15/15) (Claude Code · five short validated tasks; n = 15; 95% interval 80% to 100%). Claude Opus 5.5: 100% (15/15) (Claude Code · five short validated tasks; n = 15; 95% interval 80% to 100%). The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them.
How do Claude Haiku 4.5 and Claude Opus 5.5 compare on total time per call?
Claude Haiku 4.5: 4.43 s (Claude Code · five short validated tasks; n = 15; run range 3.2 s to 23.6 s). Claude Opus 5.5: 2.75 s (Claude Code · five short validated tasks; n = 15; run range 2.5 s to 8.9 s). The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.