{"i":4,"comparison":{"slug":"claude-haiku-4-5-vs-claude-opus-5-5","a":"claude-haiku-4-5","b":"claude-opus-5-5","title":"Claude Haiku 4.5 vs Claude Opus 5.5","seoTitle":"Claude Haiku 4.5 vs Claude Opus 5.5: measured benchmarks","description":"Claude Haiku 4.5 vs Claude Opus 5.5: 25 measured metrics from 4 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","verdict":"Claude Haiku 4.5 and Claude Opus 5.5 share 25 measured metrics and 14 list-price calculations from 5 studies. Claude Opus 5.5 leads on 3 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Pass rate on 4 harder tasks (Lenient (format misses counted)), 50% (6/12) vs 0% (0/12). On those rows the 95% intervals do not overlap. The other rows are 7 ties and 29 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",4.43,2.75,"seconds","4.43 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[3.16,23.57],[2.47,8.91],"\u0001"],["Time to first useful output",3.63,1.92,"seconds","3.63 s","1.92 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; Claude Opus 5.5 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.78,22.27],[1.56,7.23],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",0,1401,"tokens","0","1,401","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",3790,680,"tokens","3,790","680","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",367,64,"tokens","367","64","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00566,0.00688,"usd","$0.0057","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; Claude Opus 5.5 $0.0059 to $0.022); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00513,0.01804],[0.00592,0.02226],true],["List-price cost per passing answer (calculation)",0.00836,0.01009,"usd","$0.0084","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Opus 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Opus 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,9.18,"seconds","39.0 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[15.27,75.13],[4.24,27.21],"\u0001"],["Time to first useful output on hard tasks",35.54,6.78,"seconds","35.5 s","6.78 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; Claude Opus 5.5 2.39 s to 21.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[12.88,70.31],[2.39,21.77],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",5064,945,"tokens","5,064","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.02824,"usd","$0.067","$0.028","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.028, 2.4x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",91.68,54.79,"percent","91.7%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[29.92,95.6],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.024492,0.012528,"usd","$0.024","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.001795,0.008003,"usd","$0.0018","$0.0080","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.00451,0.00771,"usd","$0.0045","$0.0077","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",91.68,54.79,"percent","91.7%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[29.92,95.6],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",90.19,0,"percent","90.2%","0%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[73.1,97.59],[0,93.33],true],["Time to first text: a 250-line answer, six models",4,1.97,"seconds","4.00 s","1.97 s","unclear","Only 4 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.84 s to 6.38 s; Claude Opus 5.5 1.70 s to 2.35 s), but 4 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[2.84,6.38],[1.7,2.35],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",153.2,155.5,"tokens","153","156","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[152.6,216.1],[154.6,156.4],true],["Output speed in characters per second after the first text (calculation)",547,347,"count","547","347","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","\u0001","speed-anatomy-chars-per-second",3,4,"Claude Code","Claude Code","range","minmax",[546,548],[345,349],true],["Time to first text as the prompt grows: 1k",1.93,1.51,"seconds","1.93 s","1.51 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 1.85 s to 2.04 s; Claude Opus 5.5 1.46 s to 2.01 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.85,2.04],[1.46,2.01],true],["Time to first text as the prompt grows: 16k",2.27,1.74,"seconds","2.27 s","1.74 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.22 s to 2.47 s; Claude Opus 5.5 1.70 s to 2.97 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[2.22,2.47],[1.7,2.97],true],["Time to first text as the prompt grows: 64k",2.78,1.79,"seconds","2.78 s","1.79 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.45 s to 2.89 s; Claude Opus 5.5 1.72 s to 3.72 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[2.45,2.89],[1.72,3.72],true],["Total time per call by prompt size (1k prompt)",2.34,1.83,"seconds","2.34 s","1.83 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.22 s to 2.46 s; Claude Opus 5.5 1.82 s to 2.41 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.22,2.46],[1.82,2.41],"\u0001"],["Total time per call by prompt size (16k prompt)",2.79,2.36,"seconds","2.79 s","2.36 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.58 s to 2.84 s; Claude Opus 5.5 2.11 s to 3.40 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.58,2.84],[2.11,3.4],"\u0001"],["Total time per call by prompt size (64k prompt)",3.13,2.35,"seconds","3.13 s","2.35 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 3.28 s; Claude Opus 5.5 2.26 s to 4.29 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.84,3.28],[2.26,4.29],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,0.5556,"rate","100% (9/9)","56% (5/9)","tie","The 95% intervals overlap (Claude Haiku 4.5 70% to 100%; Claude Opus 5.5 27% to 81%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Claude Code","ci95","ci95",[0.7009,1],[0.2667,0.8112],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0,0.4167,"rate","0% (0/12)","42% (5/12)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 24%; Claude Opus 5.5 19% to 68%), so this sample cannot separate them.","harder-tasks-head-to-head",12,"harder-h2h-pass-rate",12,12,"Claude Code","Claude Code","ci95","ci95",[0,0.2425],[0.1933,0.6805],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0,0.5,"rate","0% (0/12)","50% (6/12)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; Claude Opus 5.5 25% to 75%).","harder-tasks-head-to-head",12,"harder-h2h-pass-rate",12,12,"Claude Code","Claude Code","ci95","ci95",[0,0.2425],[0.2538,0.7462],"\u0001"],["Calls that tried a tool although tools were off",0.0833,0.4167,"rate","8% (1/12)","42% (5/12)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head",12,"harder-h2h-tool-attempts",12,12,"Claude Code","Claude Code","ci95","ci95",[0.0149,0.3539],[0.1933,0.6805],"\u0001"],["Strict pass rate by task: 10x10 nonogram",0,1,"rate","0% (0/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Opus 5.5 44% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head",3,"harder-h2h-pass-by-task",3,3,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.4385,1],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0,"rate","0% (0/3)","0% (0/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Opus 5.5 0% to 56%), so this sample cannot separate them.","harder-tasks-head-to-head",3,"harder-h2h-pass-by-task",3,3,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0,0.5615],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,0.3333,"rate","0% (0/3)","33% (1/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them.","harder-tasks-head-to-head",3,"harder-h2h-pass-by-task",3,3,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.0615,0.7923],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0,0.3333,"rate","0% (0/3)","33% (1/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them.","harder-tasks-head-to-head",3,"harder-h2h-pass-by-task",3,3,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.0615,0.7923],"\u0001"],["Total time per call on harder tasks",108.98,80.34,"seconds","109.0 s","80.3 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 25.7 s to 223.9 s; Claude Opus 5.5 3.82 s to 279.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",10,9,"Claude Code","Claude Code","range","minmax",[25.73,223.95],[3.82,279.5],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",12508,8420,"tokens","12,508","8,420","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",10,9,"Claude Code","Claude Code","range","minmax",[2965,26532],[279,40044],"\u0001"]]}}}