{"i":1,"comparison":{"slug":"claude-haiku-4-5-vs-claude-sonnet-5-5","a":"claude-haiku-4-5","b":"claude-sonnet-5-5","title":"Claude Haiku 4.5 vs Claude Sonnet 5.5","seoTitle":"Claude Haiku 4.5 vs Claude Sonnet 5.5: measured benchmarks","description":"Claude Haiku 4.5 vs Claude Sonnet 5.5: 113 measured metrics from 12 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","verdict":"Claude Haiku 4.5 and Claude Sonnet 5.5 share 113 measured metrics and 40 list-price calculations from 14 studies. Claude Sonnet 5.5 leads on 21 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); and 18 more. On those rows the 95% intervals, run ranges and p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 58 ties and 74 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 2 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,0.8,"rate","100% (15/15)","80% (12/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.5481,0.9295],"\u0001"],["Total time per call",4.43,2.31,"seconds","4.43 s","2.31 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Sonnet 5.5 2.17 s to 7.73 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[3.16,23.57],[2.17,7.73],"\u0001"],["Time to first useful output",3.63,1.56,"seconds","3.63 s","1.56 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; Claude Sonnet 5.5 0.99 s to 6.39 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.78,22.27],[0.99,6.39],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",0,1401,"tokens","0","1,401","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",3790,685,"tokens","3,790","685","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",367,107,"tokens","367","107","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00566,0.0036,"usd","$0.0057","$0.0036","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; Claude Sonnet 5.5 $0.0034 to $0.010); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00513,0.01804],[0.00342,0.01021],true],["List-price cost per passing answer (calculation)",0.00836,0.00624,"usd","$0.0084","$0.0062","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.0062) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Sonnet 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,7.75,"seconds","39.0 s","7.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Sonnet 5.5 2.26 s to 34.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[15.27,75.13],[2.26,34.79],"\u0001"],["Time to first useful output on hard tasks",35.54,5.95,"seconds","35.5 s","5.95 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; Claude Sonnet 5.5 0.86 s to 30.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[12.88,70.31],[0.86,30.57],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",5064,1050,"tokens","5,064","1,050","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.01435,"usd","$0.067","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.014, 4.7x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",0,1,"rate","0% (0/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","ci95","ci95",[0,0.2775],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",0.1,1,"rate","10% (1/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","ci95","ci95",[0.0179,0.4042],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: how many different answers (Exact number)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (JSON object)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (Code fix)",6,3,"count","6","3","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: time per call (Exact number)",5.06,6.89,"seconds","5.06 s","6.89 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 4.42 s to 6.20 s; Claude Sonnet 5.5 5.81 s to 7.81 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","range","minmax",[4.42,6.2],[5.81,7.81],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",7.03,2.89,"seconds","7.03 s","2.89 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.28 s to 12.3 s; Claude Sonnet 5.5 2.68 s to 5.30 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","range","minmax",[5.28,12.27],[2.68,5.3],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",5.95,2.67,"seconds","5.95 s","2.67 s","b","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; Claude Sonnet 5.5 2.32 s to 4.34 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","range","minmax",[4.89,7.33],[2.32,4.34],"\u0001"],["Full pass rate by kind of memory: No memory",0.2,0.6,"rate","20% (2/10)","60% (9/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 6% to 51%; Claude Sonnet 5.5 36% to 80%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.0567,0.5098],[0.3575,0.8018],"\u0001"],["Full pass rate by kind of memory: /init CLAUDE.md",0.2,0.6,"rate","20% (2/10)","60% (9/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 6% to 51%; Claude Sonnet 5.5 36% to 80%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.0567,0.5098],[0.3575,0.8018],"\u0001"],["Full pass rate by kind of memory: Curated, 11 lines",0.7,1,"rate","70% (7/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 40% to 89%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.3968,0.8922],[0.7961,1],"\u0001"],["Full pass rate by kind of memory: Raw notes, 60 lines",0.6,0.9333,"rate","60% (6/10)","93% (14/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 31% to 83%; Claude Sonnet 5.5 70% to 99%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.3127,0.8318],[0.7018,0.9881],"\u0001"],["Full pass rate by kind of memory: Dreamed notes",0.7,1,"rate","70% (7/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 40% to 89%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.3968,0.8922],[0.7961,1],"\u0001"],["Full pass rate by kind of memory: Handbook, 210 lines",0.3,1,"rate","30% (3/10)","100% (15/15)","b","The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%).","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.1078,0.6032],[0.7961,1],"\u0001"],["Full pass rate by kind of memory: Stop hook only",0.8,0.8,"rate","80% (8/10)","80% (12/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.4902,0.9433],[0.5481,0.9295],"\u0001"],["Full pass rate by kind of memory: Curated + hook",0.9,1,"rate","90% (9/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 60% to 98%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.5958,0.9821],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: No memory",0,0.4,"rate","0% (0/10)","40% (6/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 20% to 64%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0,0.2775],[0.1982,0.6425],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md",0.1,0.6667,"rate","10% (1/10)","67% (10/15)","b","The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 42% to 85%).","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.0179,0.4042],[0.4171,0.8482],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Curated, 11 lines",0.8,1,"rate","80% (8/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.4902,0.9433],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Raw notes, 60 lines",0.6,1,"rate","60% (6/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 31% to 83%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.3127,0.8318],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Dreamed notes",0.8,1,"rate","80% (8/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.4902,0.9433],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines",0.3,1,"rate","30% (3/10)","100% (15/15)","b","The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%).","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.1078,0.6032],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Stop hook only",0.8,0.6667,"rate","80% (8/10)","67% (10/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 42% to 85%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.4902,0.9433],[0.4171,0.8482],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Curated + hook",1,1,"rate","100% (10/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.7225,1],[0.7961,1],"\u0001"],["A stale README command: who still ran it?: No memory",1,0.8,"rate","100% (10/10)","80% (12/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.7225,1],[0.5481,0.9295],"\u0001"],["A stale README command: who still ran it?: /init CLAUDE.md",1,0.8667,"rate","100% (10/10)","87% (13/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 62% to 96%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.7225,1],[0.6212,0.9626],"\u0001"],["A stale README command: who still ran it?: Curated, 11 lines",0,0,"rate","0% (0/10)","0% (0/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0,0.2775],[0,0.2039],"\u0001"],["A stale README command: who still ran it?: Raw notes, 60 lines",1,0,"rate","100% (10/10)","0% (0/15)","b","The 95% intervals do not overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 0% to 20%).","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.7225,1],[0,0.2039],"\u0001"],["A stale README command: who still ran it?: Dreamed notes",0.1,0,"rate","10% (1/10)","0% (0/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.0179,0.4042],[0,0.2039],"\u0001"],["A stale README command: who still ran it?: Handbook, 210 lines",0,0,"rate","0% (0/10)","0% (0/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0,0.2775],[0,0.2039],"\u0001"],["A stale README command: who still ran it?: Stop hook only",1,0.6,"rate","100% (10/10)","60% (9/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 36% to 80%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.7225,1],[0.3575,0.8018],"\u0001"],["A stale README command: who still ran it?: Curated + hook",0,0,"rate","0% (0/10)","0% (0/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0,0.2775],[0,0.2039],"\u0001"],["List-price cost per fully correct result (calculation): No memory",0.3786,0.1386,"usd","$0.38","$0.14","unclear","No interval or range was recorded for either side, so the gap ($0.38 vs $0.14, 2.7x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",2,9,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): /init CLAUDE.md",0.4317,0.1359,"usd","$0.43","$0.14","unclear","No interval or range was recorded for either side, so the gap ($0.43 vs $0.14, 3.2x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",2,9,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Curated, 11 lines",0.1095,0.0818,"usd","$0.11","$0.082","unclear","No interval or range was recorded for either side, so the gap ($0.11 vs $0.082) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",7,15,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Raw notes, 60 lines",0.1255,0.1009,"usd","$0.13","$0.10","unclear","No interval or range was recorded for either side, so the gap ($0.13 vs $0.10) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",6,14,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Dreamed notes",0.1153,0.0896,"usd","$0.12","$0.090","unclear","No interval or range was recorded for either side, so the gap ($0.12 vs $0.090) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",7,15,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Handbook, 210 lines",0.2615,0.1006,"usd","$0.26","$0.10","unclear","No interval or range was recorded for either side, so the gap ($0.26 vs $0.10, 2.6x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",3,15,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Stop hook only",0.1419,0.1278,"usd","$0.14","$0.13","unclear","No interval or range was recorded for either side, so the gap ($0.14 vs $0.13) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",8,12,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Curated + hook",0.096,0.0843,"usd","$0.096","$0.084","unclear","No interval or range was recorded for either side, so the gap ($0.096 vs $0.084) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",9,15,"","","\u0001","\u0001","\u0001","\u0001",true],["Time per session: No memory",54.2,18,"seconds","54.2 s","18.0 s","unclear","No interval or range was recorded for either side, so the gap (54.2 s vs 18.0 s, 3.0x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: /init CLAUDE.md",52.9,19,"seconds","52.9 s","19.0 s","unclear","No interval or range was recorded for either side, so the gap (52.9 s vs 19.0 s, 2.8x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Curated, 11 lines",51.7,21.9,"seconds","51.7 s","21.9 s","unclear","No interval or range was recorded for either side, so the gap (51.7 s vs 21.9 s, 2.4x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Raw notes, 60 lines",51,26.8,"seconds","51.0 s","26.8 s","unclear","No interval or range was recorded for either side, so the gap (51.0 s vs 26.8 s) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Dreamed notes",51.8,27.2,"seconds","51.8 s","27.2 s","unclear","No interval or range was recorded for either side, so the gap (51.8 s vs 27.2 s) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Handbook, 210 lines",49.9,23.6,"seconds","49.9 s","23.6 s","unclear","No interval or range was recorded for either side, so the gap (49.9 s vs 23.6 s, 2.1x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Stop hook only",68.5,27.3,"seconds","68.5 s","27.3 s","unclear","No interval or range was recorded for either side, so the gap (68.5 s vs 27.3 s, 2.5x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Curated + hook",52.8,22,"seconds","52.8 s","22.0 s","unclear","No interval or range was recorded for either side, so the gap (52.8 s vs 22.0 s, 2.4x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Typed routing decisions answered exactly right",0.8902,0.939,"rate","89% (73/82)","94% (77/82)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 94%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","routing-jev-vs-llm",82,"routing-exact-decisions",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.8044,0.9412],[0.8651,0.9737],"\u0001"],["Per-question accuracy",0.9433,0.9742,"rate","94% (183/194)","97% (189/194)","tie","The 95% intervals overlap (Claude Haiku 4.5 90% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them.","routing-jev-vs-llm",194,"routing-key-accuracy",194,194,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.9013,0.968],[0.9411,0.9889],"\u0001"],["Exact rate by decision type: Failure class",0.9444,1,"rate","94% (17/18)","100% (18/18)","tie","The 95% intervals overlap (Claude Haiku 4.5 74% to 99%; Claude Sonnet 5.5 82% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",18,"routing-exact-by-decision",18,18,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.7424,0.9901],[0.8241,1],"\u0001"],["Exact rate by decision type: Message intent",1,1,"rate","100% (20/20)","100% (20/20)","tie","The 95% intervals overlap (Claude Haiku 4.5 84% to 100%; Claude Sonnet 5.5 84% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",20,"routing-exact-by-decision",20,20,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.8389,1],[0.8389,1],"\u0001"],["Exact rate by decision type: Is it a rule?",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Haiku 4.5 76% to 100%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",12,"routing-exact-by-decision",12,12,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Exact rate by decision type: Context shape",0.75,0.8438,"rate","75% (24/32)","84% (27/32)","tie","The 95% intervals overlap (Claude Haiku 4.5 58% to 87%; Claude Sonnet 5.5 68% to 93%), so this sample cannot separate them.","routing-jev-vs-llm",32,"routing-exact-by-decision",32,32,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.5789,0.8675],[0.6825,0.9314],"\u0001"],["Cost per 1,000 routing decisions",8.924,4.996,"usd","$8.92","$5.00","unclear","No interval or range was recorded for either side, so the gap ($8.92 vs $5.00) is not tested against run-to-run variation.","routing-jev-vs-llm",82,"routing-cost-per-1000",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Time per routing decision (Wall time (CLI))",12674,2598,"ms","12,674 ms","2,598 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,674 ms to 34,413 ms; Claude Sonnet 5.5 2,598 ms to 4,298 ms); not a confidence interval.","routing-jev-vs-llm",82,"routing-decision-latency",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","range","p50-p95",[12674,34413],[2598,4298],"\u0001"],["Time per routing decision (Model time (API))",10734,1599,"ms","10,734 ms","1,599 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,734 ms to 32,072 ms; Claude Sonnet 5.5 1,599 ms to 2,574 ms); not a confidence interval.","routing-jev-vs-llm",82,"routing-decision-latency",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","range","p50-p95",[10734,32072],[1599,2574],"\u0001"],["Time to make one routing decision",12543,2597,"ms","12,543 ms","2,597 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,543 ms to 34,481 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval.","routing-overhead",82,"router-overhead-decision-latency",82,82,"thinking on · via Claude Code · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","range","p50-p95",[12543,34481],[2597,4298],"\u0001"],["Where an LLM router’s time goes: model vs CLI (Model API time)",10508,1596,"ms","10,508 ms","1,596 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,508 ms to 32,132 ms; Claude Sonnet 5.5 1,596 ms to 2,583 ms); not a confidence interval.","routing-overhead",82,"router-overhead-cli-vs-model-time",82,82,"thinking on · via Claude Code · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","range","p50-p95",[10508,32132],[1596,2583],"\u0001"],["Where an LLM router’s time goes: model vs CLI (CLI and harness time)",1698,973,"ms","1,698 ms","973 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 1,698 ms to 2,677 ms; Claude Sonnet 5.5 973 ms to 1,277 ms); not a confidence interval.","routing-overhead",82,"router-overhead-cli-vs-model-time",82,82,"thinking on · via Claude Code · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","range","p50-p95",[1698,2677],[973,1277],"\u0001"],["Routing calls that returned a decision",1,1,"rate","100% (82/82)","100% (82/82)","tie","The 95% intervals overlap (Claude Haiku 4.5 96% to 100%; Claude Sonnet 5.5 96% to 100%), so this sample cannot separate them.","routing-overhead",82,"router-overhead-completed",82,82,"thinking on · via Claude Code · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","ci95","ci95",[0.9552,1],[0.9552,1],"\u0001"],["Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))",441.74,247.3,"usd","$441.74","$247.30","unclear","No interval or range was recorded for either side, so the gap ($441.74 vs $247.30) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))",62.47,34.97,"usd","$62.47","$34.97","unclear","No interval or range was recorded for either side, so the gap ($62.47 vs $34.97) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Every model call routed (49.5 per task))",620.8785,128.5515,"seconds","620.9 s","128.6 s","unclear","No interval or range was recorded for either side, so the gap (620.9 s vs 128.6 s, 4.8x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Only System One decisions (7 per task))",87.801,18.179,"seconds","87.8 s","18.2 s","unclear","No interval or range was recorded for either side, so the gap (87.8 s vs 18.2 s, 4.8x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate: single call vs agent loop on eight hard tasks",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).","single-call-vs-agent-loop",24,"agent-loop-pass-rate",24,24,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Strict passes per task: single call vs agent loop: Interval merge fix",1,1,"rate","100% (3/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.4385,1],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: DST day-length fix",0.3333,1,"rate","33% (1/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 6% to 79%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.0615,0.7923],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: CSV parser",0.6667,1,"rate","67% (2/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.2077,0.9385],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: Event-loop order",0,1,"rate","0% (0/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0,0.5615],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: Room schedule",0,1,"rate","0% (0/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0,0.5615],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: SemVer regex",1,1,"rate","100% (3/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.4385,1],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: Money refactor",0.6667,1,"rate","67% (2/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.2077,0.9385],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: SQL report",0,1,"rate","0% (0/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0,0.5615],[0.4385,1],"\u0001"],["Total time per attempt: single call vs agent loop",39.01,7.75,"seconds","39.0 s","7.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Sonnet 5.5 2.26 s to 34.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","single-call-vs-agent-loop",24,"agent-loop-total-time",24,24,"Claude Code · single call","Claude Code · single call","range","minmax",[15.27,75.13],[2.26,34.79],"\u0001"],["Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))",3941,2281,"tokens","3,941","2,281","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop",24,"agent-loop-tokens",24,24,"Claude Code · single call","Claude Code · single call","range","minmax",[3879,4221],[2234,2669],"\u0001"],["Tokens per attempt: single call vs agent loop (Output tokens)",5064,1050,"tokens","5,064","1,050","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop",24,"agent-loop-tokens",24,24,"Claude Code · single call","Claude Code · single call","range","minmax",[1899,9321],[176,3895],"\u0001"],["Tool calls per agent-loop attempt",3,0,"count","3","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","\u0001","agent-loop-tool-calls",24,16,"Claude Code · agent loop","Claude Code · agent loop","range","minmax",[2,18],[0,3],"\u0001"],["List-price cost per strict pass: single call vs agent loop (calculation)",0.0672,0.01435,"usd","$0.067","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.014, 4.7x) is not tested against run-to-run variation.","single-call-vs-agent-loop",24,"agent-loop-cost-per-pass",24,24,"Claude Code · single call","Claude Code · single call","\u0001","\u0001","\u0001","\u0001",true],["Haiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right))",0.8659,0.939,"rate","87% (71/82)","94% (77/82)","tie","The 95% intervals overlap (Claude Haiku 4.5 78% to 92%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","haiku-thinking-on-off",82,"haiku-thinking-router-exact",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","ci95","ci95",[0.7755,0.9234],[0.8651,0.9737],"\u0001"],["Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)",0.9124,0.9742,"rate","91% (177/194)","97% (189/194)","tie","The 95% intervals overlap (Claude Haiku 4.5 86% to 94%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them.","haiku-thinking-on-off",194,"haiku-thinking-router-exact",194,194,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","ci95","ci95",[0.8642,0.9446],[0.9411,0.9889],"\u0001"],["Haiku thinking study: time per routing decision (Wall time (CLI))",4.66,2.6,"seconds","4.66 s","2.60 s","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 4.66 s to 8.18 s; Claude Sonnet 5.5 2.60 s to 4.30 s); not a confidence interval.","haiku-thinking-on-off",82,"haiku-thinking-router-latency",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","range","p50-p95",[4.66,8.18],[2.6,4.3],"\u0001"],["Haiku thinking study: time per routing decision (Model time (API))",3.79,1.6,"seconds","3.79 s","1.60 s","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 3.79 s to 7.43 s; Claude Sonnet 5.5 1.60 s to 2.58 s); not a confidence interval.","haiku-thinking-on-off",82,"haiku-thinking-router-latency",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","range","p50-p95",[3.79,7.43],[1.6,2.58],"\u0001"],["Haiku thinking study: thinking and visible output tokens per routing decision (Thinking tokens)",0,2,"tokens","0","2","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","haiku-thinking-on-off",82,"haiku-thinking-router-tokens",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","\u0001","\u0001","\u0001","\u0001","\u0001"],["Haiku thinking study: thinking and visible output tokens per routing decision (Visible output tokens)",366,105,"tokens","366","105","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","haiku-thinking-on-off",82,"haiku-thinking-router-tokens",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","\u0001","\u0001","\u0001","\u0001","\u0001"],["Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)",3.364,7.324,"usd","$3.36","$7.32","unclear","No interval or range was recorded for either side, so the gap ($3.36 vs $7.32, 2.2x) is not tested against run-to-run variation.","haiku-thinking-on-off",82,"haiku-thinking-router-cost",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","\u0001","\u0001","\u0001","\u0001",true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",0,1,"rate","0% (0/24)","100% (12/12)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; Claude Sonnet 5.5 76% to 100%).","json-schema-vs-instructions","\u0001","structured-output-pass-rate",24,12,"Claude Code · instructions","Claude Code · instructions","ci95","ci95",[0,0.138],[0.7575,1],true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))",0.7083,1,"rate","71% (17/24)","100% (12/12)","tie","The 95% intervals overlap (Claude Haiku 4.5 51% to 85%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions","\u0001","structured-output-pass-rate",24,12,"Claude Code · instructions","Claude Code · instructions","ci95","ci95",[0.5083,0.8509],[0.7575,1],true],["What each call produced: strict pass, format miss, wrong values or error (Strict pass)",0,12,"count","0","12","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Format miss)",17,0,"count","17","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Wrong values)",7,0,"count","7","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Error)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per call, instructions vs schema mode",9.52,3.52,"seconds","9.52 s","3.52 s","b","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 5.67 s to 17.0 s; Claude Sonnet 5.5 2.67 s to 4.12 s). A range is not a confidence interval.","json-schema-vs-instructions","\u0001","structured-output-time",24,12,"Claude Code · instructions","Claude Code · instructions","range","minmax",[5.67,17],[2.67,4.12],"\u0001"],["Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)",1128,368,"tokens","1,128","368","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-tokens",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Unseen routing decisions answered exactly right",0.7857,0.875,"rate","79% (44/56)","88% (49/56)","tie","The 95% intervals overlap (Claude Haiku 4.5 66% to 87%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-exact",56,56,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6618,0.8729],[0.7637,0.9381],"\u0001"],["Per-question accuracy on unseen decisions",0.816,0.92,"rate","82% (102/125)","92% (115/125)","tie","The 95% intervals overlap (Claude Haiku 4.5 74% to 87%; Claude Sonnet 5.5 86% to 96%), so this sample cannot separate them.","routing-holdout",125,"routing-holdout-key-accuracy",125,125,"Claude Code","Claude Code · effort low","ci95","ci95",[0.739,0.8741],[0.859,0.956],"\u0001"],["Exact rate on unseen decisions, by decision type: Failure class",0.9286,1,"rate","93% (13/14)","100% (14/14)","tie","The 95% intervals overlap (Claude Haiku 4.5 69% to 99%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.7847,1],"\u0001"],["Exact rate on unseen decisions, by decision type: Message intent",0.9286,1,"rate","93% (13/14)","100% (14/14)","tie","The 95% intervals overlap (Claude Haiku 4.5 69% to 99%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.7847,1],"\u0001"],["Exact rate on unseen decisions, by decision type: Is it a rule?",0.9286,0.9286,"rate","93% (13/14)","93% (13/14)","tie","The 95% intervals overlap (Claude Haiku 4.5 69% to 99%; Claude Sonnet 5.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Context shape",0.3571,0.5714,"rate","36% (5/14)","57% (8/14)","tie","The 95% intervals overlap (Claude Haiku 4.5 16% to 61%; Claude Sonnet 5.5 33% to 79%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"Claude Code","Claude Code · effort low","ci95","ci95",[0.1634,0.6124],[0.3259,0.7862],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm))",0.8902,0.939,"rate","89% (73/82)","94% (77/82)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 94%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","routing-holdout",82,"routing-holdout-tuned-vs-unseen",82,82,"Claude Code","Claude Code · effort low","ci95","ci95",[0.8044,0.9412],[0.8651,0.9737],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Unseen holdout)",0.7857,0.875,"rate","79% (44/56)","88% (49/56)","tie","The 95% intervals overlap (Claude Haiku 4.5 66% to 87%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-tuned-vs-unseen",56,56,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6618,0.8729],[0.7637,0.9381],"\u0001"],["Time per routing decision, by route (Wall time)",9.444,2.359,"seconds","9.44 s","2.36 s","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 9.44 s to 25.4 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.","routing-holdout",56,"routing-holdout-latency",56,56,"Claude Code","Claude Code · effort low","range","p50-p95",[9.444,25.413],[2.359,3.657],"\u0001"],["Time per routing decision, by route (Model time (API, CLI-reported))",7.522,1.485,"seconds","7.52 s","1.49 s","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 7.52 s to 23.9 s; Claude Sonnet 5.5 1.49 s to 2.38 s); not a confidence interval.","routing-holdout",56,"routing-holdout-latency",56,56,"Claude Code","Claude Code · effort low","range","p50-p95",[7.522,23.913],[1.485,2.377],"\u0001"],["Cost per 1,000 unseen routing decisions",7.129,7.244,"usd","$7.13","$7.24","unclear","No interval or range was recorded for either side, so the gap ($7.13 vs $7.24) is not tested against run-to-run variation.","routing-holdout",56,"routing-holdout-cost-per-1000",56,56,"Claude Code","Claude Code · effort low","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",91.68,54.54,"percent","91.7%","54.5%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[0,95.91],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.024492,0.006665,"usd","$0.024","$0.0067","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.001795,0.003672,"usd","$0.0018","$0.0037","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.00451,0.004012,"usd","$0.0045","$0.0040","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",91.68,54.54,"percent","91.7%","54.5%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[0,95.91],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",90.19,0,"percent","90.2%","0%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[73.1,97.59],[0,72.75],true],["Time to first text: a 250-line answer, six models",4,1.96,"seconds","4.00 s","1.96 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; Claude Sonnet 5.5 0.88 s to 4.09 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[2.84,6.38],[0.88,4.09],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",153.2,231.7,"tokens","153","232","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[152.6,216.1],[230.3,233],true],["Output speed in characters per second after the first text (calculation)",547,517,"count","547","517","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","\u0001","speed-anatomy-chars-per-second",3,4,"Claude Code","Claude Code","range","minmax",[546,548],[513,519],true],["Time to first text as the prompt grows: 1k",1.93,1.45,"seconds","1.93 s","1.45 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 1.85 s to 2.04 s; Claude Sonnet 5.5 1.23 s to 1.72 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.85,2.04],[1.23,1.72],true],["Time to first text as the prompt grows: 16k",2.27,1.78,"seconds","2.27 s","1.78 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.47 s; Claude Sonnet 5.5 1.64 s to 2.11 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[2.22,2.47],[1.64,2.11],true],["Time to first text as the prompt grows: 64k",2.78,3.07,"seconds","2.78 s","3.07 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.45 s to 2.89 s; Claude Sonnet 5.5 1.38 s to 3.61 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[2.45,2.89],[1.38,3.61],true],["Total time per call by prompt size (1k prompt)",2.34,1.78,"seconds","2.34 s","1.78 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.46 s; Claude Sonnet 5.5 1.57 s to 2.12 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.22,2.46],[1.57,2.12],"\u0001"],["Total time per call by prompt size (16k prompt)",2.79,2.1,"seconds","2.79 s","2.10 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.58 s to 2.84 s; Claude Sonnet 5.5 1.98 s to 2.48 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.58,2.84],[1.98,2.48],"\u0001"],["Total time per call by prompt size (64k prompt)",3.13,3.44,"seconds","3.13 s","3.44 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 3.28 s; Claude Sonnet 5.5 1.74 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.84,3.28],[1.74,4.38],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,1,"rate","100% (9/9)","100% (9/9)","tie","The 95% intervals overlap (Claude Haiku 4.5 70% to 100%; Claude Sonnet 5.5 70% to 100%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Claude Code","ci95","ci95",[0.7009,1],[0.7009,1],"\u0001"],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Interval merge fix",0.01789,0.00557,"usd","$0.018","$0.0056","unclear","No interval or range was recorded for either side, so the gap ($0.018 vs $0.0056, 3.2x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): DST day length",0.0293,0.02532,"usd","$0.029","$0.025","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.025) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): CSV parser",0.02919,0.01464,"usd","$0.029","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.015) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Event-loop order",0.03708,0.01588,"usd","$0.037","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.037 vs $0.016, 2.3x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Room schedule",0.03578,0.01243,"usd","$0.036","$0.012","unclear","No interval or range was recorded for either side, so the gap ($0.036 vs $0.012, 2.9x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SemVer regex",0.04217,0.00514,"usd","$0.042","$0.0051","unclear","No interval or range was recorded for either side, so the gap ($0.042 vs $0.0051, 8.2x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Money refactor",0.02039,0.00961,"usd","$0.020","$0.0096","unclear","No interval or range was recorded for either side, so the gap ($0.020 vs $0.0096, 2.1x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SQLite report query",0.03648,0.01719,"usd","$0.036","$0.017","unclear","No interval or range was recorded for either side, so the gap ($0.036 vs $0.017, 2.1x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on 4 harder tasks (Strict pass)",0,0.375,"rate","0% (0/12)","38% (6/16)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 24%; Claude Sonnet 5.5 18% to 61%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Claude Code","ci95","ci95",[0,0.2425],[0.1848,0.6136],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0,0.375,"rate","0% (0/12)","38% (6/16)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 24%; Claude Sonnet 5.5 18% to 61%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Claude Code","ci95","ci95",[0,0.2425],[0.1848,0.6136],"\u0001"],["Calls that tried a tool although tools were off",0.0833,0.3125,"rate","8% (1/12)","31% (5/16)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-tool-attempts",12,16,"Claude Code","Claude Code","ci95","ci95",[0.0149,0.3539],[0.1416,0.556],"\u0001"],["Strict pass rate by task: 10x10 nonogram",0,1,"rate","0% (0/3)","100% (4/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 51% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.5101,1],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0,"rate","0% (0/3)","0% (0/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 0% to 49%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0,0.4899],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,0,"rate","0% (0/3)","0% (0/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 0% to 49%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0,0.4899],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0,0.5,"rate","0% (0/3)","50% (2/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 15% to 85%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.15,0.85],"\u0001"],["Total time per call on harder tasks",108.98,70.43,"seconds","109.0 s","70.4 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 25.7 s to 223.9 s; Claude Sonnet 5.5 4.32 s to 210.1 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",10,12,"Claude Code","Claude Code","range","minmax",[25.73,223.95],[4.32,210.08],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",12508,9287,"tokens","12,508","9,287","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",10,12,"Claude Code","Claude Code","range","minmax",[2965,26532],[407,27921],"\u0001"]]}}}