{"i":37,"comparison":{"slug":"claude-haiku-4-5-vs-claude-fable-5-1","a":"claude-haiku-4-5","b":"claude-fable-5-1","title":"Claude Haiku 4.5 vs Claude Fable 5.1","seoTitle":"Claude Haiku 4.5 vs Claude Fable 5.1: measured benchmarks","description":"Claude Haiku 4.5 vs Claude Fable 5.1: 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","verdict":"Claude Haiku 4.5 and Claude Fable 5.1 share 12 measured metrics and 11 list-price calculations from 4 studies. Claude Fable 5.1 leads on 2 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24). On those rows the 95% intervals do not overlap. The other rows are 1 tie and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Fable 5.1 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",4.43,1.94,"seconds","4.43 s","1.94 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Fable 5.1 1.41 s to 9.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[3.16,23.57],[1.41,9.83],"\u0001"],["Time to first useful output",3.63,1.2,"seconds","3.63 s","1.20 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; Claude Fable 5.1 0.95 s to 7.90 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.78,22.27],[0.95,7.9],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",0,2760,"tokens","0","2,760","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",3790,473,"tokens","3,790","473","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",367,64,"tokens","367","64","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00566,0.00987,"usd","$0.0057","$0.0099","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; Claude Fable 5.1 $0.0049 to $0.058); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00513,0.01804],[0.0049,0.05843],true],["List-price cost per passing answer (calculation)",0.00836,0.02054,"usd","$0.0084","$0.021","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.021, 2.5x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Fable 5.1 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Fable 5.1 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,16.13,"seconds","39.0 s","16.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Fable 5.1 4.46 s to 90.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[15.27,75.13],[4.46,90],"\u0001"],["Time to first useful output on hard tasks",35.54,11.63,"seconds","35.5 s","11.6 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; Claude Fable 5.1 2.00 s to 85.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[12.88,70.31],[2,85.33],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",5064,1366,"tokens","5,064","1,366","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.09331,"usd","$0.067","$0.093","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.093) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",91.68,64.24,"percent","91.7%","64.2%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[23.44,97.19],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.024492,0.053696,"usd","$0.024","$0.054","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.001795,0.018702,"usd","$0.0018","$0.019","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.00451,0.02091,"usd","$0.0045","$0.021","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",91.68,64.24,"percent","91.7%","64.2%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[23.44,97.19],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",90.19,0,"percent","90.2%","0%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[73.1,97.59],[0,74.01],true],["Time to first text: a 250-line answer, six models",4,4.43,"seconds","4.00 s","4.43 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; Claude Fable 5.1 2.27 s to 4.64 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[2.84,6.38],[2.27,4.64],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",153.2,122.6,"tokens","153","123","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[152.6,216.1],[120.9,131.4],true],["Output speed in characters per second after the first text (calculation)",547,273,"count","547","273","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","\u0001","speed-anatomy-chars-per-second",3,4,"Claude Code","Claude Code","range","minmax",[546,548],[270,293],true]]}}}