{"i":0,"comparison":{"slug":"claude-sonnet-5-5-vs-claude-opus-5-5","a":"claude-sonnet-5-5","b":"claude-opus-5-5","title":"Claude Sonnet 5.5 vs Claude Opus 5.5","seoTitle":"Claude Sonnet 5.5 vs Claude Opus 5.5: measured benchmarks","description":"Claude Sonnet 5.5 vs Claude Opus 5.5: 35 measured metrics from 8 studies, with sample sizes, intervals and every failure counted.","verdict":"Claude Sonnet 5.5 and Claude Opus 5.5 share 35 measured metrics and 31 list-price calculations from 10 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 16 ties and 50 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved on the same 3 SWE-bench Verified instances (interim)",0.3333,0.6667,"rate","33% (1/3)","67% (2/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 6% to 79%; Claude Opus 5.5 21% to 94%), so this sample cannot separate them.","swe-bench-opus-vs-sonnet",3,"swebench-opus-sonnet-resolved",3,3,"Agent · older builds · SWE-bench Verified, interim paired probe","Agent · new build · SWE-bench Verified, interim paired probe","ci95","ci95",[0.0615,0.7923],[0.2077,0.9385],"\u0001"],["List-price cost per attempt (calculation)",2.88,7.59,"usd","$2.88","$7.59","unclear","No interval or range was recorded for either side, so the gap ($2.88 vs $7.59, 2.6x) is not tested against run-to-run variation.","swe-bench-opus-vs-sonnet",3,"swebench-opus-sonnet-cost-per-attempt",3,3,"Agent · older builds · SWE-bench Verified, interim paired probe","Agent · new build · SWE-bench Verified, interim paired probe","\u0001","\u0001","\u0001","\u0001",true],["Worker time per attempt",9.37,20.26,"minutes","9.4 min","20.3 min","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.7 min to 15.0 min; Claude Opus 5.5 9.8 min to 25.3 min); the medians alone do not show a reliable difference. A range is not a confidence interval.","swe-bench-opus-vs-sonnet",3,"swebench-opus-sonnet-minutes",3,3,"Agent · older builds · SWE-bench Verified, interim paired probe","Agent · new build · SWE-bench Verified, interim paired probe","range","minmax",[4.74,15],[9.76,25.29],"\u0001"],["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,2.75,"seconds","2.31 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.17,7.73],[2.47,8.91],"\u0001"],["Time to first useful output",1.56,1.92,"seconds","1.56 s","1.92 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; Claude Opus 5.5 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.99,6.39],[1.56,7.23],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1401,1401,"tokens","1,401","1,401","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",685,680,"tokens","685","680","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",107,64,"tokens","107","64","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.0036,0.00688,"usd","$0.0036","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 $0.0034 to $0.010; Claude Opus 5.5 $0.0059 to $0.022); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00342,0.01021],[0.00592,0.02226],true],["List-price cost per passing answer (calculation)",0.00624,0.01009,"usd","$0.0062","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Opus 5.5 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Opus 5.5 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,9.18,"seconds","7.75 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[2.26,34.79],[4.24,27.21],"\u0001"],["Time to first useful output on hard tasks",5.95,6.78,"seconds","5.95 s","6.78 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.86 s to 30.6 s; Claude Opus 5.5 2.39 s to 21.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[0.86,30.57],[2.39,21.77],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1050,945,"tokens","1,050","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.02824,"usd","$0.014","$0.028","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.028) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; Claude Opus 5.5 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",23.1,56.9,"seconds","23.1 s","56.9 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; Claude Opus 5.5 29.8 s to 185.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","range","minmax",[18.7,44.5],[29.8,185.8],"\u0001"],["Tool calls per coding session",7.5,7.5,"calls","7.5","7.5","tie","Same value. More or fewer is not better by itself for this metric.","coding-agents-head-to-head",12,"coding-agents-tool-calls",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","range","minmax",[3,14],[5,14],"\u0001"],["List-price cost per passing coding session (calculation)",0.085,0.2229,"usd","$0.085","$0.22","unclear","No interval or range was recorded for either side, so the gap ($0.085 vs $0.22, 2.6x) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 81% to 100%; Claude Opus 5.5 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.97,9.18,"seconds","7.97 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 21.6 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[2.26,21.61],[4.24,27.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",1054,945,"tokens","1,054","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01398,0.02893,"usd","$0.014","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.029, 2.1x) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded)",0.135003,0.255057,"usd","$0.14","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.14 vs $0.26) is not tested against run-to-run variation.","caching-consistency",15,"caching-cost-with-without",15,15,"Claude Code · calculation: 5-turn cached sessions over a fixed ledger","Claude Code · calculation: 5-turn cached sessions over a fixed ledger","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price)",0.269788,0.544228,"usd","$0.27","$0.54","unclear","No interval or range was recorded for either side, so the gap ($0.27 vs $0.54, 2.0x) is not tested against run-to-run variation.","caching-consistency",15,"caching-cost-with-without",15,15,"Claude Code · calculation: 5-turn cached sessions over a fixed ledger","Claude Code · calculation: 5-turn cached sessions over a fixed ledger","\u0001","\u0001","\u0001","\u0001",true],["Time per turn: first turn vs later turns in a cached session (Turn 1 (writes the ledger to the cache))",1.64,1.9,"seconds","1.64 s","1.90 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.58 s to 1.79 s; Claude Opus 5.5 1.78 s to 4.36 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",3,"caching-latency-first-vs-later",3,3,"Claude Code · 5-turn cached sessions over a fixed ledger","Claude Code · 5-turn cached sessions over a fixed ledger","range","minmax",[1.58,1.79],[1.78,4.36],"\u0001"],["Time per turn: first turn vs later turns in a cached session (Turns 2-5 (read the ledger from the cache))",1.61,2.4,"seconds","1.61 s","2.40 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.35 s to 5.63 s; Claude Opus 5.5 1.63 s to 12.7 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",12,"caching-latency-first-vs-later",12,12,"Claude Code · 5-turn cached sessions over a fixed ledger","Claude Code · 5-turn cached sessions over a fixed ledger","range","minmax",[1.35,5.63],[1.63,12.67],"\u0001"],["Cost of a reused prefix with and without the cache, by session length (calculation): 1 turn",15.66,31.31,"usd","$15.66","$31.31","unclear","No interval or range was recorded for either side, so the gap ($15.66 vs $31.31) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 2 turns",31.32,62.62,"usd","$31.32","$62.62","unclear","No interval or range was recorded for either side, so the gap ($31.32 vs $62.62) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 3 turns",46.99,93.94,"usd","$46.99","$93.94","unclear","No interval or range was recorded for either side, so the gap ($46.99 vs $93.94) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 5 turns",78.31,156.56,"usd","$78.31","$156.56","unclear","No interval or range was recorded for either side, so the gap ($78.31 vs $156.56) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 10 turns",156.62,313.12,"usd","$156.62","$313.12","unclear","No interval or range was recorded for either side, so the gap ($156.62 vs $313.12) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 20 turns",313.24,626.24,"usd","$313.24","$626.24","unclear","No interval or range was recorded for either side, so the gap ($313.24 vs $626.24) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (No cache (the same either way))",156.62,313.12,"usd","$156.62","$313.12","unclear","No interval or range was recorded for either side, so the gap ($156.62 vs $313.12) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-session-split","\u0001","\u0001","","cache read $0.2 per M","\u0001","\u0001","\u0001","\u0001",true],["One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (One 10-turn session, 1-hour cache)",45.42,76.71,"usd","$45.42","$76.71","unclear","No interval or range was recorded for either side, so the gap ($45.42 vs $76.71) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-session-split","\u0001","\u0001","","cache read $0.2 per M","\u0001","\u0001","\u0001","\u0001",true],["One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix))",313.24,626.24,"usd","$313.24","$626.24","unclear","No interval or range was recorded for either side, so the gap ($313.24 vs $626.24) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-session-split","\u0001","\u0001","","cache read $0.2 per M","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",54.54,54.79,"percent","54.5%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[0,95.91],[29.92,95.6],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.006665,0.012528,"usd","$0.0067","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.003672,0.008003,"usd","$0.0037","$0.0080","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.004012,0.00771,"usd","$0.0040","$0.0077","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.006299,0.013104,"usd","$0.0063","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.013978,0.028925,"usd","$0.014","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.54,54.79,"percent","54.5%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[0,95.91],[29.92,95.6],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",0,0,"percent","0%","0%","tie","Same value. More or fewer is not better by itself for this metric.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[0,72.75],[0,93.33],true],["Time to first text: a 250-line answer, six models",1.96,1.97,"seconds","1.96 s","1.97 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; Claude Opus 5.5 1.70 s to 2.35 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[0.88,4.09],[1.7,2.35],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",231.7,155.5,"tokens","232","156","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[230.3,233],[154.6,156.4],true],["Output speed in characters per second after the first text (calculation)",517,347,"count","517","347","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Code","Claude Code","range","minmax",[513,519],[345,349],true],["Time to first text as the prompt grows: 1k",1.45,1.51,"seconds","1.45 s","1.51 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.23 s to 1.72 s; Claude Opus 5.5 1.46 s to 2.01 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.23,1.72],[1.46,2.01],true],["Time to first text as the prompt grows: 16k",1.78,1.74,"seconds","1.78 s","1.74 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.64 s to 2.11 s; Claude Opus 5.5 1.70 s to 2.97 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.64,2.11],[1.7,2.97],true],["Time to first text as the prompt grows: 64k",3.07,1.79,"seconds","3.07 s","1.79 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.38 s to 3.61 s; Claude Opus 5.5 1.72 s to 3.72 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.38,3.61],[1.72,3.72],true],["Total time per call by prompt size (1k prompt)",1.78,1.83,"seconds","1.78 s","1.83 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.57 s to 2.12 s; Claude Opus 5.5 1.82 s to 2.41 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[1.57,2.12],[1.82,2.41],"\u0001"],["Total time per call by prompt size (16k prompt)",2.1,2.36,"seconds","2.10 s","2.36 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.98 s to 2.48 s; Claude Opus 5.5 2.11 s to 3.40 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[1.98,2.48],[2.11,3.4],"\u0001"],["Total time per call by prompt size (64k prompt)",3.44,2.35,"seconds","3.44 s","2.35 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.74 s to 4.38 s; Claude Opus 5.5 2.26 s to 4.29 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[1.74,4.38],[2.26,4.29],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,0.5556,"rate","100% (9/9)","56% (5/9)","tie","The 95% intervals overlap (Claude Sonnet 5.5 70% to 100%; Claude Opus 5.5 27% to 81%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Claude Code","ci95","ci95",[0.7009,1],[0.2667,0.8112],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0.375,0.4167,"rate","38% (6/16)","42% (5/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; Claude Opus 5.5 19% to 68%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",16,12,"Claude Code","Claude Code","ci95","ci95",[0.1848,0.6136],[0.1933,0.6805],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0.375,0.5,"rate","38% (6/16)","50% (6/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; Claude Opus 5.5 25% to 75%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",16,12,"Claude Code","Claude Code","ci95","ci95",[0.1848,0.6136],[0.2538,0.7462],"\u0001"],["Calls that tried a tool although tools were off",0.3125,0.4167,"rate","31% (5/16)","42% (5/12)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-tool-attempts",16,12,"Claude Code","Claude Code","ci95","ci95",[0.1416,0.556],[0.1933,0.6805],"\u0001"],["Strict pass rate by task: 10x10 nonogram",1,1,"rate","100% (4/4)","100% (3/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 51% to 100%; Claude Opus 5.5 44% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",4,3,"Claude Code","Claude Code","ci95","ci95",[0.5101,1],[0.4385,1],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0,"rate","0% (0/4)","0% (0/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 0% to 49%; Claude Opus 5.5 0% to 56%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",4,3,"Claude Code","Claude Code","ci95","ci95",[0,0.4899],[0,0.5615],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,0.3333,"rate","0% (0/4)","33% (1/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 0% to 49%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",4,3,"Claude Code","Claude Code","ci95","ci95",[0,0.4899],[0.0615,0.7923],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0.5,0.3333,"rate","50% (2/4)","33% (1/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 15% to 85%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",4,3,"Claude Code","Claude Code","ci95","ci95",[0.15,0.85],[0.0615,0.7923],"\u0001"],["Total time per call on harder tasks",70.43,80.34,"seconds","70.4 s","80.3 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.32 s to 210.1 s; Claude Opus 5.5 3.82 s to 279.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",12,9,"Claude Code","Claude Code","range","minmax",[4.32,210.08],[3.82,279.5],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",9287,8420,"tokens","9,287","8,420","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",12,9,"Claude Code","Claude Code","range","minmax",[407,27921],[279,40044],"\u0001"],["List-price cost per strict pass on harder tasks (calculation)",0.23843,0.59333,"usd","$0.24","$0.59","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $0.59, 2.5x) is not tested against run-to-run variation.","harder-tasks-head-to-head","\u0001","harder-h2h-cost-per-pass",16,12,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}}}