{"i":5,"comparison":{"slug":"claude-code-cli-vs-codex-cli","a":"claude-code-cli","b":"codex-cli","title":"Claude Code vs Codex CLI","seoTitle":"Claude Code vs Codex CLI: measured benchmarks","description":"Claude Code vs Codex CLI: 66 measured metrics from 11 studies (Pass rate on five validated tasks; Total time per call; more), with sample sizes and intervals.","verdict":"Claude Code and Codex CLI share 66 measured metrics and 22 list-price calculations from 12 studies. Claude Code leads on 7 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 4 more. Codex CLI leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 31 ties and 49 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Each run pairs a CLI with a model, so these rows cannot separate the CLI from the model; the contexts name both. Some rows rest on small samples (n = 2 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Code 55% to 93%; Codex CLI 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,5.65,"seconds","2.31 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.17 s to 7.73 s; Codex CLI 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","range","minmax",[2.17,7.73],[4.1,25.46],"\u0001"],["Time to first useful output",1.56,5.05,"seconds","1.56 s","5.05 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 0.99 s to 6.39 s; Codex CLI 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","range","minmax",[0.99,6.39],[3.36,17.82],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1401,5180,"tokens","1,401","5,180","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",685,6943,"tokens","685","6,943","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",107,42,"tokens","107","42","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.0036,0.01018,"usd","$0.0036","$0.010","unclear","The run ranges (fastest to slowest) overlap (Claude Code $0.0034 to $0.010; Codex CLI $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","range","minmax",[0.00342,0.01021],[0.0054,0.02686],true],["List-price cost per passing answer (calculation)",0.00624,0.01564,"usd","$0.0062","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,13.11,"seconds","7.75 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","range","minmax",[2.26,34.79],[8.54,61.6],"\u0001"],["Time to first useful output on hard tasks",5.95,10.23,"seconds","5.95 s","10.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 0.86 s to 30.6 s; Codex CLI 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-first-useful-latency",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","range","minmax",[0.86,30.57],[6.09,40.41],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1050,335,"tokens","1,050","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head","\u0001","hard-h2h-output-tokens",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",23.1,113.4,"seconds","23.1 s","113.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[18.7,44.5],[78.5,221.9],"\u0001"],["Tool calls per coding session",7.5,12.5,"calls","7.5","12.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","coding-agents-head-to-head",12,"coding-agents-tool-calls",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[3,14],[8,18],"\u0001"],["List-price cost per passing coding session (calculation)",0.085,0.0978,"usd","$0.085","$0.098","unclear","No interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Code 81% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.63,13.11,"seconds","7.63 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.71 s to 24.0 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","range","minmax",[2.71,24.01],[8.54,61.6],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",770,335,"tokens","770","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01352,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: how many different answers (Exact number)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (JSON object)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (Code fix)",3,6,"count","3","6","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: time per call (Exact number)",6.89,13.38,"seconds","6.89 s","13.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","range","minmax",[5.81,7.81],[12.29,17.97],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",2.89,6.42,"seconds","2.89 s","6.42 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.68 s to 5.30 s; Codex CLI 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","range","minmax",[2.68,5.3],[5.25,8.26],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",2.67,11.29,"seconds","2.67 s","11.3 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","range","minmax",[2.32,4.34],[9.08,14.85],"\u0001"],["CLI start-up tax on a one-word answer (First output event)",563,489,"ms","563 ms","489 ms","unclear","The run ranges (fastest to slowest) overlap (Claude Code 519 ms to 726 ms; Codex CLI 354 ms to 1,304 ms); the medians alone do not show a reliable difference. A range is not a confidence interval.","routing-overhead",5,"cli-startup-tax",5,5,"Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","default model · CLI start-up, one-word prompt, 5 runs","range","minmax",[519,726],[354,1304],"\u0001"],["CLI start-up tax on a one-word answer (First model output)",1461,5059,"ms","1,461 ms","5,059 ms","a","The run ranges (fastest to slowest) do not overlap (Claude Code 1,206 ms to 2,308 ms; Codex CLI 4,391 ms to 5,478 ms). A range is not a confidence interval. Samples are small (5 runs per side).","routing-overhead",5,"cli-startup-tax",5,5,"Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","default model · CLI start-up, one-word prompt, 5 runs","range","minmax",[1206,2308],[4391,5478],"\u0001"],["CLI start-up tax on a one-word answer (Total wall time)",2529,5999,"ms","2,529 ms","5,999 ms","a","The run ranges (fastest to slowest) do not overlap (Claude Code 2,273 ms to 3,382 ms; Codex CLI 5,367 ms to 6,506 ms). A range is not a confidence interval. Samples are small (5 runs per side).","routing-overhead",5,"cli-startup-tax",5,5,"Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","default model · CLI start-up, one-word prompt, 5 runs","range","minmax",[2273,3382],[5367,6506],"\u0001"],["Input tokens a CLI sends for a one-word answer",6761,17051,"tokens","6,761","17,051","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","routing-overhead",5,"cli-startup-input-tokens",5,5,"Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","default model · CLI start-up, one-word prompt, 5 runs","\u0001","\u0001","\u0001","\u0001","\u0001"],["Repairing a scheduler: Claude Code vs Codex vs API (Total time)",15,61.16,"seconds","15.0 s","61.2 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 13.9 s to 15.9 s; Codex CLI 59.9 s to 69.5 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs","GPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[13.89,15.89],[59.9,69.51],"\u0001"],["Repairing a scheduler: Claude Code vs Codex vs API (First useful output)",7.55,15.56,"seconds","7.55 s","15.6 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 6.77 s to 7.63 s; Codex CLI 13.7 s to 23.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs","GPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[6.77,7.63],[13.65,23.04],"\u0001"],["Output tokens to repair the scheduler (Output tokens)",2227,1181,"tokens","2,227","1,181","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",3,"scheduler-repair-output-tokens",3,3,"Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs","GPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs","\u0001","\u0001","\u0001","\u0001","\u0001"],["Strict pass rate: single call vs agent loop on eight hard tasks",1,0.625,"rate","100% (24/24)","63% (10/16)","a","The 95% intervals do not overlap (Claude Code 86% to 100%; Codex CLI 39% to 82%).","single-call-vs-agent-loop","\u0001","agent-loop-pass-rate",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.862,1],[0.3864,0.8152],"\u0001"],["Strict passes per task: single call vs agent loop: Interval merge fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001"],["Strict passes per task: single call vs agent loop: DST day-length fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001"],["Strict passes per task: single call vs agent loop: CSV parser",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001"],["Strict passes per task: single call vs agent loop: Event-loop order",1,0,"rate","100% (3/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0,0.6576],"\u0001"],["Strict passes per task: single call vs agent loop: Room schedule",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001"],["Strict passes per task: single call vs agent loop: SemVer regex",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001"],["Strict passes per task: single call vs agent loop: Money refactor",1,0,"rate","100% (3/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0,0.6576],"\u0001"],["Strict passes per task: single call vs agent loop: SQL report",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001"],["Total time per attempt: single call vs agent loop",7.75,5.16,"seconds","7.75 s","5.16 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","single-call-vs-agent-loop","\u0001","agent-loop-total-time",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","range","minmax",[2.26,34.79],[3.59,11.32],"\u0001"],["Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))",2281,11582,"tokens","2,281","11,582","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","\u0001","agent-loop-tokens",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","range","minmax",[2234,2669],[11526,11818],"\u0001"],["Tokens per attempt: single call vs agent loop (Output tokens)",1050,345,"tokens","1,050","345","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","\u0001","agent-loop-tokens",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","range","minmax",[176,3895],[36,634],"\u0001"],["Tool calls per agent-loop attempt",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","single-call-vs-agent-loop","\u0001","agent-loop-tool-calls",16,14,"Claude Sonnet 5.5 · agent loop","GPT-6 Luna · agent loop","range","minmax",[0,3],[0,1],"\u0001"],["List-price cost per strict pass: single call vs agent loop (calculation)",0.01435,0.00116,"usd","$0.014","$0.0012","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation.","single-call-vs-agent-loop","\u0001","agent-loop-cost-per-pass",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","\u0001","\u0001","\u0001","\u0001",true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions",12,"structured-output-pass-rate",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","ci95","ci95",[0.7575,1],[0.7575,1],true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions",12,"structured-output-pass-rate",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","ci95","ci95",[0.7575,1],[0.7575,1],true],["What each call produced: strict pass, format miss, wrong values or error (Strict pass)",12,12,"count","12","12","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Format miss)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Wrong values)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Error)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per call, instructions vs schema mode",3.52,6.21,"seconds","3.52 s","6.21 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 2.67 s to 4.12 s; Codex CLI 4.20 s to 12.3 s). A range is not a confidence interval.","json-schema-vs-instructions",12,"structured-output-time",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","range","minmax",[2.67,4.12],[4.2,12.27],"\u0001"],["Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)",368,117,"tokens","368","117","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions",12,"structured-output-tokens",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Reasoning share of output tokens per call on hard tasks (calculation)",54.54,46.33,"percent","54.5%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-share",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[0,95.91],[11.42,86.85],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.006665,0.002273,"usd","$0.0067","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.003672,0.003025,"usd","$0.0037","$0.0030","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.004012,0.020339,"usd","$0.0040","$0.020","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005946,0.002273,"usd","$0.0059","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Sonnet 5.5 · effort medium","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.01352,0.025637,"usd","$0.014","$0.026","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Sonnet 5.5 · effort medium","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.54,46.33,"percent","54.5%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-short-vs-hard",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[0,95.91],[11.42,86.85],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",0,41.05,"percent","0%","41%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[0,72.75],[0,71.43],true],["Time to first text: a 250-line answer, six models",1.96,3.52,"seconds","1.96 s","3.52 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 0.88 s to 4.09 s; Codex CLI 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[0.88,4.09],[2.75,4.42],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",231.7,79.6,"tokens","232","80","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[230.3,233],[71.6,80.5],true],["Output speed in characters per second after the first text (calculation)",517,323,"count","517","323","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[513,519],[291,327],true],["Time to first text as the prompt grows: 1k",1.45,3.36,"seconds","1.45 s","3.36 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.23 s to 1.72 s; Codex CLI 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.23,1.72],[3.36,4.75],true],["Time to first text as the prompt grows: 16k",1.78,4.02,"seconds","1.78 s","4.02 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.64 s to 2.11 s; Codex CLI 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.64,2.11],[3.3,4.28],true],["Time to first text as the prompt grows: 64k",3.07,3.93,"seconds","3.07 s","3.93 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 1.38 s to 3.61 s; Codex CLI 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.38,3.61],[3.42,4.38],true],["Total time per call by prompt size (1k prompt)",1.78,3.43,"seconds","1.78 s","3.43 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.57 s to 2.12 s; Codex CLI 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.57,2.12],[3.43,4.92],"\u0001"],["Total time per call by prompt size (16k prompt)",2.1,4.14,"seconds","2.10 s","4.14 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.98 s to 2.48 s; Codex CLI 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.98,2.48],[3.96,4.68],"\u0001"],["Total time per call by prompt size (64k prompt)",3.44,3.96,"seconds","3.44 s","3.96 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 1.74 s to 4.38 s; Codex CLI 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.74,4.38],[3.47,4.44],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,1,"rate","100% (9/9)","100% (9/9)","tie","The 95% intervals overlap (Claude Code 70% to 100%; Codex CLI 70% to 100%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","ci95","ci95",[0.7009,1],[0.7009,1],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0.375,0.6875,"rate","38% (6/16)","69% (11/16)","tie","The 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head",16,"harder-h2h-pass-rate",16,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.1848,0.6136],[0.444,0.8584],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0.375,0.6875,"rate","38% (6/16)","69% (11/16)","tie","The 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head",16,"harder-h2h-pass-rate",16,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.1848,0.6136],[0.444,0.8584],"\u0001"],["Calls that tried a tool although tools were off",0.3125,0,"rate","31% (5/16)","0% (0/16)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head",16,"harder-h2h-tool-attempts",16,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.1416,0.556],[0,0.1936],"\u0001"],["Strict pass rate by task: 10x10 nonogram",1,0.75,"rate","100% (4/4)","75% (3/4)","tie","The 95% intervals overlap (Claude Code 51% to 100%; Codex CLI 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.5101,1],[0.3006,0.9544],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0.25,"rate","0% (0/4)","25% (1/4)","tie","The 95% intervals overlap (Claude Code 0% to 49%; Codex CLI 5% to 70%), so this sample cannot separate them.","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0,0.4899],[0.0456,0.6994],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,1,"rate","0% (0/4)","100% (4/4)","b","The 95% intervals do not overlap (Claude Code 0% to 49%; Codex CLI 51% to 100%).","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0,0.4899],[0.5101,1],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0.5,0.75,"rate","50% (2/4)","75% (3/4)","tie","The 95% intervals overlap (Claude Code 15% to 85%; Codex CLI 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.15,0.85],[0.3006,0.9544],"\u0001"],["Total time per call on harder tasks",70.43,120.24,"seconds","70.4 s","120.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 4.32 s to 210.1 s; Codex CLI 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",12,13,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[4.32,210.08],[46.24,273.46],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",9287,4994,"tokens","9,287","4,994","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",12,13,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[407,27921],[2099,13413],"\u0001"],["List-price cost per strict pass on harder tasks (calculation)",0.23843,0.08293,"usd","$0.24","$0.083","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $0.083, 2.9x) is not tested against run-to-run variation.","harder-tasks-head-to-head",16,"harder-h2h-cost-per-pass",16,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true]]}}}