{"i":11,"comparison":{"slug":"claude-haiku-4-5-vs-gpt-6-1-sol-codex-cli","a":"claude-haiku-4-5","b":"gpt-6-1-sol-codex-cli","title":"Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)","seoTitle":"Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI): benchmarks","description":"Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI): 40 measured metrics from 6 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","verdict":"Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) share 40 measured metrics and 16 list-price calculations from 7 studies. Claude Haiku 4.5 leads on 2 rows: Same prompt, 10 times: time per call (Exact number), 5.06 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 5.95 s vs 11.3 s. GPT-6.1 Sol (Codex CLI) leads on 6 rows: Pass rate on eight hard tasks (Strict pass), 100% (16/16) vs 46% (11/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); Same prompt, 10 times: strict pass rate (JSON object), 100% (10/10) vs 10% (1/10); and 3 more. On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 13 ties and 35 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",4.43,5.65,"seconds","4.43 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[3.16,23.57],[4.1,25.46],"\u0001"],["Time to first useful output",3.63,5.05,"seconds","3.63 s","5.05 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[2.78,22.27],[3.36,17.82],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",0,5180,"tokens","0","5,180","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",3790,6943,"tokens","3,790","6,943","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",367,42,"tokens","367","42","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00566,0.01018,"usd","$0.0057","$0.010","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.00513,0.01804],[0.0054,0.02686],true],["List-price cost per passing answer (calculation)",0.00836,0.01564,"usd","$0.0084","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.016) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (16/16)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; GPT-6.1 Sol (Codex CLI) 81% to 100%).","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Haiku 4.5 47% to 82%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,13.11,"seconds","39.0 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[15.27,75.13],[8.54,61.6],"\u0001"],["Time to first useful output on hard tasks",35.54,10.23,"seconds","35.5 s","10.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-first-useful-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[12.88,70.31],[6.09,40.41],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",5064,335,"tokens","5,064","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head","\u0001","hard-h2h-output-tokens",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.02564,"usd","$0.067","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.026, 2.6x) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",0,1,"rate","0% (0/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; GPT-6.1 Sol (Codex CLI) 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0,0.2775],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",0.1,1,"rate","10% (1/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; GPT-6.1 Sol (Codex CLI) 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.0179,0.4042],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: how many different answers (Exact number)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (JSON object)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (Code fix)",6,6,"count","6","6","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: time per call (Exact number)",5.06,13.38,"seconds","5.06 s","13.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.42 s to 6.20 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[4.42,6.2],[12.29,17.97],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",7.03,6.42,"seconds","7.03 s","6.42 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.28 s to 12.3 s; GPT-6.1 Sol (Codex CLI) 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[5.28,12.27],[5.25,8.26],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",5.95,11.29,"seconds","5.95 s","11.3 s","a","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[4.89,7.33],[9.08,14.85],"\u0001"],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",0,1,"rate","0% (0/24)","100% (12/12)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; GPT-6.1 Sol (Codex CLI) 76% to 100%).","json-schema-vs-instructions","\u0001","structured-output-pass-rate",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","ci95","ci95",[0,0.138],[0.7575,1],true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))",0.7083,1,"rate","71% (17/24)","100% (12/12)","tie","The 95% intervals overlap (Claude Haiku 4.5 51% to 85%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions","\u0001","structured-output-pass-rate",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","ci95","ci95",[0.5083,0.8509],[0.7575,1],true],["What each call produced: strict pass, format miss, wrong values or error (Strict pass)",0,12,"count","0","12","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Format miss)",17,0,"count","17","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Wrong values)",7,0,"count","7","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Error)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per call, instructions vs schema mode",9.52,6.21,"seconds","9.52 s","6.21 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.67 s to 17.0 s; GPT-6.1 Sol (Codex CLI) 4.20 s to 12.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","json-schema-vs-instructions","\u0001","structured-output-time",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","range","minmax",[5.67,17],[4.2,12.27],"\u0001"],["Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)",1128,117,"tokens","1,128","117","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-tokens",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Reasoning share of output tokens per call on hard tasks (calculation)",91.68,46.33,"percent","91.7%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-share",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[76.46,99.27],[11.42,86.85],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.024492,0.002273,"usd","$0.024","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.001795,0.003025,"usd","$0.0018","$0.0030","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.00451,0.020339,"usd","$0.0045","$0.020","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",91.68,46.33,"percent","91.7%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-short-vs-hard",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[76.46,99.27],[11.42,86.85],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",90.19,41.05,"percent","90.2%","41%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Codex CLI · effort medium","range","minmax",[73.1,97.59],[0,71.43],true],["Time to first text: a 250-line answer, six models",4,3.52,"seconds","4.00 s","3.52 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[2.84,6.38],[2.75,4.42],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",153.2,79.6,"tokens","153","80","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[152.6,216.1],[71.6,80.5],true],["Output speed in characters per second after the first text (calculation)",547,323,"count","547","323","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","\u0001","speed-anatomy-chars-per-second",3,4,"Claude Code","Codex CLI · effort low","range","minmax",[546,548],[291,327],true],["Time to first text as the prompt grows: 1k",1.93,3.36,"seconds","1.93 s","3.36 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 1.85 s to 2.04 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.85,2.04],[3.36,4.75],true],["Time to first text as the prompt grows: 16k",2.27,4.02,"seconds","2.27 s","4.02 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.47 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.22,2.47],[3.3,4.28],true],["Time to first text as the prompt grows: 64k",2.78,3.93,"seconds","2.78 s","3.93 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.45 s to 2.89 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.45,2.89],[3.42,4.38],true],["Total time per call by prompt size (1k prompt)",2.34,3.43,"seconds","2.34 s","3.43 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.46 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.22,2.46],[3.43,4.92],"\u0001"],["Total time per call by prompt size (16k prompt)",2.79,4.14,"seconds","2.79 s","4.14 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.58 s to 2.84 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.58,2.84],[3.96,4.68],"\u0001"],["Total time per call by prompt size (64k prompt)",3.13,3.96,"seconds","3.13 s","3.96 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.84 s to 3.28 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.84,3.28],[3.47,4.44],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,1,"rate","100% (9/9)","100% (9/9)","tie","The 95% intervals overlap (Claude Haiku 4.5 70% to 100%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Codex CLI · effort low","ci95","ci95",[0.7009,1],[0.7009,1],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0,0.6875,"rate","0% (0/12)","69% (11/16)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%).","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.2425],[0.444,0.8584],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0,0.6875,"rate","0% (0/12)","69% (11/16)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%).","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.2425],[0.444,0.8584],"\u0001"],["Calls that tried a tool although tools were off",0.0833,0,"rate","8% (1/12)","0% (0/16)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-tool-attempts",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.0149,0.3539],[0,0.1936],"\u0001"],["Strict pass rate by task: 10x10 nonogram",0,0.75,"rate","0% (0/3)","75% (3/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.3006,0.9544],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0.25,"rate","0% (0/3)","25% (1/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.0456,0.6994],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,1,"rate","0% (0/3)","100% (4/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 51% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.5101,1],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0,0.75,"rate","0% (0/3)","75% (3/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.3006,0.9544],"\u0001"],["Total time per call on harder tasks",108.98,120.24,"seconds","109.0 s","120.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 25.7 s to 223.9 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",10,13,"Claude Code","Codex CLI · effort medium","range","minmax",[25.73,223.95],[46.24,273.46],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",12508,4994,"tokens","12,508","4,994","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",10,13,"Claude Code","Codex CLI · effort medium","range","minmax",[2965,26532],[2099,13413],"\u0001"]]}}}