{"i":10,"comparison":{"slug":"claude-opus-5-5-vs-gpt-6-1-sol-codex-cli","a":"claude-opus-5-5","b":"gpt-6-1-sol-codex-cli","title":"Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)","seoTitle":"Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): benchmarks","description":"Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): 31 measured metrics from 6 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","verdict":"Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) share 31 measured metrics and 19 list-price calculations from 7 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 12 ties and 38 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.71,5.6,"seconds","2.71 s","5.60 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.45 s to 11.8 s; GPT-6.1 Sol (Codex CLI) 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[2.45,11.78],[4.05,19.52],"\u0001"],["Time to first useful output",2.04,5.32,"seconds","2.04 s","5.32 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.40 s to 9.94 s; GPT-6.1 Sol (Codex CLI) 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[1.4,9.94],[3.64,16.37],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1463,6716,"tokens","1,463","6,716","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",619,5406,"tokens","619","5,406","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",78,42,"tokens","78","42","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00694,0.01047,"usd","$0.0069","$0.010","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 $0.0059 to $0.027; GPT-6.1 Sol (Codex CLI) $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[0.00592,0.02708],[0.0066,0.02812],true],["List-price cost per passing answer (calculation)",0.01049,0.01322,"usd","$0.010","$0.013","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.013) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",11.03,18.12,"seconds","11.0 s","18.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.63 s to 63.0 s; GPT-6.1 Sol (Codex CLI) 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","range","minmax",[3.63,63],[11.67,92.21],"\u0001"],["Time to first useful output on hard tasks",7.13,12.69,"seconds","7.13 s","12.7 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.15 s to 56.2 s; GPT-6.1 Sol (Codex CLI) 8.93 s to 75.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-first-useful-latency",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","range","minmax",[2.15,56.23],[8.93,75.91],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1052,436,"tokens","1,052","436","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head","\u0001","hard-h2h-output-tokens",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.03337,0.01514,"usd","$0.033","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.033 vs $0.015, 2.2x) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Opus 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",56.9,113.4,"seconds","56.9 s","113.4 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 29.8 s to 185.8 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[29.8,185.8],[78.5,221.9],"\u0001"],["Tool calls per coding session",7.5,12.5,"calls","7.5","12.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","coding-agents-head-to-head",12,"coding-agents-tool-calls",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[5,14],[8,18],"\u0001"],["List-price cost per passing coding session (calculation)",0.2229,0.0978,"usd","$0.22","$0.098","unclear","No interval or range was recorded for either side, so the gap ($0.22 vs $0.098, 2.3x) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",9.72,13.11,"seconds","9.72 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 4.78 s to 31.4 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","range","minmax",[4.78,31.36],[8.54,61.6],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",853,335,"tokens","853","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02947,0.02564,"usd","$0.029","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.026) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",54.43,57.01,"percent","54.4%","57%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-share",24,16,"Claude Code · effort high","Codex CLI · effort high","range","minmax",[36.14,96.23],[29.19,90.8],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.017969,0.004114,"usd","$0.018","$0.0041","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code · effort high","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.00799,0.002905,"usd","$0.0080","$0.0029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code · effort high","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.007407,0.008117,"usd","$0.0074","$0.0081","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code · effort high","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.01344,0.002273,"usd","$0.013","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.029475,0.025637,"usd","$0.029","$0.026","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.43,57.01,"percent","54.4%","57%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-short-vs-hard",24,16,"Claude Code · effort high","Codex CLI · effort high","range","minmax",[36.14,96.23],[29.19,90.8],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",43.59,58.06,"percent","43.6%","58.1%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code · effort high","Codex CLI · effort high","range","minmax",[0,93.33],[0,75.76],true],["Time to first text: a 250-line answer, six models",1.97,3.52,"seconds","1.97 s","3.52 s","unclear","Only 4 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.35 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s), but 4 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[1.7,2.35],[2.75,4.42],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",155.5,79.6,"tokens","156","80","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[154.6,156.4],[71.6,80.5],true],["Output speed in characters per second after the first text (calculation)",347,323,"count","347","323","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[345,349],[291,327],true],["Time to first text as the prompt grows: 1k",1.51,3.36,"seconds","1.51 s","3.36 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.46 s to 2.01 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.46,2.01],[3.36,4.75],true],["Time to first text as the prompt grows: 16k",1.74,4.02,"seconds","1.74 s","4.02 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.97 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.7,2.97],[3.3,4.28],true],["Time to first text as the prompt grows: 64k",1.79,3.93,"seconds","1.79 s","3.93 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.72 s to 3.72 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.72,3.72],[3.42,4.38],true],["Total time per call by prompt size (1k prompt)",1.83,3.43,"seconds","1.83 s","3.43 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.82 s to 2.41 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.82,2.41],[3.43,4.92],"\u0001"],["Total time per call by prompt size (16k prompt)",2.36,4.14,"seconds","2.36 s","4.14 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 2.11 s to 3.40 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.11,3.4],[3.96,4.68],"\u0001"],["Total time per call by prompt size (64k prompt)",2.35,3.96,"seconds","2.35 s","3.96 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.26 s to 4.29 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.26,4.29],[3.47,4.44],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",0.5556,1,"rate","56% (5/9)","100% (9/9)","tie","The 95% intervals overlap (Claude Opus 5.5 27% to 81%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Codex CLI · effort low","ci95","ci95",[0.2667,0.8112],[0.7009,1],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0.4167,0.6875,"rate","42% (5/12)","69% (11/16)","tie","The 95% intervals overlap (Claude Opus 5.5 19% to 68%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.1933,0.6805],[0.444,0.8584],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0.5,0.6875,"rate","50% (6/12)","69% (11/16)","tie","The 95% intervals overlap (Claude Opus 5.5 25% to 75%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.2538,0.7462],[0.444,0.8584],"\u0001"],["Calls that tried a tool although tools were off",0.4167,0,"rate","42% (5/12)","0% (0/16)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-tool-attempts",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.1933,0.6805],[0,0.1936],"\u0001"],["Strict pass rate by task: 10x10 nonogram",1,0.75,"rate","100% (3/3)","75% (3/4)","tie","The 95% intervals overlap (Claude Opus 5.5 44% to 100%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.4385,1],[0.3006,0.9544],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0.25,"rate","0% (0/3)","25% (1/4)","tie","The 95% intervals overlap (Claude Opus 5.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.0456,0.6994],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0.3333,1,"rate","33% (1/3)","100% (4/4)","tie","The 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 51% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.0615,0.7923],[0.5101,1],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0.3333,0.75,"rate","33% (1/3)","75% (3/4)","tie","The 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.0615,0.7923],[0.3006,0.9544],"\u0001"],["Total time per call on harder tasks",80.34,120.24,"seconds","80.3 s","120.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.82 s to 279.5 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",9,13,"Claude Code","Codex CLI · effort medium","range","minmax",[3.82,279.5],[46.24,273.46],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",8420,4994,"tokens","8,420","4,994","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",9,13,"Claude Code","Codex CLI · effort medium","range","minmax",[279,40044],[2099,13413],"\u0001"],["List-price cost per strict pass on harder tasks (calculation)",0.59333,0.08293,"usd","$0.59","$0.083","unclear","No interval or range was recorded for either side, so the gap ($0.59 vs $0.083, 7.2x) is not tested against run-to-run variation.","harder-tasks-head-to-head","\u0001","harder-h2h-cost-per-pass",12,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true]]}}}