{"i":35,"comparison":{"slug":"claude-sonnet-5-5-vs-gpt-6-luna-codex-cli","a":"claude-sonnet-5-5","b":"gpt-6-luna-codex-cli","title":"Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI)","seoTitle":"Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI): benchmarks","description":"Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI): 14 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","verdict":"Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. Claude Sonnet 5.5 leads on 1 row: Strict pass rate: single call vs agent loop on eight hard tasks, 100% (24/24) vs 63% (10/16). On those rows the 95% intervals do not overlap. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation","n"],"$r":[["Strict pass rate: single call vs agent loop on eight hard tasks",1,0.625,"rate","100% (24/24)","63% (10/16)","a","The 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%).","single-call-vs-agent-loop","agent-loop-pass-rate",24,16,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.862,1],[0.3864,0.8152],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Interval merge fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: DST day-length fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: CSV parser",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Event-loop order",1,0,"rate","100% (3/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0,0.6576],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Room schedule",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: SemVer regex",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Money refactor",1,0,"rate","100% (3/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0,0.6576],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: SQL report",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Total time per attempt: single call vs agent loop",7.75,5.16,"seconds","7.75 s","5.16 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","single-call-vs-agent-loop","agent-loop-total-time",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[2.26,34.79],[3.59,11.32],"\u0001","\u0001"],["Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))",2281,11582,"tokens","2,281","11,582","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tokens",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[2234,2669],[11526,11818],"\u0001","\u0001"],["Tokens per attempt: single call vs agent loop (Output tokens)",1050,345,"tokens","1,050","345","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tokens",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[176,3895],[36,634],"\u0001","\u0001"],["Tool calls per agent-loop attempt",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","single-call-vs-agent-loop","agent-loop-tool-calls",16,14,"Claude Code · agent loop","Codex CLI · agent loop","range","minmax",[0,3],[0,1],"\u0001","\u0001"],["List-price cost per strict pass: single call vs agent loop (calculation)",0.01435,0.00116,"usd","$0.014","$0.0012","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation.","single-call-vs-agent-loop","agent-loop-cost-per-pass",24,16,"Claude Code · single call","Codex CLI · single call","\u0001","\u0001","\u0001","\u0001",true,"\u0001"],["Time to first text: a 250-line answer, six models",1.96,3.3,"seconds","1.96 s","3.30 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy","speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[0.88,4.09],[3.19,3.47],"\u0001",4],["Output speed after the first text: visible tokens per second (calculation)",231.7,129.1,"tokens","232","129","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[230.3,233],[55.5,259.1],true,4],["Output speed in characters per second after the first text (calculation)",517,524,"count","517","524","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","speed-anatomy-chars-per-second",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[513,519],[225,1052],true,4]]}}}