{"i":34,"comparison":{"slug":"claude-haiku-4-5-vs-gpt-6-luna-codex-cli","a":"claude-haiku-4-5","b":"gpt-6-luna-codex-cli","title":"Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI)","seoTitle":"Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI): benchmarks","description":"Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI): 14 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","verdict":"Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. GPT-6 Luna (Codex CLI) leads on 1 row: Total time per attempt: single call vs agent loop, 5.16 s vs 39.0 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation","n"],"$r":[["Strict pass rate: single call vs agent loop on eight hard tasks",0.4583,0.625,"rate","46% (11/24)","63% (10/16)","tie","The 95% intervals overlap (Claude Haiku 4.5 28% to 65%; GPT-6 Luna (Codex CLI) 39% to 82%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-pass-rate",24,16,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.2789,0.6493],[0.3864,0.8152],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Interval merge fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: DST day-length fix",0.3333,1,"rate","33% (1/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 6% to 79%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.0615,0.7923],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: CSV parser",0.6667,1,"rate","67% (2/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.2077,0.9385],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Event-loop order",0,0,"rate","0% (0/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0,0.5615],[0,0.6576],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Room schedule",0,0.5,"rate","0% (0/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0,0.5615],[0.0945,0.9055],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: SemVer regex",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Money refactor",0.6667,0,"rate","67% (2/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.2077,0.9385],[0,0.6576],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: SQL report",0,1,"rate","0% (0/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0,0.5615],[0.3424,1],"\u0001","\u0001"],["Total time per attempt: single call vs agent loop",39.01,5.16,"seconds","39.0 s","5.16 s","b","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s). A range is not a confidence interval.","single-call-vs-agent-loop","agent-loop-total-time",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[15.27,75.13],[3.59,11.32],"\u0001","\u0001"],["Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))",3941,11582,"tokens","3,941","11,582","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tokens",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[3879,4221],[11526,11818],"\u0001","\u0001"],["Tokens per attempt: single call vs agent loop (Output tokens)",5064,345,"tokens","5,064","345","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tokens",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[1899,9321],[36,634],"\u0001","\u0001"],["Tool calls per agent-loop attempt",3,0,"count","3","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tool-calls",24,14,"Claude Code · agent loop","Codex CLI · agent loop","range","minmax",[2,18],[0,1],"\u0001","\u0001"],["List-price cost per strict pass: single call vs agent loop (calculation)",0.0672,0.00116,"usd","$0.067","$0.0012","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.0012, 58x) is not tested against run-to-run variation.","single-call-vs-agent-loop","agent-loop-cost-per-pass",24,16,"Claude Code · single call","Codex CLI · single call","\u0001","\u0001","\u0001","\u0001",true,"\u0001"],["Time to first text: a 250-line answer, six models",4,3.3,"seconds","4.00 s","3.30 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy","speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[2.84,6.38],[3.19,3.47],"\u0001",4],["Output speed after the first text: visible tokens per second (calculation)",153.2,129.1,"tokens","153","129","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[152.6,216.1],[55.5,259.1],true,4],["Output speed in characters per second after the first text (calculation)",547,524,"count","547","524","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","speed-anatomy-chars-per-second",3,4,"Claude Code","Codex CLI · effort low","range","minmax",[546,548],[225,1052],true,"\u0001"]]}}}