{"i":5,"study":{"slug":"coding-agents-head-to-head","title":"Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks","seoTitle":"Claude Code vs Codex CLI: 6 coding tasks, hidden tests","description":"36 graded sessions: Claude Code with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol. All passed every hidden test; time, tool calls and diffs differ.","question":"On small real repository tasks graded by hidden tests, how do coding-agent CLIs compare when they run with their normal file and shell tools?","answer":"All 3 agents passed every hidden check in every session (12/12, 12/12, 12/12; 95% Wilson 76–100% each), so this task set cannot separate them on quality. Sonnet 5.5 in Claude Code was fastest (median 23.1 s; its sessions took 18.7 s to 44.5 s), Opus 5.5 in Claude Code took a median 56.9 s, GPT-6.1 Sol in Codex CLI took a median 113.4 s; the run ranges of Sonnet 5.5 in Claude Code and GPT-6.1 Sol in Codex CLI do not overlap. Codex made more tool calls (median 12.5 vs 7.5 and 7.5) and larger diffs (median 94 lines vs 45 and 77.5, mostly added tests), and every Codex session also followed the tester’s global AGENTS.md (10 of 12 wrote a work log nobody asked for), so its time and diff include extra work. Opus used 1.9× the output tokens of Sonnet in the same CLI (medians). Gemini CLI was not run: it needed a browser login.","date":"2026-10-06","updated":"2026-10-06","tags":["claude-code","codex-cli","coding-agents","hidden-tests","head-to-head"],"caveats":["Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.","Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.","Each run pairs a CLI with a model (Claude Code with Claude models, Codex CLI with GPT-6.1 Sol), so the results cannot separate the CLI from the model.","n = 12 sessions per agent (2 per task). Time ranges are the fastest and slowest sessions, not confidence intervals; the two lanes shared one machine.","Tokens are as each CLI reports them, with different tokenizers and context handling: compare tokens inside Claude Code (Sonnet vs Opus), not across vendors. Costs are list-price calculations on subscription sessions, not invoices."],"sourceIds":["agent-coding-agents","calc-repricing","price-anthropic","price-openai"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["coding-agents-pass-all","Sessions that passed every hidden check, all three agents",1,"rate","100% (36/36)",36,[0.9036,1],"The ceiling: 36 of 36 means this task set cannot rank the agents on quality."],["coding-agents-sessions","Graded sessions (6 tasks × 2 repetitions × 3 agents)",36,"count","36",36,"\u0001","0 timeouts, 0 errors or usage-limit stops, nothing retried. Gemini CLI not run: it asked for a browser login."],["coding-agents-median-time-fastest","Median time per session, Sonnet 5.5 in Claude Code",23.1,"seconds","23.1 s",12,"\u0001","Fastest 18.7 s, slowest 44.5 s (a range of 12 sessions, not an interval)."],["coding-agents-time-ratio","Median time, GPT-6.1 Sol in Codex CLI vs Sonnet 5.5 in Claude Code (ratio of medians)",4.92,"ratio","4.9×",12,"\u0001","113.4 s vs 23.1 s. The run ranges do not overlap. Codex time includes the work its standing instructions asked for (see the caveats)."],["coding-agents-tester-notes","Codex sessions that read the tester’s global notes (standing instructions)",12,"count","12 of 12",12,"\u0001","10 of 12 also wrote a WORKLOG.md and 9 reported a commit attempt; no task asked for either. No Claude Code session did any of this."],["coding-agents-outside-edits","Edits that landed outside the task repository",0,"count","0",36,"\u0001","3 attempted writes to a temp folder (1 Sonnet 5.5, 2 Opus 5.5) were refused by the Claude Code permission check."],["coding-agents-cost-total","List-price estimate of all 36 sessions (calculation, not an invoice)",4.87,"usd","$4.87",36,"\u0001","\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","whisker","yLabel","viz","series","note","sourceIds","xLabel"],"$r":[["coding-agents-pass-rate","Coding sessions that passed every hidden check","A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals","dot-range","rate","ci95","Passed","IntervalDotPlot",[{"name":"Passed every hidden check","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.7575,1,12],["Claude Opus 5.5 · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",1,0.7575,1,12]]}}],"6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.",["agent-coding-agents"],"\u0001"],["coding-agents-wall-time","Time per coding session","Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)","dot-range","seconds","minmax","Seconds","LatencyLanes",[{"name":"Wall time per session","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",23.1,18.7,44.5,12,true],["Claude Opus 5.5 · Claude Code",56.9,29.8,185.8,12,"\u0001"],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",113.4,78.5,221.9,12,"\u0001"]]}}],"CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.",["agent-coding-agents"],"\u0001"],["coding-agents-time-by-task","Time per coding task","Median of 2 sessions per task and agent, seconds","grouped-bar","seconds","\u0001","Seconds","SmallMultiplesBars",{"$k":["name","points"],"$r":[["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["Pagination fix",19.1,2],["CLI --top flag",43.6,2],["Invoice refactor",24.1,2],["LRU cache",23.6,2],["Queue race",22.7,2],["Strict TypeScript types",27.6,2]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["Pagination fix",39.3,2],["CLI --top flag",55.3,2],["Invoice refactor",58.4,2],["LRU cache",62.4,2],["Queue race",118.6,2],["Strict TypeScript types",58,2]]}],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",{"$k":["label","value","n"],"$r":[["Pagination fix",82.7,2],["CLI --top flag",101.2,2],["Invoice refactor",129.1,2],["LRU cache",203.1,2],["Queue race",98.6,2],["Strict TypeScript types",143.6,2]]}]]},"2 sessions per cell, so one slow session moves a bar: Opus 5.5 in Claude Code took 51.5 s and 185.8 s on Queue race.",["agent-coding-agents"],"Task"],["coding-agents-tool-calls","Tool calls per coding session","Median; whiskers = fewest and most of 12 sessions (not an interval)","dot-range","calls","minmax","Tool calls","IntervalDotPlot",[{"name":"Tool calls per session","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",7.5,3,14,12],["Claude Opus 5.5 · Claude Code",7.5,5,14,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",12.5,8,18,12]]}}],"Claude Code counts its tool calls (Bash, Read, Edit, Write, Glob, Grep). Codex CLI counts shell commands and file changes; it has no separate read tool, so it reads files with shell commands. Turns are not compared: Codex reports one turn per run.",["agent-coding-agents"],"\u0001"],["coding-agents-token-mix","Tokens per session, as each CLI reports them","Mean per session by kind · not comparable across vendors","stacked-bar","tokens","\u0001","Tokens","TokenStack",{"$k":["name","points"],"$r":[["Cache read",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",73846,12],["Claude Opus 5.5 · Claude Code",96964,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",174560,12]]}],["Cache write",[{"label":"Claude Sonnet 5.5 · Claude Code","value":9163,"n":12},{"label":"Claude Opus 5.5 · Claude Code","value":11378,"n":12}]],["Uncached input",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",13,12],["Claude Opus 5.5 · Claude Code",16,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",19811,12]]}],["Output",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",3359,12],["Claude Opus 5.5 · Claude Code",5621,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",4072,12]]}]]},"Claude Code reports uncached input, cache reads, cache writes and output. Codex CLI reports input (its cached part included), the cached part and output (reasoning included), and no cache writes. Different tokenizers and context handling: compare Sonnet with Opus here, not Claude with Codex.",["agent-coding-agents"],"\u0001"],["coding-agents-diff-lines","Lines changed per session, by kind of file","Mean lines added plus deleted per session, against the base commit","stacked-bar","count","\u0001","Lines","TokenStack",{"$k":["name","points"],"$r":[["Code the task is about",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",48.4,12],["Claude Opus 5.5 · Claude Code",56.8,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",52.7,12]]}],["Tests",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",2.4,12],["Claude Opus 5.5 · Claude Code",19.4,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",70.4,12]]}],["README and package.json",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.8,12],["Claude Opus 5.5 · Claude Code",3.3,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",4.6,12]]}],["Work log and notes nobody asked for",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0,12],["Claude Opus 5.5 · Claude Code",0,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",9.4,12]]}]]},"Counted from each session’s patch, new files included. Median lines per session: Sonnet 5.5 in Claude Code 45, Opus 5.5 in Claude Code 77.5, GPT-6.1 Sol in Codex CLI 94. Codex wrote a WORKLOG.md in 10 of 12 sessions; no task asked for one.",["agent-coding-agents"],"\u0001"],["coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)","Reported tokens of all 12 sessions × list price, divided by the passes","bar","usd","\u0001","USD per pass","CostBars",[{"name":"List-price cost per pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.085,12],["Claude Opus 5.5 · Claude Code",0.2229,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",0.0978,12]]}}],"Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.",["agent-coding-agents","calc-repricing","price-anthropic","price-openai"],"\u0001"]]},"related":["hard-model-head-to-head","model-head-to-head","swe-bench-opus-vs-sonnet","effort-ladder"],"hero":{"statIds":["coding-agents-time-ratio"]}}}