Claude Code vs Codex CLI on hidden tests: when every agent passes, what differs?
36 hidden-test sessions: Sonnet 5.5 and Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI. All passed. Time, tool calls, diffs and one confound differ.
TL;DR
- We ran 36 coding sessions: 6 small Node.js repository tasks × 2 repetitions × 3 agents. The agents: Claude Code with Sonnet 5.5, Claude Code with Opus 5.5, and Codex CLI with GPT-6.1 Sol at medium effort. Each agent had its normal file and shell tools. Hidden tests graded each session after it ended. Nothing was retried.
- Every session passed. 12 of 12 per agent (95% Wilson interval 76% to 100% each). These tasks cannot rank the agents on quality.
- Time separates them. Median time per session: Sonnet 23.1 s, Opus 56.9 s, Codex 113.4 s. Sonnet's slowest session (44.5 s) was shorter than the fastest Codex session (78.5 s). That is the only gap that passes our comparison rules.
- We found a confound in our own run. Codex CLI read the tester's global
AGENTS.mdin 12 of 12 sessions, although we passed--ignore-user-config. 10 of 12 wrote aWORKLOG.mdthat no task asked for, and 9 reported a commit attempt. Codex time, tool calls and diffs include that work. - The agents wrote about the same amount of task code. The diffs differ in tests: a mean of 2.4 test lines per session for Sonnet, 19.4 for Opus and 70.4 for Codex.
- Opus did more work than Sonnet in the same CLI, with the same result: 1.9x the output tokens by median (1.7x by mean, the value the token chart shows) and 2.6x the list-price cost per pass ($0.22 against $0.085, a calculation).
Every session, task and control: /benchmarks/coding-agents-head-to-head.
Why we ran a coding-agent test
Our earlier head-to-heads called each model with its tools off: one prompt, one reply, one validator. That measures a model through its CLI. A coding agent does more. It reads files, runs tests, edits code and decides when it is done.
So this time, each CLI got a real repository, a task and its normal tools. We graded the result with tests that the agent could not see. The question: when all three agents can do the job, what is different?
The setup
Six small Node.js repositories, one task each:
- Pagination fix: a bug that spans two files (offset, page count, next and previous flags, validation).
- CLI
--topflag: add a flag with strict validation and exact error text. - Invoice refactor: remove duplication and keep every quirk, floating-point and validation order included.
- LRU cache: implement a written spec (recency, eviction callback, TTL with an injected clock, O(1) get and set).
- Queue race: fix a concurrency race and an
onIdlehang in an async task queue. - Strict TypeScript types: make
tscpass on strict settings with noanyand no suppressions.
Each repository had 1 to 4 visible tests. The hidden checks lived outside the repository and ran only after the session ended.
Controls before the first session. Every base repository failed its hidden checks. Every reference patch passed all of them. So a pass means the agent fixed the task, and a failure would mean it did not.
The agents:
- Claude Code 2.1.286, headless, with the
sonnetandopusmodels at the default effort. Tools: Bash, Read, Edit, Write, Glob and Grep. No web tools, MCP servers or sub-agents. The OS sandbox was on, writes were limited to the repository, and the network was off. Setting sources were off. - Codex CLI
execwith GPT-6.1 Sol at medium effort, the workspace-write sandbox, no network and user config ignored. Its version was not recorded.
The two lanes ran at the same time on one machine, with different accounts. Each session had a 10-minute timeout. Gemini CLI 0.16.0 asked for a browser login, so it was not run and is not counted.
Quality: 36 of 36, a ceiling
All three agents passed every hidden check in every session. A perfect 12 of 12 has a 95% interval of 76% to 100%, so the three results sit in the same band. On every comparison page, the pass-rate row from this study is a tie: Claude Code vs Codex CLI, Sonnet vs GPT-6.1 Sol, Opus vs GPT-6.1 Sol and Sonnet vs Opus.
A ceiling is not a ranking. It says the tasks were too easy to separate these agents. Our hard head-to-head and our effort ladder hit the same ceiling. To find a quality gap, we need harder tasks.
Time: the one gap that holds
The chart at the top shows the median and the fastest and slowest of 12 sessions per agent:
- Claude Code with Sonnet 5.5: 23.1 s (18.7 s to 44.5 s)
- Claude Code with Opus 5.5: 56.9 s (29.8 s to 185.8 s)
- Codex CLI with GPT-6.1 Sol: 113.4 s (78.5 s to 221.9 s)
A range is not a confidence interval. Our rule for time is strict: a row names a winner only when the two run ranges do not overlap and each side has at least 5 runs. Sonnet's range and the Codex range do not overlap, with 12 runs each. So Claude Code vs Codex CLI and Sonnet vs GPT-6.1 Sol mark Claude Code with Sonnet as faster on this row. The median ratio is 4.9x. Opus overlaps both other agents, so its rows are unclear. That is also why the chart draws no finish order: each lane overlaps the Opus lane, though the Sonnet and Codex lanes do not overlap each other.
Per task, each bar is the median of only 2 sessions:
- Sonnet had the lowest time on all 6 tasks.
- Codex had the highest time on 5 of 6. Its longest cell was the LRU cache, at 203.1 s.
- Queue race is the exception. Opus took 51.5 s and 185.8 s there, so its median of 118.6 s is above the Codex median of 98.6 s. One slow session moves a 2-session bar.
Part of the Codex time is not the task. The next section explains why.
The confound we found in our own run
We ran Codex CLI with --ignore-user-config. We expected it to start without the tester's personal settings. It did not. That flag does not switch off the global AGENTS.md file, which holds the tester's standing instructions.
The session logs show the effect:
- 12 of 12 Codex sessions read the tester's notes with a shell command.
- 10 of 12 wrote a
WORKLOG.md. No task asked for one. - 9 of 12 reported a commit attempt. No task asked for that either.
- No Claude Code session did any of this. Claude Code ran with setting sources off.
What it changes: the Codex time, tool calls and lines changed include work that the tester's notes asked for. We cannot separate that work from the task work inside a session.
What it does not change: the pass result. The hidden tests ran on the code, and every Codex session passed.
We found the confound in the analysis, after the run. We did not drop the Codex lane, and we did not rerun it quietly. The study discloses the confound as its first caveat, and every Codex label on the site carries the qualifier "tester's AGENTS.md", so every Codex fact and comparison row shows it. A clean rerun with the notes switched off would remove the confound. We have not run it yet.
A lesson for your own setup: a coding CLI can load instructions that you did not pass on the command line. Check which instruction files your agent reads before you benchmark it or put it in CI. Codex followed the notes faithfully. In your own repository that is a feature. In a benchmark it is a confound.
How each agent works: tool calls and diffs
Median tool calls per session: Sonnet 7.5, Opus 7.5, Codex 12.5. The ranges overlap (3 to 14, 5 to 14 and 8 to 18). The two CLIs also count differently. Claude Code counts each tool call. Codex counts shell commands and file changes, and it has no separate read tool, so it reads files with shell commands. More calls is not better or worse by itself, so no comparison page names a winner on this row.
The diffs tell a clearer story. Mean lines added plus deleted per session, against the base commit:
- Code the task is about: Sonnet 48.4, Opus 56.8, Codex 52.7. About the same.
- Tests: Sonnet 2.4, Opus 19.4, Codex 70.4.
- README and package.json: Sonnet 0.8, Opus 3.3, Codex 4.6.
- Work log and notes nobody asked for: Sonnet 0, Opus 0, Codex 9.4.
The median diff was 45 lines for Sonnet, 77.5 for Opus and 94 for Codex. So the size gap is mostly tests. Sonnet almost never added a test. Codex added the most. We cannot tell how much of the Codex test code the tester's notes asked for. The extra tests did not change any grade: the hidden tests decided every pass.
Tokens: compare inside one CLI
Each CLI reports tokens in its own way, with its own tokenizer. Compare Sonnet with Opus here, not Claude with Codex.
- Inside Claude Code, Opus used more of every kind. Mean per session: cache reads 96,964 against 73,846, cache writes 11,378 against 9,163, output 5,621 against 3,359. That is 1.7x by mean; by medians, Opus wrote 1.9x the output tokens of Sonnet.
- Codex CLI reported a mean of 174,560 cached input tokens, 19,811 uncached input tokens and 4,072 output tokens (reasoning included). It reports no cache writes. Its input includes its own system prompt and, in this run, the tester's
AGENTS.md.
Why does a one-line change take tens of thousands of input tokens? Each agent turn resends the conversation, and the CLI adds its own system prompt and tool definitions. See the hidden context tax and tokens per call, explained.
Cost per pass (a calculation)
Reported tokens of all 12 sessions × list price, divided by the passes. Both CLIs ran on flat subscriptions, so this is not a bill.
- Claude Code with Sonnet 5.5: $0.085 per pass
- Codex CLI with GPT-6.1 Sol: $0.098 per pass
- Claude Code with Opus 5.5: $0.22 per pass
All 36 sessions together come to $4.87 at list price. Claude cache writes are priced at 2x input, as in our other studies; Claude Code's own estimate gives the same totals. These values have no interval, so the comparison pages mark the cost rows "unclear": the gap is stated, not tested.
Opus cost 2.6x as much as Sonnet per pass (our calculation). Opus lists at twice Sonnet's input and output price, with the same cache-read price. So the extra cost is more than the price gap: Opus also used more tokens for the same result. Our Opus vs Sonnet SWE-bench probe shows the same 2.6x on real GitHub issues.
What this means if you choose a coding agent
- On small, well-specified tasks with tests, all three finished the job. Choose on time, cost and how much extra each agent does.
- Claude Code with Sonnet 5.5 was the fastest and the cheapest per pass here. It also wrote the smallest diffs, with almost no new tests. If you want tests with each change, ask for them in the task.
- Opus 5.5 did not pass more. It took longer and cost 2.6x per pass. That matches every other Sonnet vs Opus result we have.
- Codex CLI follows standing instructions. Know which ones it loads. Its time and diff here include work that the tester's notes asked for.
- These are small tasks. Repository size, task length and harder bugs can change the order. Run your own tasks before you decide.
How we measured
- Protocol declared after the controls and before the first session. 6 tasks × 2 repetitions × 3 agents = 36 sessions; 0 timeouts, 0 errors or usage-limit stops, nothing retried.
- Pass = every hidden check passes. No partial credit.
- Recorded per session: wall time from CLI start to exit, tool calls, tokens as each CLI reports them, files touched, lines changed against the base commit, commits, and edits outside the repository.
- Edits outside the repository: 0. Claude Code's permission check refused 3 attempted writes to a temp folder (1 Sonnet, 2 Opus).
- Costs: reported tokens × list price; a calculation, not an invoice.
Caveats
- The Codex confound. Codex CLI ran with the tester's global
AGENTS.md. Its time, tool calls and lines changed include that work. - A ceiling. Every agent passed every session. These tasks cannot separate the agents on quality.
- CLI and model together. Each run pairs a CLI with a model (Claude Code with Claude models, Codex CLI with GPT-6.1 Sol). The results cannot separate the CLI from the model.
- Small samples. n = 12 sessions per agent, 2 per task. Time ranges are the fastest and slowest sessions, not confidence intervals. The two lanes shared one machine.
- Tokens differ by vendor. Different tokenizers and context handling: compare tokens inside Claude Code only.
What to read next
- Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look
- Claude vs Codex on hard tasks: GPT-6.1 Sol joins the hard set
- Claude Code vs Codex CLI: the hidden context tax
- Claude Code vs Codex CLI vs the API: latency
- Claude Sonnet vs Opus: when is Opus worth the price?
- Model pages: Claude Code, Codex CLI, GPT-6.1 Sol (Codex CLI)
Compare agents on your own repository
Agent records the route, the model, the time, the tokens, the cost and the validation result of every step of your work. Try Agent and see which agent fits which task in your own repository.