• Head to head
  • Coding Agents
  • Claude Code
  • Codex CLI

Claude Code vs Codex CLI on hidden tests: when every agent passes, what differs?

36 hidden-test sessions: Sonnet 5.5 and Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI. All passed. Time, tool calls, diffs and one confound differ.

TL;DR

  • We ran 36 coding sessions: 6 small Node.js repository tasks × 2 repetitions × 3 agents. The agents: Claude Code with Sonnet 5.5, Claude Code with Opus 5.5, and Codex CLI with GPT-6.1 Sol at medium effort. Each agent had its normal file and shell tools. Hidden tests graded each session after it ended. Nothing was retried.
  • Every session passed. 12 of 12 per agent (95% Wilson interval 76% to 100% each). These tasks cannot rank the agents on quality.
  • Time separates them. Median time per session: Sonnet 23.1 s, Opus 56.9 s, Codex 113.4 s. Sonnet's slowest session (44.5 s) was shorter than the fastest Codex session (78.5 s). That is the only gap that passes our comparison rules.
  • We found a confound in our own run. Codex CLI read the tester's global AGENTS.md in 12 of 12 sessions, although we passed --ignore-user-config. 10 of 12 wrote a WORKLOG.md that no task asked for, and 9 reported a commit attempt. Codex time, tool calls and diffs include that work.
  • The agents wrote about the same amount of task code. The diffs differ in tests: a mean of 2.4 test lines per session for Sonnet, 19.4 for Opus and 70.4 for Codex.
  • Opus did more work than Sonnet in the same CLI, with the same result: 1.9x the output tokens by median (1.7x by mean, the value the token chart shows) and 2.6x the list-price cost per pass ($0.22 against $0.085, a calculation).

Every session, task and control: /benchmarks/coding-agents-head-to-head.

Entrance: medians race at 81× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

3 rows. Slowest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 113 s (range 78.5 s–222 s, n 12). Fastest Claude Sonnet 5.5 · Claude Code 23.1 s (range 18.7 s–44.5 s, n 12). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 12 per row

Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)

CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Why we ran a coding-agent test

Our earlier head-to-heads called each model with its tools off: one prompt, one reply, one validator. That measures a model through its CLI. A coding agent does more. It reads files, runs tests, edits code and decides when it is done.

So this time, each CLI got a real repository, a task and its normal tools. We graded the result with tests that the agent could not see. The question: when all three agents can do the job, what is different?

The setup

Six small Node.js repositories, one task each:

  • Pagination fix: a bug that spans two files (offset, page count, next and previous flags, validation).
  • CLI --top flag: add a flag with strict validation and exact error text.
  • Invoice refactor: remove duplication and keep every quirk, floating-point and validation order included.
  • LRU cache: implement a written spec (recency, eviction callback, TTL with an injected clock, O(1) get and set).
  • Queue race: fix a concurrency race and an onIdle hang in an async task queue.
  • Strict TypeScript types: make tsc pass on strict settings with no any and no suppressions.

Each repository had 1 to 4 visible tests. The hidden checks lived outside the repository and ran only after the session ended.

Controls before the first session. Every base repository failed its hidden checks. Every reference patch passed all of them. So a pass means the agent fixed the task, and a failure would mean it did not.

The agents:

  • Claude Code 2.1.286, headless, with the sonnet and opus models at the default effort. Tools: Bash, Read, Edit, Write, Glob and Grep. No web tools, MCP servers or sub-agents. The OS sandbox was on, writes were limited to the repository, and the network was off. Setting sources were off.
  • Codex CLI exec with GPT-6.1 Sol at medium effort, the workspace-write sandbox, no network and user config ignored. Its version was not recorded.

The two lanes ran at the same time on one machine, with different accounts. Each session had a 10-minute timeout. Gemini CLI 0.16.0 asked for a browser login, so it was not run and is not counted.

Quality: 36 of 36, a ceiling

Every rate is 95% or more
Claude Sonnet 5.5
Claude Opus 5.5
GPT-6.1 Sol (medium, tester’s…

Every interval overlaps every other: this chart does not order these rows.

3 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 12 per row3 of 3 at 100%: this task set cannot separate them.

A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals

6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

All three agents passed every hidden check in every session. A perfect 12 of 12 has a 95% interval of 76% to 100%, so the three results sit in the same band. On every comparison page, the pass-rate row from this study is a tie: Claude Code vs Codex CLI, Sonnet vs GPT-6.1 Sol, Opus vs GPT-6.1 Sol and Sonnet vs Opus.

A ceiling is not a ranking. It says the tasks were too easy to separate these agents. Our hard head-to-head and our effort ladder hit the same ceiling. To find a quality gap, we need harder tasks.

Time: the one gap that holds

The chart at the top shows the median and the fastest and slowest of 12 sessions per agent:

  • Claude Code with Sonnet 5.5: 23.1 s (18.7 s to 44.5 s)
  • Claude Code with Opus 5.5: 56.9 s (29.8 s to 185.8 s)
  • Codex CLI with GPT-6.1 Sol: 113.4 s (78.5 s to 221.9 s)

A range is not a confidence interval. Our rule for time is strict: a row names a winner only when the two run ranges do not overlap and each side has at least 5 runs. Sonnet's range and the Codex range do not overlap, with 12 runs each. So Claude Code vs Codex CLI and Sonnet vs GPT-6.1 Sol mark Claude Code with Sonnet as faster on this row. The median ratio is 4.9x. Opus overlaps both other agents, so its rows are unclear. That is also why the chart draws no finish order: each lane overlaps the Opus lane, though the Sonnet and Codex lanes do not overlap each other.

Time per coding task

Median of 2 sessions per task and agent, seconds

Claude Sonnet 5.5 · Claude Code

Pagination fix
CLI --top flag
Invoice refactor
LRU cache
Queue race
Strict TypeScript types

Claude Opus 5.5 · Claude Code

Pagination fix
CLI --top flag
Invoice refactor
LRU cache
Queue race
Strict TypeScript types

GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Pagination fix
CLI --top flag
Invoice refactor
LRU cache
Queue race
Strict TypeScript types

One panel per series, all on the same axis.

6 tasks, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest CLI --top flag 43.6 s (n 2). Fastest Pagination fix 19.1 s (n 2). Claude Opus 5.5 · Claude Code: slowest Queue race 119 s (n 2). Fastest Pagination fix 39.3 s (n 2).

Notesn = 2 per row

2 sessions per cell, so one slow session moves a bar: Opus 5.5 in Claude Code took 51.5 s and 185.8 s on Queue race.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Per task, each bar is the median of only 2 sessions:

  • Sonnet had the lowest time on all 6 tasks.
  • Codex had the highest time on 5 of 6. Its longest cell was the LRU cache, at 203.1 s.
  • Queue race is the exception. Opus took 51.5 s and 185.8 s there, so its median of 118.6 s is above the Codex median of 98.6 s. One slow session moves a 2-session bar.

Part of the Codex time is not the task. The next section explains why.

The confound we found in our own run

We ran Codex CLI with --ignore-user-config. We expected it to start without the tester's personal settings. It did not. That flag does not switch off the global AGENTS.md file, which holds the tester's standing instructions.

The session logs show the effect:

  • 12 of 12 Codex sessions read the tester's notes with a shell command.
  • 10 of 12 wrote a WORKLOG.md. No task asked for one.
  • 9 of 12 reported a commit attempt. No task asked for that either.
  • No Claude Code session did any of this. Claude Code ran with setting sources off.

What it changes: the Codex time, tool calls and lines changed include work that the tester's notes asked for. We cannot separate that work from the task work inside a session.

What it does not change: the pass result. The hidden tests ran on the code, and every Codex session passed.

We found the confound in the analysis, after the run. We did not drop the Codex lane, and we did not rerun it quietly. The study discloses the confound as its first caveat, and every Codex label on the site carries the qualifier "tester's AGENTS.md", so every Codex fact and comparison row shows it. A clean rerun with the notes switched off would remove the confound. We have not run it yet.

A lesson for your own setup: a coding CLI can load instructions that you did not pass on the command line. Check which instruction files your agent reads before you benchmark it or put it in CI. Codex followed the notes faithfully. In your own repository that is a feature. In a benchmark it is a confound.

How each agent works: tool calls and diffs

Claude Sonnet 5.5
Claude Opus 5.5
GPT-6.1 Sol (medium, tester’s…

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 12.5 (range 8–18, n 12). Lowest Claude Opus 5.5 · Claude Code 7.5 (range 5–14, n 12). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n = 12 per row

Median; whiskers = fewest and most of 12 sessions (not an interval)

Claude Code counts its tool calls (Bash, Read, Edit, Write, Glob, Grep). Codex CLI counts shell commands and file changes; it has no separate read tool, so it reads files with shell commands. Turns are not compared: Codex reports one turn per run.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Median tool calls per session: Sonnet 7.5, Opus 7.5, Codex 12.5. The ranges overlap (3 to 14, 5 to 14 and 8 to 18). The two CLIs also count differently. Claude Code counts each tool call. Codex counts shell commands and file changes, and it has no separate read tool, so it reads files with shell commands. More calls is not better or worse by itself, so no comparison page names a winner on this row.

  • Code the task is about
  • Tests
  • README and package.json
  • Work log and notes nobody asked for
Bar length is the total; segments are its parts.
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

3 rows, 4 series: Code the task is about, Tests, README and package.json, Work log and notes nobody asked for. Code the task is about: highest Claude Opus 5.5 · Claude Code 56.8 (n 12). Lowest Claude Sonnet 5.5 · Claude Code 48.4 (n 12). Tests: highest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 70.4 (n 12). Lowest Claude Sonnet 5.5 · Claude Code 2.4 (n 12).

Notesn = 12 per row

Mean lines added plus deleted per session, against the base commit

Counted from each session’s patch, new files included. Median lines per session: Sonnet 5.5 in Claude Code 45, Opus 5.5 in Claude Code 77.5, GPT-6.1 Sol in Codex CLI 94. Codex wrote a WORKLOG.md in 10 of 12 sessions; no task asked for one.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

The diffs tell a clearer story. Mean lines added plus deleted per session, against the base commit:

  • Code the task is about: Sonnet 48.4, Opus 56.8, Codex 52.7. About the same.
  • Tests: Sonnet 2.4, Opus 19.4, Codex 70.4.
  • README and package.json: Sonnet 0.8, Opus 3.3, Codex 4.6.
  • Work log and notes nobody asked for: Sonnet 0, Opus 0, Codex 9.4.

The median diff was 45 lines for Sonnet, 77.5 for Opus and 94 for Codex. So the size gap is mostly tests. Sonnet almost never added a test. Codex added the most. We cannot tell how much of the Codex test code the tester's notes asked for. The extra tests did not change any grade: the hidden tests decided every pass.

Tokens: compare inside one CLI

  • Cache read
  • Cache write
  • Uncached input
  • Output
Bar length is the total; segments are its parts.
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

3 rows, 4 series: Cache read, Cache write, Uncached input, Output. Cache read: highest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 174,560 (n 12). Lowest Claude Sonnet 5.5 · Claude Code 73,846 (n 12). Cache write: highest Claude Opus 5.5 · Claude Code 11,378 (n 12). Lowest Claude Sonnet 5.5 · Claude Code 9,163 (n 12).

Notesn = 12 per row

Mean per session by kind · not comparable across vendors

Claude Code reports uncached input, cache reads, cache writes and output. Codex CLI reports input (its cached part included), the cached part and output (reasoning included), and no cache writes. Different tokenizers and context handling: compare Sonnet with Opus here, not Claude with Codex.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Each CLI reports tokens in its own way, with its own tokenizer. Compare Sonnet with Opus here, not Claude with Codex.

  • Inside Claude Code, Opus used more of every kind. Mean per session: cache reads 96,964 against 73,846, cache writes 11,378 against 9,163, output 5,621 against 3,359. That is 1.7x by mean; by medians, Opus wrote 1.9x the output tokens of Sonnet.
  • Codex CLI reported a mean of 174,560 cached input tokens, 19,811 uncached input tokens and 4,072 output tokens (reasoning included). It reports no cache writes. Its input includes its own system prompt and, in this run, the tester's AGENTS.md.

Why does a one-line change take tens of thousands of input tokens? Each agent turn resends the conversation, and the CLI adds its own system prompt and tool definitions. See the hidden context tax and tokens per call, explained.

Cost per pass (a calculation)

Calculation
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Opus 5.5 · Claude Code $0.22 (n 12). Lowest Claude Sonnet 5.5 · Claude Code $0.085 (n 12).

Notesn = 12 per row

Reported tokens of all 12 sessions × list price, divided by the passes

Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.

Sources: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Reported tokens of all 12 sessions × list price, divided by the passes. Both CLIs ran on flat subscriptions, so this is not a bill.

  • Claude Code with Sonnet 5.5: $0.085 per pass
  • Codex CLI with GPT-6.1 Sol: $0.098 per pass
  • Claude Code with Opus 5.5: $0.22 per pass

All 36 sessions together come to $4.87 at list price. Claude cache writes are priced at 2x input, as in our other studies; Claude Code's own estimate gives the same totals. These values have no interval, so the comparison pages mark the cost rows "unclear": the gap is stated, not tested.

Opus cost 2.6x as much as Sonnet per pass (our calculation). Opus lists at twice Sonnet's input and output price, with the same cache-read price. So the extra cost is more than the price gap: Opus also used more tokens for the same result. Our Opus vs Sonnet SWE-bench probe shows the same 2.6x on real GitHub issues.

What this means if you choose a coding agent

  • On small, well-specified tasks with tests, all three finished the job. Choose on time, cost and how much extra each agent does.
  • Claude Code with Sonnet 5.5 was the fastest and the cheapest per pass here. It also wrote the smallest diffs, with almost no new tests. If you want tests with each change, ask for them in the task.
  • Opus 5.5 did not pass more. It took longer and cost 2.6x per pass. That matches every other Sonnet vs Opus result we have.
  • Codex CLI follows standing instructions. Know which ones it loads. Its time and diff here include work that the tester's notes asked for.
  • These are small tasks. Repository size, task length and harder bugs can change the order. Run your own tasks before you decide.

How we measured

  • Protocol declared after the controls and before the first session. 6 tasks × 2 repetitions × 3 agents = 36 sessions; 0 timeouts, 0 errors or usage-limit stops, nothing retried.
  • Pass = every hidden check passes. No partial credit.
  • Recorded per session: wall time from CLI start to exit, tool calls, tokens as each CLI reports them, files touched, lines changed against the base commit, commits, and edits outside the repository.
  • Edits outside the repository: 0. Claude Code's permission check refused 3 attempted writes to a temp folder (1 Sonnet, 2 Opus).
  • Costs: reported tokens × list price; a calculation, not an invoice.

Caveats

  • The Codex confound. Codex CLI ran with the tester's global AGENTS.md. Its time, tool calls and lines changed include that work.
  • A ceiling. Every agent passed every session. These tasks cannot separate the agents on quality.
  • CLI and model together. Each run pairs a CLI with a model (Claude Code with Claude models, Codex CLI with GPT-6.1 Sol). The results cannot separate the CLI from the model.
  • Small samples. n = 12 sessions per agent, 2 per task. Time ranges are the fastest and slowest sessions, not confidence intervals. The two lanes shared one machine.
  • Tokens differ by vendor. Different tokenizers and context handling: compare tokens inside Claude Code only.

Compare agents on your own repository

Agent records the route, the model, the time, the tokens, the cost and the validation result of every step of your work. Try Agent and see which agent fits which task in your own repository.

The data behind this post

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.