• Claude Code
  • Codex CLI
  • Coding Agents
  • Hidden Tests
  • Head to head

Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks

On small real repository tasks graded by hidden tests, how do coding-agent CLIs compare when they run with their normal file and shell tools?

Published · 7 charts · Download the data or a carousel

4.9×

n = 12

Median time, GPT-6.1 Sol in Codex CLI vs Sonnet 5.5 in Claude Code (ratio of medians)

113.4 s vs 23.1 s. The run ranges do not overlap. Codex time includes the work its standing instructions asked for (see the caveats).

The answer

All 3 agents passed every hidden check in every session (12/12, 12/12, 12/12; 95% Wilson 76–100% each), so this task set cannot separate them on quality. Sonnet 5.5 in Claude Code was fastest (median 23.1 s; its sessions took 18.7 s to 44.5 s), Opus 5.5 in Claude Code took a median 56.9 s, GPT-6.1 Sol in Codex CLI took a median 113.4 s; the run ranges of Sonnet 5.5 in Claude Code and GPT-6.1 Sol in Codex CLI do not overlap. Codex made more tool calls (median 12.5 vs 7.5 and 7.5) and larger diffs (median 94 lines vs 45 and 77.5, mostly added tests), and every Codex session also followed the tester’s global AGENTS.md (10 of 12 wrote a work log nobody asked for), so its time and diff include extra work. Opus used 1.9× the output tokens of Sonnet in the same CLI (medians). Gemini CLI was not run: it needed a browser login.

Key numbers

100% (36/36)

Sessions that passed every hidden check, all three agents

95% CI 90%–100% · n = 36

36

Graded sessions (6 tasks × 2 repetitions × 3 agents)

n = 36

23.1s

Median time per session, Sonnet 5.5 in Claude Code

n = 12

12 of 12

Codex sessions that read the tester’s global notes (standing instructions)

n = 12

0

Edits that landed outside the task repository

n = 36

$4.87

List-price estimate of all 36 sessions (calculation, not an invoice)

n = 36

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Every rate is 95% or more
Claude Sonnet 5.5
Claude Opus 5.5
GPT-6.1 Sol (medium, tester’s…

Every interval overlaps every other: this chart does not order these rows.

3 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 12 per row3 of 3 at 100%: this task set cannot separate them.

A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals

6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Share card (PNG)
Entrance: medians race at 81× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

3 rows. Slowest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 113 s (range 78.5 s–222 s, n 12). Fastest Claude Sonnet 5.5 · Claude Code 23.1 s (range 18.7 s–44.5 s, n 12). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 12 per row

Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)

CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Share card (PNG)

Time per coding task

Median of 2 sessions per task and agent, seconds

Claude Sonnet 5.5 · Claude Code

Pagination fix
CLI --top flag
Invoice refactor
LRU cache
Queue race
Strict TypeScript types

Claude Opus 5.5 · Claude Code

Pagination fix
CLI --top flag
Invoice refactor
LRU cache
Queue race
Strict TypeScript types

GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Pagination fix
CLI --top flag
Invoice refactor
LRU cache
Queue race
Strict TypeScript types

One panel per series, all on the same axis.

6 tasks, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest CLI --top flag 43.6 s (n 2). Fastest Pagination fix 19.1 s (n 2). Claude Opus 5.5 · Claude Code: slowest Queue race 119 s (n 2). Fastest Pagination fix 39.3 s (n 2).

Notesn = 2 per row

2 sessions per cell, so one slow session moves a bar: Opus 5.5 in Claude Code took 51.5 s and 185.8 s on Queue race.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Share card (PNG)
Claude Sonnet 5.5
Claude Opus 5.5
GPT-6.1 Sol (medium, tester’s…

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 12.5 (range 8–18, n 12). Lowest Claude Opus 5.5 · Claude Code 7.5 (range 5–14, n 12). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n = 12 per row

Median; whiskers = fewest and most of 12 sessions (not an interval)

Claude Code counts its tool calls (Bash, Read, Edit, Write, Glob, Grep). Codex CLI counts shell commands and file changes; it has no separate read tool, so it reads files with shell commands. Turns are not compared: Codex reports one turn per run.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Share card (PNG)
  • Cache read
  • Cache write
  • Uncached input
  • Output
Bar length is the total; segments are its parts.
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

3 rows, 4 series: Cache read, Cache write, Uncached input, Output. Cache read: highest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 174,560 (n 12). Lowest Claude Sonnet 5.5 · Claude Code 73,846 (n 12). Cache write: highest Claude Opus 5.5 · Claude Code 11,378 (n 12). Lowest Claude Sonnet 5.5 · Claude Code 9,163 (n 12).

Notesn = 12 per row

Mean per session by kind · not comparable across vendors

Claude Code reports uncached input, cache reads, cache writes and output. Codex CLI reports input (its cached part included), the cached part and output (reasoning included), and no cache writes. Different tokenizers and context handling: compare Sonnet with Opus here, not Claude with Codex.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Share card (PNG)
  • Code the task is about
  • Tests
  • README and package.json
  • Work log and notes nobody asked for
Bar length is the total; segments are its parts.
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

3 rows, 4 series: Code the task is about, Tests, README and package.json, Work log and notes nobody asked for. Code the task is about: highest Claude Opus 5.5 · Claude Code 56.8 (n 12). Lowest Claude Sonnet 5.5 · Claude Code 48.4 (n 12). Tests: highest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 70.4 (n 12). Lowest Claude Sonnet 5.5 · Claude Code 2.4 (n 12).

Notesn = 12 per row

Mean lines added plus deleted per session, against the base commit

Counted from each session’s patch, new files included. Median lines per session: Sonnet 5.5 in Claude Code 45, Opus 5.5 in Claude Code 77.5, GPT-6.1 Sol in Codex CLI 94. Codex wrote a WORKLOG.md in 10 of 12 sessions; no task asked for one.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Share card (PNG)
Calculation
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Opus 5.5 · Claude Code $0.22 (n 12). Lowest Claude Sonnet 5.5 · Claude Code $0.085 (n 12).

Notesn = 12 per row

Reported tokens of all 12 sessions × list price, divided by the passes

Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.

Sources: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)

Tables

Every task and agent: sessions that passed (of 2)

Task by column: each cell is k of n from the table

All 18 cells are n of n: this task set cannot separate them.

TaskSonnet 5.5 in Claude Code1Opus 5.5 in Claude Code2GPT-6.1 Sol in Codex CLI3
Pagination fix2/22/22/2
CLI --top flag2/22/22/2
Invoice refactor2/22/22/2
LRU cache2/22/22/2
Queue race2/22/22/2
Strict TypeScript types2/22/22/2
  1. Sonnet 5.5 in Claude Code
  2. Opus 5.5 in Claude Code
  3. GPT-6.1 Sol in Codex CLI

Every cell is at its maximum, so all cells share one calm shade. Totals and other columns are in the Table view.

Counts from the table, not an intervaln = 2 per cell

6 rows by 3 columns. 18 of 18 k/n cells are full (n of n) and 0 are zero.

The six tasks and their controls

TaskWhat the agent had to doHidden checksBase repositoryReference patch
Pagination fixFix a pagination bug that spans two files (offset, page count, next and previous flags, validation) to the README rules13 hidden testsfails (tests 1/13)passes (tests 13/13)
CLI --top flagAdd a --top <n> flag with strict validation and exact error text to a small CLI11 hidden tests that run the CLIfails (tests 2/11)passes (tests 11/11)
Invoice refactorRemove duplication without a change in behaviour, floating-point and validation-order quirks included20 golden behaviour tests and 4 structure testsfails (tests 21/24)passes (tests 24/24)
LRU cacheImplement an LRU cache to a written spec (recency, eviction callback, TTL with an injected clock, SameValueZero keys, O(1) get and set)16 hidden testsfails (tests 0/16)passes (tests 16/16)
Queue raceFix a concurrency race and an onIdle hang in an async task queue10 hidden testsfails (tests 4/10)passes (tests 10/10)
Strict TypeScript typesMake tsc pass on strict settings with no any and no suppressions; the usage file and the config stay unchangedstatic checks, tsc with a hidden usage file (11 @ts-expect-error probes) and 4 runtime testsfails (tests 4/4; static checks fail)passes (tests 4/4)

Method

  1. Six small Node.js repositories: a pagination bug across two files, a CLI flag with exact error text, a refactor that must keep every quirk, an LRU cache to a written spec, a queue race and strict TypeScript types. Each has 1 to 4 visible tests; the hidden checks live outside the repository and run only after the session ends.
  2. Controls before the first session (Node v25.2.1): every base repository fails its hidden checks and every reference patch passes all of them.
  3. Claude Code 2.1.286 headless with the sonnet and opus models at the default effort; tools Bash, Read, Edit, Write, Glob, Grep; no web tools, MCP servers or sub-agents; the OS sandbox on, writes limited to the repository, no network; setting sources off. Codex CLI exec with GPT-6.1 Sol at medium effort, workspace-write sandbox, no network, user config ignored (its version was not recorded).
  4. 6 tasks × 2 repetitions × 3 agents = 36 sessions with a 10-minute timeout; nothing retried. The Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts.
  5. Pass = every hidden check passes; no partial credit. Recorded per session: wall time, tool calls, tokens as reported, files touched, lines changed against the base commit, commits and edits outside the repository.
  6. Gemini CLI 0.16.0 was probed first: it asked for a browser login, so it was not run and is not counted. The protocol was declared after the controls and before the first session.

Caveats

  • Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.
  • Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  • Each run pairs a CLI with a model (Claude Code with Claude models, Codex CLI with GPT-6.1 Sol), so the results cannot separate the CLI from the model.
  • n = 12 sessions per agent (2 per task). Time ranges are the fastest and slowest sessions, not confidence intervals; the two lanes shared one machine.
  • Tokens are as each CLI reports them, with different tokenizers and context handling: compare tokens inside Claude Code (Sonnet vs Opus), not across vendors. Costs are list-price calculations on subscription sessions, not invoices.

Sources

  • Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

    Our recorded runs ·

    Six small Node.js repositories with hidden tests; controls before the first session (every base fails, every reference passes). Claude Code 2.1.286 with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol at medium effort, 2 repetitions per task, OS sandboxes without network, protocol declared before the first session, every session kept. Gemini CLI was probed and not run (browser login).

    Raw data: coding-agents/sessions.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/coding-agents-head-to-head.

Models and comparisons in this study

More studies

All benchmarks
  • Prompt Caching
  • Cache Reuse

Does a new Claude Code session reuse the prompt cache of an earlier one?

30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.

0of 2 (95% interval 0% to 66%) · Later sessions with at least 50% of turn-1 input cached, A: new folder each time · n = 2

4 chartsUpdated October 7, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.