• Head to head
  • Hard Tasks
  • Codex
  • Codex CLI

Claude vs Codex on hard tasks: GPT-6.1 Sol joins the hard set

GPT-6.1 Sol in the Codex CLI passed 16/16 hard tasks at medium and high effort. Sonnet, Opus and Fable passed 24/24. What separates them: time and cost.

TL;DR

  • Our first hard-task run had no Codex result: all 30 Codex CLI attempts were blocked before any model call. A later batch reached the model: 32 Codex calls on the same 8 tasks and validators.
  • GPT-6.1 Sol through the Codex CLI passed 16 of 16 at medium effort and 16 of 16 at high effort (95% interval 81% to 100%). No format misses and no wrong answers.
  • Claude Sonnet 5.5, Opus 5.5, Opus 5.5 (high) and Fable 5.1 each passed 24 of 24 (86% to 100%). Pass rate does not separate these six configurations.
  • One pass-rate row does separate: GPT-6.1 Sol beats Claude Haiku 4.5 on strict pass rate, 16/16 against 11/24 (46%). The intervals do not overlap.
  • Median time per call: Sonnet 7.7 s, GPT-6.1 Sol (medium) 13.1 s, GPT-6.1 Sol (high) 18.1 s. The per-call ranges overlap, so speed is not a tested ranking.
  • At list price (a calculation), the lowest cost per strict pass is still Sonnet at $0.0143. GPT-6.1 Sol at high effort is close at $0.0151.

Every call and validator result: /benchmarks/hard-model-head-to-head.

Live story · 44 sHaiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.

Transcript
  1. Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
  2. 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  3. Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  4. Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  5. Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  6. Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  7. Open benchmarks: intervals, sources and every failure kept.

What changed since the first hard run

When tasks get hard reported 120 Claude Code calls and 30 blocked Codex attempts. Those 30 attempts are still in the record, as "blocked, not scored".

The new Codex batch used the same 8 tasks, the same prompts and the same validators. Two repetitions per task at two efforts gives 16 calls per configuration, against 24 for each Claude configuration. The batch stopped after 2 calls when its process ended. Before we resumed it, we wrote the rule into the protocol: the resume skips every task, repetition and effort already in the batch file. So no call was repeated or replaced. No run hit a usage limit.

The study now holds 152 calls that reached a model. 139 passed strictly (91%, 95% interval 86% to 95%).

Pass rate: six perfect scores

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

  • Claude Sonnet 5.5 · Claude Code: 24/24.
  • Claude Opus 5.5 · Claude Code: 24/24. At high effort: 24/24.
  • Claude Fable 5.1 · Claude Code: 24/24.
  • GPT-6.1 Sol (medium) · Codex CLI: 16/16. At high effort: 16/16.
  • Claude Haiku 4.5 · Claude Code: 11/24 strict, 16/24 on a lenient reading (5 format misses, 8 wrong answers).

A perfect 16/16 has a wider interval (81% to 100%) than a perfect 24/24 (86% to 100%). Both overlap, so the hard set still has a ceiling for these six. On the comparison pages, every Claude-vs-Sol pass-rate row is a tie: Sonnet vs Sol, Opus vs Sol and Fable vs Sol.

The exception is Haiku vs Sol. On strict pass rate, Haiku's interval (28% to 65%) does not overlap Sol's (81% to 100%), so Sol wins that row. On the lenient reading, Haiku reaches 47% to 82%, which overlaps, so that row is a tie. Part of Haiku's gap is format; part is wrong answers.

Speed: Claude Code was faster here, but not by a tested margin

Entrance: medians race at 28× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

7 rows. Slowest Claude Haiku 4.5 · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 16–24 per row

Median per configuration; whiskers = fastest and slowest call

One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Median total time per call:

  • Sonnet 7.7 s, Opus 9.2 s, Opus (high) 11.0 s
  • GPT-6.1 Sol (medium) 13.1 s
  • Fable 16.1 s
  • GPT-6.1 Sol (high) 18.1 s
  • Haiku 39.0 s

The whiskers are the fastest and slowest single call, not a confidence interval. Sonnet ran from 2.26 s to 34.79 s; GPT-6.1 Sol (medium) from 8.54 s to 61.6 s. The ranges overlap, so the comparison pages mark these rows "unclear", not a win.

Entrance: medians race at 25× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

7 rows. Slowest Claude Haiku 4.5 · Claude Code 35.5 s (range 12.9 s–70.3 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 6 s (range 0.9 s–30.6 s, n 24). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 16–24 per row

Median per configuration; whiskers = fastest and slowest call

One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Time to first useful output shows the same order: Sonnet 5.95 s, GPT-6.1 Sol (medium) 10.23 s, GPT-6.1 Sol (high) 12.69 s.

Part of the gap is the CLI, not the model. On a one-word answer, the Codex CLI took 6.00 s and sent 17,051 input tokens; Claude Code with Haiku took 2.53 s and sent 6,761 (5 runs each, different models). There, the ranges do not overlap: Claude Code vs Codex CLI.

  • First output event
  • First model output
  • Total wall time
Entrance: medians race at 4.3× real timeMotion reduced: press Replay to animateThe slowest median is 6 s. The clock runs at the recorded speed.
Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Median of 5 runs; whiskers = fastest and slowest run

Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

Tokens: GPT-6.1 Sol writes less

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).

Notesn 16–24 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Median output tokens per call: GPT-6.1 Sol 335 at medium and 436 at high, against 945 to 1,366 for the Claude configurations that passed everything, and 5,064 for Haiku. Reasoning tokens are part of those figures where the CLI reports them: Sol 150 and 225, Sonnet 585.

Fewer tokens is not better by itself. Here it means Sol reached the same strict pass rate with shorter replies.

Cost per strict pass (a calculation)

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Reported tokens × list price for every call, divided by strict passes. The calls ran on subscriptions, so this is not a bill.

  • Claude Sonnet 5.5: $0.0143
  • GPT-6.1 Sol (high): $0.0151
  • GPT-6.1 Sol (medium): $0.0256
  • Claude Opus 5.5: $0.0282
  • Claude Opus 5.5 (high): $0.0334
  • Claude Haiku 4.5: $0.0672
  • Claude Fable 5.1: $0.0933

Sol at high effort cost less per pass than Sol at medium. The output-token medians do not explain that (436 vs 335), and we report the figure without a cause. These costs have no interval, so read the gaps as a direction.

Calculation
  • Claude Code
  • Codex CLI
Better: upper left

Haloed: on the frontier (1 of 7). A point in the shaded area is no better on either axis than a haloed point.

List-price calculation, not a run. 7 points: Strict pass rate against USD per strict pass (list-price calculation). USD per strict pass (list-price calculation) runs from $0.014 to $0.093; Strict pass rate from 46% to 100%. Highlighted: Claude Sonnet 5.5 · Claude Code.

Notesn 16–24 per point

Strict pass rate against list-price cost per strict pass

Upper-left is better. Highlighted points are on the frontier: no other configuration passes at least as often for at most the same cost per pass. Frontier: Claude Sonnet 5.5 · Claude Code. Costs are calculations from tokens. Pass rates with their 95% intervals are in the pass-rate chart.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

The quality-vs-cost frontier holds one point: Sonnet. No other configuration passes as often for less.

And on the easy set?

The five-task study also gained Codex calls: a declared top-up of rep 3 at medium and high. GPT-6.1 Sol now has n = 15 at medium and high, and n = 10 at low. It passed every call. The whole five-task study now holds 127 of 130 passes. The only non-passes are 3 Sonnet format misses on one task. Details: /benchmarks/model-head-to-head.

So, Claude or Codex?

On this evidence:

  • On quality, it is a tie. Sonnet, Opus, Fable and GPT-6.1 Sol all passed every hard task we have. We need harder tasks to find a difference.
  • On time, Claude Code with Sonnet was faster in this run, and some of the gap is CLI start-up. The ranges overlap, so this is not a tested result.
  • On cost per pass, Sonnet and Sol at high effort are close, and both are below Opus and Fable.
  • Avoid Haiku 4.5 at its CLI default for hard reasoning. It is the only configuration that lost a pass-rate row.

Full records: GPT-6.1 Sol (Codex CLI), Codex CLI, Claude Code and Sonnet 5.5.

How we measured

  • Protocols declared before the first call. Two amendments were declared before they took effect: the Codex top-up on the easy set, and the resume rule for an interrupted batch.
  • Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call, one call at a time per account.
  • Attempts: Claude Code 120 attempts, 120 reached a model. Codex CLI 62 attempts, 32 reached a model, 30 blocked (earlier batch). Nothing was trimmed or retried.
  • Strict pass: the whole reply, trimmed, passes the validator. Format miss: the strict check fails, but a lenient extractor finds an answer that passes. Never counted as a pass.

Caveats

  • Route + model. Each row pairs a CLI with a model. A Claude-vs-GPT row compares route + model pairs, not models alone.
  • Different days. The Claude and Codex batches ran on different days on the same host, each on its own subscription.
  • Smaller Codex cells: n = 16 per configuration, 2 calls per task.
  • Ceiling: six configurations at 100%.
  • Costs are calculations, not bills.

See your own pass rates

Agent validates each step it delivers and records which model did the work, how long it took and what it cost. Try Agent and see the receipts for your own tasks.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.