Claude vs Codex on hard tasks: GPT-6.1 Sol joins the hard set
GPT-6.1 Sol in the Codex CLI passed 16/16 hard tasks at medium and high effort. Sonnet, Opus and Fable passed 24/24. What separates them: time and cost.
TL;DR
- Our first hard-task run had no Codex result: all 30 Codex CLI attempts were blocked before any model call. A later batch reached the model: 32 Codex calls on the same 8 tasks and validators.
- GPT-6.1 Sol through the Codex CLI passed 16 of 16 at medium effort and 16 of 16 at high effort (95% interval 81% to 100%). No format misses and no wrong answers.
- Claude Sonnet 5.5, Opus 5.5, Opus 5.5 (high) and Fable 5.1 each passed 24 of 24 (86% to 100%). Pass rate does not separate these six configurations.
- One pass-rate row does separate: GPT-6.1 Sol beats Claude Haiku 4.5 on strict pass rate, 16/16 against 11/24 (46%). The intervals do not overlap.
- Median time per call: Sonnet 7.7 s, GPT-6.1 Sol (medium) 13.1 s, GPT-6.1 Sol (high) 18.1 s. The per-call ranges overlap, so speed is not a tested ranking.
- At list price (a calculation), the lowest cost per strict pass is still Sonnet at $0.0143. GPT-6.1 Sol at high effort is close at $0.0151.
Every call and validator result: /benchmarks/hard-model-head-to-head.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks
139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.
Transcript
- Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
- 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
- Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
- Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
- Open benchmarks: intervals, sources and every failure kept.
What changed since the first hard run
When tasks get hard reported 120 Claude Code calls and 30 blocked Codex attempts. Those 30 attempts are still in the record, as "blocked, not scored".
The new Codex batch used the same 8 tasks, the same prompts and the same validators. Two repetitions per task at two efforts gives 16 calls per configuration, against 24 for each Claude configuration. The batch stopped after 2 calls when its process ended. Before we resumed it, we wrote the rule into the protocol: the resume skips every task, repetition and effort already in the batch file. So no call was repeated or replaced. No run hit a usage limit.
The study now holds 152 calls that reached a model. 139 passed strictly (91%, 95% interval 86% to 95%).
Pass rate: six perfect scores
- Claude Sonnet 5.5 · Claude Code: 24/24.
- Claude Opus 5.5 · Claude Code: 24/24. At high effort: 24/24.
- Claude Fable 5.1 · Claude Code: 24/24.
- GPT-6.1 Sol (medium) · Codex CLI: 16/16. At high effort: 16/16.
- Claude Haiku 4.5 · Claude Code: 11/24 strict, 16/24 on a lenient reading (5 format misses, 8 wrong answers).
A perfect 16/16 has a wider interval (81% to 100%) than a perfect 24/24 (86% to 100%). Both overlap, so the hard set still has a ceiling for these six. On the comparison pages, every Claude-vs-Sol pass-rate row is a tie: Sonnet vs Sol, Opus vs Sol and Fable vs Sol.
The exception is Haiku vs Sol. On strict pass rate, Haiku's interval (28% to 65%) does not overlap Sol's (81% to 100%), so Sol wins that row. On the lenient reading, Haiku reaches 47% to 82%, which overlaps, so that row is a tie. Part of Haiku's gap is format; part is wrong answers.
Speed: Claude Code was faster here, but not by a tested margin
Median total time per call:
- Sonnet 7.7 s, Opus 9.2 s, Opus (high) 11.0 s
- GPT-6.1 Sol (medium) 13.1 s
- Fable 16.1 s
- GPT-6.1 Sol (high) 18.1 s
- Haiku 39.0 s
The whiskers are the fastest and slowest single call, not a confidence interval. Sonnet ran from 2.26 s to 34.79 s; GPT-6.1 Sol (medium) from 8.54 s to 61.6 s. The ranges overlap, so the comparison pages mark these rows "unclear", not a win.
Time to first useful output shows the same order: Sonnet 5.95 s, GPT-6.1 Sol (medium) 10.23 s, GPT-6.1 Sol (high) 12.69 s.
Part of the gap is the CLI, not the model. On a one-word answer, the Codex CLI took 6.00 s and sent 17,051 input tokens; Claude Code with Haiku took 2.53 s and sent 6,761 (5 runs each, different models). There, the ranges do not overlap: Claude Code vs Codex CLI.
Tokens: GPT-6.1 Sol writes less
Median output tokens per call: GPT-6.1 Sol 335 at medium and 436 at high, against 945 to 1,366 for the Claude configurations that passed everything, and 5,064 for Haiku. Reasoning tokens are part of those figures where the CLI reports them: Sol 150 and 225, Sonnet 585.
Fewer tokens is not better by itself. Here it means Sol reached the same strict pass rate with shorter replies.
Cost per strict pass (a calculation)
Reported tokens × list price for every call, divided by strict passes. The calls ran on subscriptions, so this is not a bill.
- Claude Sonnet 5.5: $0.0143
- GPT-6.1 Sol (high): $0.0151
- GPT-6.1 Sol (medium): $0.0256
- Claude Opus 5.5: $0.0282
- Claude Opus 5.5 (high): $0.0334
- Claude Haiku 4.5: $0.0672
- Claude Fable 5.1: $0.0933
Sol at high effort cost less per pass than Sol at medium. The output-token medians do not explain that (436 vs 335), and we report the figure without a cause. These costs have no interval, so read the gaps as a direction.
The quality-vs-cost frontier holds one point: Sonnet. No other configuration passes as often for less.
And on the easy set?
The five-task study also gained Codex calls: a declared top-up of rep 3 at medium and high. GPT-6.1 Sol now has n = 15 at medium and high, and n = 10 at low. It passed every call. The whole five-task study now holds 127 of 130 passes. The only non-passes are 3 Sonnet format misses on one task. Details: /benchmarks/model-head-to-head.
So, Claude or Codex?
On this evidence:
- On quality, it is a tie. Sonnet, Opus, Fable and GPT-6.1 Sol all passed every hard task we have. We need harder tasks to find a difference.
- On time, Claude Code with Sonnet was faster in this run, and some of the gap is CLI start-up. The ranges overlap, so this is not a tested result.
- On cost per pass, Sonnet and Sol at high effort are close, and both are below Opus and Fable.
- Avoid Haiku 4.5 at its CLI default for hard reasoning. It is the only configuration that lost a pass-rate row.
Full records: GPT-6.1 Sol (Codex CLI), Codex CLI, Claude Code and Sonnet 5.5.
How we measured
- Protocols declared before the first call. Two amendments were declared before they took effect: the Codex top-up on the easy set, and the resume rule for an interrupted batch.
- Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call, one call at a time per account.
- Attempts: Claude Code 120 attempts, 120 reached a model. Codex CLI 62 attempts, 32 reached a model, 30 blocked (earlier batch). Nothing was trimmed or retried.
- Strict pass: the whole reply, trimmed, passes the validator. Format miss: the strict check fails, but a lenient extractor finds an answer that passes. Never counted as a pass.
Caveats
- Route + model. Each row pairs a CLI with a model. A Claude-vs-GPT row compares route + model pairs, not models alone.
- Different days. The Claude and Codex batches ran on different days on the same host, each on its own subscription.
- Smaller Codex cells: n = 16 per configuration, 2 calls per task.
- Ceiling: six configurations at 100%.
- Costs are calculations, not bills.
What to read next
- Does reasoning effort buy quality? Sonnet, Opus and GPT-6.1 Sol from low to high
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable
- Claude Code vs Codex CLI: the hidden context tax
- What does a router cost you?
See your own pass rates
Agent validates each step it delivers and records which model did the work, how long it took and what it cost. Try Agent and see the receipts for your own tasks.