GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
On a task set built so that Claude Sonnet 5.5 did not pass it every time, does pass rate separate GPT-6.1 Sol (Codex CLI) from Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 (Claude Code)?
Published · 8 charts · Download the data or a carousel
39%
The answer
22 of 56 counted calls passed strictly (39%, 95% Wilson interval 28% to 52%) on 4 tasks. The Sonnet pilot did not pass these twice. GPT-6.1 Sol (medium) passed 11/16 strictly (69%, 95% interval 44% to 86%). Opus 5.5 passed 5/12 strictly (42%, 95% interval 19% to 68%). Sonnet 5.5 passed 6/16 strictly (38%, 95% interval 18% to 61%). Haiku 4.5 passed 0/12 strictly (0%, 95% interval 0% to 24%). GPT-6.1 Sol (medium) is ahead of Haiku 4.5 (the 95% intervals do not overlap). The other 5 of 6 pairs overlap, so this set cannot rank them. The lenient reading counts format misses. Opus 5.5 is also ahead of Haiku 4.5 on the lenient reading. Opus 5.5 6/12 (n = 12, 95% Wilson interval 25.4% to 74.6%); Haiku 4.5 0/12 (n = 12, 95% Wilson interval 0.0% to 24.2%). The intervals miss by 1.1 points (calculation), so this is fragile. Haiku 4.5 passed none (a floor for that configuration on this set). No configuration passed every call across the full set. Some per-task cells still hit a ceiling; see the task chart. Of 34 non-passes, 1 was a format miss with the right answer. Another 21 were wrong answers. 12 gave no answer; 10 hit the 300 s timeout. 11 of 56 calls tried a tool although tools were off. Per configuration: GPT-6.1 Sol (medium) 0 of 16, Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12. None of them passed. The Codex runner also tells the model not to call tools; the Claude Code runner does not. Median total time per completed call (wrong answers and format misses included; timeouts and tool-call parse errors excluded; ranges are not intervals): GPT-6.1 Sol (medium) 120.2 s (n = 13, range 46.2 s to 273.5 s). Opus 5.5 80.3 s (n = 9, range 3.8 s to 279.5 s). Sonnet 5.5 70.4 s (n = 12, range 4.3 s to 210.1 s). Haiku 4.5 109.0 s (n = 10, range 25.7 s to 223.9 s). Every configuration’s fastest-to-slowest range overlaps every other, so the medians describe this run and are not a tested ranking. Cost is a list-price calculation; the calls ran on subscriptions. The lowest recorded lower bound was GPT-6.1 Sol (medium) · Codex CLI: $0.083 (a lower bound: 3 timed-out calls report no tokens; $0.100 if each had cost a median call, an assumption). Selection effect: the study picked tasks with mixed or failed Sonnet pilot results. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls. The counted calls are new calls. Selection can produce this pattern; the run does not establish its cause.
Key numbers
41% (23/56)
Counted calls with a correct answer, format misses included (lenient reading)
95% CI 29%–54% · n = 56
69% (11/16)
GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks
95% CI 44%–86% · n = 16
42% (5/12)
Opus 5.5 (Claude Code): strict pass rate on the harder tasks
95% CI 19%–68% · n = 12
38% (6/16)
Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks
95% CI 18%–61% · n = 16
0% (0/12)
Haiku 4.5 (Claude Code): strict pass rate on the harder tasks
95% CI 0%–24% · n = 12
1
Non-passes that were format misses, not wrong answers
of 34 non-passes (21 wrong answers, 12 no answer) · n = 34
20% (11/56)
Counted calls that tried a tool although tools were off (a behaviour, not a quality score)
95% CI 11%–32% · n = 56
18% (10/56)
Counted calls that ran past the 300 s limit and gave no answer
95% CI 10%–30% · n = 56
81% (26/32)
Pilot (not scored): strict passes of Claude Sonnet 5.5 on all candidate tasks
95% CI 65%–91% · n = 32
25% (2/8)
Pilot (not scored): strict passes of Claude Sonnet 5.5 on the tasks that were kept
95% CI 7.1%–59% · n = 8
4
Candidate tasks kept for the counted set
of 16 (Sonnet passed 12 candidates 2 of 2, including 10 of 10 code, SQL, spec, numeric and simulation tasks) · n = 16
$0.0829
GPT-6.1 Sol (medium) · Codex CLI
Lowest recorded cost lower bound per strict pass (calculation)
(lower bound) · n = 16
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 69% | 69% | Strict pass: 44%–86%; Lenient (format misses counted): 44%–86% | 16 |
| Claude Opus 5.5 · Claude Code | 42% | 50% | Strict pass: 19%–68%; Lenient (format misses counted): 25%–75% | 12 |
| Claude Sonnet 5.5 · Claude Code | 38% | 38% | Strict pass: 18%–61%; Lenient (format misses counted): 18%–61% | 16 |
| Claude Haiku 4.5 · Claude Code | 0% | 0% | Strict pass: 0%–24%; Lenient (format misses counted): 0%–24% | 12 |
4 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest GPT-6.1 Sol (medium) · Codex CLI 69% (95% interval 44%–86%, n 16). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–24%, n 12). Not all intervals overlap. Lenient (format misses counted): highest GPT-6.1 Sol (medium) · Codex CLI 69% (95% interval 44%–86%, n 16). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–24%, n 12). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 12–16 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row. Counted calls are new calls.
Source: Harder tasks head-to-head
- Strict pass
- Format miss (correct answer, wrong format)
- Wrong answer
- No answer (timeout or error)
One square per call; counts at the right are exact and in legend order.
| Item | Strict pass | Format miss (correct answer, wrong format) | Wrong answer | No answer (timeout or error) | n |
|---|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 11 | 0 | 2 | 3 | 16 |
| Claude Opus 5.5 · Claude Code | 5 | 1 | 3 | 3 | 12 |
| Claude Sonnet 5.5 · Claude Code | 6 | 0 | 6 | 4 | 16 |
| Claude Haiku 4.5 · Claude Code | 0 | 0 | 10 | 2 | 12 |
4 rows, 4 series: Strict pass, Format miss (correct answer, wrong format), Wrong answer, No answer (timeout or error). Strict pass: highest GPT-6.1 Sol (medium) · Codex CLI 11 (n 16). Lowest Claude Haiku 4.5 · Claude Code 0 (n 12). Format miss (correct answer, wrong format): highest Claude Opus 5.5 · Claude Code 1 (n 12). Lowest Claude Haiku 4.5 · Claude Code 0 (n 12).
Notesn 12–16 per row
Counts per configuration: strict passes, format misses, wrong answers and calls with no answer
A format miss fails strictly but has an extracted answer that passes the same validator. Extra working or grids can cause this outcome. It is not a pass. A call with no answer is a timeout or an error; it counts as a non-pass.
Source: Harder tasks head-to-head
Every interval overlaps every other: this chart does not order these rows.
| Item | Tool attempt | 95% interval | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 0% | 0%–19% | 16 |
| Claude Opus 5.5 · Claude Code | 42% | 19%–68% | 12 |
| Claude Sonnet 5.5 · Claude Code | 31% | 14%–56% | 16 |
| Claude Haiku 4.5 · Claude Code | 8.3% | 1.5%–35% | 12 |
4 rows. Highest Claude Opus 5.5 · Claude Code 42% (95% interval 19%–68%, n 12). Lowest GPT-6.1 Sol (medium) · Codex CLI 0% (95% interval 0%–19%, n 16). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 12–16 per row
Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse
Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator. It is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.
Source: Harder tasks head-to-head
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | GPT-6.1 Sol (medium) · Codex CLI | Claude Opus 5.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Haiku 4.5 · Claude Code | 95% interval | n |
|---|---|---|---|---|---|---|
| 10x10 nonogram | 75% | 100% | 100% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 30%–95%; Claude Opus 5.5 · Claude Code: 44%–100%; Claude Sonnet 5.5 · Claude Code: 51%–100%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| Sudoku, 22 givens | 25% | 0% | 0% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 4.6%–70%; Claude Opus 5.5 · Claude Code: 0%–56%; Claude Sonnet 5.5 · Claude Code: 0%–49%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| 6x6 Skyscrapers | 100% | 33% | 0% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 51%–100%; Claude Opus 5.5 · Claude Code: 6.2%–79%; Claude Sonnet 5.5 · Claude Code: 0%–49%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| Seeded shuffle output | 75% | 33% | 50% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 30%–95%; Claude Opus 5.5 · Claude Code: 6.2%–79%; Claude Sonnet 5.5 · Claude Code: 15%–85%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
4 rows, 4 series: GPT-6.1 Sol (medium) · Codex CLI, Claude Opus 5.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Haiku 4.5 · Claude Code. GPT-6.1 Sol (medium) · Codex CLI: highest 6x6 Skyscrapers 100% (95% interval 51%–100%, n 4). Lowest Sudoku, 22 givens 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Claude Opus 5.5 · Claude Code: highest 10x10 nonogram 100% (95% interval 44%–100%, n 3). Lowest Sudoku, 22 givens 0% (95% interval 0%–56%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 3–4 per row
One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide
Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.
Source: Harder tasks head-to-head
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 120 s | 46.2 s–273 s | 13 |
| Claude Opus 5.5 · Claude Code | 80.3 s | 3.8 s–280 s | 9 |
| Claude Sonnet 5.5 · Claude Code | 70.4 s | 4.3 s–210 s | 12 |
| Claude Haiku 4.5 · Claude Code | 109 s | 25.7 s–224 s | 10 |
4 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 120 s (range 46.2 s–273 s, n 13). Fastest Claude Sonnet 5.5 · Claude Code 70.4 s (range 4.3 s–210 s, n 12). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 9–13 per row
Median per configuration; whiskers = fastest and slowest call
Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session. Arena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.
Source: Harder tasks head-to-head
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | Range (lowest–highest run) | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 4,994 | 4,971 | Output tokens: 2,099–13,413; Reasoning tokens: 2,070–13,372 | 13 |
| Claude Opus 5.5 · Claude Code | 8,420 | 8,352 | Output tokens: 279–40,044; Reasoning tokens: 21–9,897 | 9 |
| Claude Sonnet 5.5 · Claude Code | 9,287 | 6,557 | Output tokens: 407–27,921; Reasoning tokens: 63–27,902 | 12 |
| Claude Haiku 4.5 · Claude Code | 12,508 | 12,483 | Output tokens: 2,965–26,532; Reasoning tokens: 2,924–26,510 | 10 |
4 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 12,508 (range 2,965–26,532, n 10). Lowest GPT-6.1 Sol (medium) · Codex CLI 4,994 (range 2,099–13,413, n 13). All run ranges overlap. Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 12,483 (range 2,924–26,510, n 10). Lowest GPT-6.1 Sol (medium) · Codex CLI 4,971 (range 2,070–13,372, n 13). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 9–13 per row
Median per configuration; reasoning tokens as the CLI reports them
Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. Claude Code used an output-token cap setting of 16,000. Some reported totals exceeded it. Codex CLI had no cap. More tokens is not better or worse by itself.
Source: Harder tasks head-to-head
Hover or focus a bar for its ratio to GPT-6.1 Sol (medium) (the highlighted row): a ratio of list-price calculations, not a measurement.
| Item | Cost per strict pass | n |
|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | $0.083 | 16 |
| Claude Sonnet 5.5 · Claude Code | $0.24 | 16 |
| Claude Opus 5.5 · Claude Code | $0.59 | 12 |
List-price calculation, not a run. 3 rows. Highest Claude Opus 5.5 · Claude Code $0.59 (n 12). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.083 (n 16).
Notesn 12–16 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. These cells show a lower bound. Assume each unpriced call cost its cell’s median priced call. This sensitivity calculation gives GPT-6.1 Sol (medium) $0.100, Opus 5.5 $0.633 and Sonnet 5.5 $0.303. Opus 5.5 figures are provisional: its cache-read price is under re-check. Highlights mark the observed frontier of these lower-bound costs. Unknown timeout costs can change it; this is not a cost ranking. Claude Haiku 4.5 · Claude Code had no strict pass, so it has no cost per pass.
Sources: Harder tasks head-to-head, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Codex CLI
- Claude Code
Haloed: on the frontier (1 of 3). A point in the shaded area is no better on either axis than a haloed point.
| Point | Series | USD per strict pass (list-price calculation) | Strict pass rate | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | Codex CLI | $0.083 | 69% | 16 |
| Claude Opus 5.5 · Claude Code | Claude Code | $0.59 | 42% | 12 |
| Claude Sonnet 5.5 · Claude Code | Claude Code | $0.24 | 38% | 16 |
List-price calculation, not a run. 3 points: Strict pass rate against USD per strict pass (list-price calculation). USD per strict pass (list-price calculation) runs from $0.083 to $0.59; Strict pass rate from 38% to 69%. Highlighted: GPT-6.1 Sol (medium) · Codex CLI.
NotesWhiskers: 95% Wilson intervaln 12–16 per point
Strict pass rate against list-price cost per strict pass
Upper-left has a higher observed pass rate and lower recorded cost per pass. Highlights mark the observed frontier of lower-bound costs. Unknown timeout costs can change it. This is not a tested ranking. Frontier: GPT-6.1 Sol (medium) · Codex CLI. Costs are calculations from tokens; calls that timed out are not priced (GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16), so those cells are lower bounds. The 95% Wilson intervals are listed below; the plot shows point estimates. Strict rates: GPT-6.1 Sol (medium) 11/16 (n = 16, 95% interval 44.4% to 85.8%); Opus 5.5 5/12 (n = 12, 95% interval 19.3% to 68.0%); Sonnet 5.5 6/16 (n = 16, 95% interval 18.5% to 61.4%).
Sources: Harder tasks head-to-head, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
Tables
Cells: configuration, calls and outcomes
| Configuration | Calls | Strict passes | 95% Wilson interval | Format misses | Wrong answers | No answer | Calls with a token report | Completed calls (timing and token n) | Median completed-call total (s) | Completed-call time range (s, not an interval) | p95 total (s) | Median output tokens |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 16 | 11/16 | 44% to 86% | 0 | 2 | 3 | 13 | 13 | 120 s | 46.24 to 273.46 | 241 s | 4,994 |
| Claude Opus 5.5 · Claude Code | 12 | 5/12 | 19% to 68% | 1 | 3 | 3 | 11 | 9 | 80.3 s | 3.82 to 279.5 | 209 s | 8,420 |
| Claude Sonnet 5.5 · Claude Code | 16 | 6/16 | 18% to 61% | 0 | 6 | 4 | 12 | 12 | 70.4 s | 4.32 to 210.08 | 209 s | 9,287 |
| Claude Haiku 4.5 · Claude Code | 12 | 0/12 | 0% to 24% | 0 | 10 | 2 | 10 | 10 | 109 s | 25.73 to 223.95 | 202 s | 12,508 |
Candidate tasks, pilot result and selection (prompts not shown)
| Task | Kind | Validator checks | Pilot round | Sonnet 5.5 pilot (strict passes) | In the counted set | Controls (wrong answers rejected, wrapped reference flagged) |
|---|---|---|---|---|---|---|
| Write a TTL and LRU cache class (random operation sequences) | code | 13 | first 12 candidates | 2/2 | no | 5/5; yes |
| Write a glob matcher (braces, **, character sets, hidden files) | code | 68 | first 12 candidates | 2/2 | no | 4/4; yes |
| Write a unified diff (minimal script, fixed tie-break, hunk headers) | code | 20 | first 12 candidates | 2/2 | no | 4/4; yes |
| Evaluate Python-style integer expressions (precedence, chained comparisons) | code | 64 | first 12 candidates | 2/2 | no | 5/5; yes |
| SQLite sessions report (gaps and islands, median, logout rule) | SQL | 7 | first 12 candidates | 2/2 | no | 5/5; yes |
| SQLite as-of price and currency report (missing days, rounding) | SQL | 14 | first 12 candidates | 2/2 | no | 5/5; yes |
| Solve a 6x6 Skyscrapers puzzle (14 clues, one solution) | reasoning | 1 | first 12 candidates | 0/2 (+2 format misses) | yes | 3/3; yes |
| Pick the most profitable jobs for two machines (one optimum) | reasoning | 1 | first 12 candidates | 2/2 | no | 3/3; yes |
| Decode UTF-8 with one U+FFFD per maximal subpart | spec | 35 | first 12 candidates | 2/2 | no | 4/4; yes |
| Next run of a cron expression in UTC (day-of-month or day-of-week rule) | spec | 37 | first 12 candidates | 2/2 | no | 4/4; yes |
| Simulate a retry queue (priorities, timeouts, backoff, ties) | simulation | 1 | first 12 candidates | 2/2 | no | 4/4; yes |
| Correctly rounded sum of doubles (ties, overflow, negative zero) | numeric | 32 | first 12 candidates | 2/2 | no | 4/4; yes |
| Solve a 10x10 nonogram (one solution) | reasoning | 1 | replacement candidates | 1/2 | yes | 4/4; yes |
| Pick the most profitable jobs for three machines (30 jobs, one optimum) | reasoning | 1 | replacement candidates | 2/2 | no | 5/5; yes |
| Predict the output of a seeded shuffle (32-bit integer arithmetic) | code reading | 1 | replacement candidates | 0/2 | yes | 4/4; yes |
| Solve a 9x9 Sudoku with 22 givens (one solution) | reasoning | 1 | replacement candidates | 1/2 | yes | 3/3; yes |
Method
- The original protocol file predates the first probe and counted call. This review checked file birth times. Later changes are amendments. This follows the hard head-to-head (/benchmarks/hard-model-head-to-head), which hit a pass-rate ceiling for the strong models.
- The study started with 16 candidate tasks, written in two rounds (12 first, then 4 replacements; code, SQL, reasoning, spec, simulation, numeric). Each is one call with no tools, an exact output format and a deterministic validator that runs in a sandbox without network. Claude Sonnet 5.5 at default effort ran each candidate twice (the pilot, 32 calls). It passed 12 of the 16 candidates 2 of 2. The study dropped them.
- The dropped candidates include 10 of 10 code, SQL, spec, numeric and simulation tasks.
- The protocol declared the selection rule before the pilot. Keep at most 8 tasks: first those Sonnet passed 1 of 2, then 0 of 2. Drop tasks with 2 of 2 passes. Only 1 of the first 12 qualified, below the threshold of 6. The study wrote 4 replacement candidates once and piloted them twice; 3 qualified.
- The counted set is 4 tasks: 10x10 nonogram; Sudoku, 22 givens; 6x6 Skyscrapers; Seeded shuffle output. The study ran no second round of replacements.
- Controls ran in stages. The first 12 candidates had controls before the first probe. After a module-export validator defect, controls ran again and stored pilot replies were re-scored without new calls. Final controls covered all 16 candidates before the replacement pilot and all counted calls. 16/16 references pass. 66/66 wrong answers fail. 16/16 wrapped references are format misses. Independent solvers checked unique reasoning answers. Python integers confirmed the code-reading answer.
- Counted cells: GPT-6.1 Sol (medium) · Codex CLI (16 calls, 4 per task). Claude Opus 5.5 · Claude Code (12 calls, 3 per task). Claude Sonnet 5.5 · Claude Code (16 calls, 4 per task). Claude Haiku 4.5 · Claude Code (12 calls, 3 per task). Rep-major, round-robin order, one call at a time per account, 300 s timeout.
- Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no other text. Format miss: the strict check fails, but an extracted answer passes the same validator. The extractor reads fenced blocks, answer lines or grid rows. For exact tasks it reads only the last answer-shaped candidate, not an earlier guess. Reported apart from wrong answers, never as a pass.
- An error or timeout counts as a non-pass. Tool use is a separate flag, not a score. It marks tool-call markup or a CLI tool-call parse error.
- Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn. The Codex runner adds a developer instruction to answer directly and not to call tools. The Claude Code runner adds no such instruction. Claude Code ran with an output-token cap setting of 16,000; the Codex CLI had none.
- Default effort means the effort flag was not passed. GPT-6.1 Sol ran at medium effort.
- Counted calls: Claude Code: 40 counted calls, 40 reached a model; Codex CLI: 16 counted calls, 16 reached a model. The study trimmed no calls. The study repeated no counted key. No run reported a usage or rate limit. The Claude CLI reported its own failed retry on tool-call parse errors; those receipts remain failures. The Claude batch stopped 1 time on a tool-call parse error. A later amendment kept that failure but allowed the lane to continue. The lane repeated no counted key.
- Uncounted probes: 2, kept apart from 32 pilot calls and 56 counted attempts. Call caps, including probes and pilots: Claude Code: 73/105 calls; Codex CLI: 17/33 calls.
- Cost per strict pass is a calculation: total reported-token cost divided by strict passes. It includes failed calls with tokens and prices cache reads and writes. Timeouts report no tokens, so the cost is a lower bound.
Caveats
- Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.
- Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
- Small samples: GPT-6.1 Sol (medium) n = 16, Opus 5.5 n = 12, Sonnet 5.5 n = 16, Haiku 4.5 n = 12. Per-task cells have 3 to 4 calls. The intervals are wide, and the calls on one task are not independent, so the intervals are likely narrower than the truth.
- Each row pairs a CLI with a model on one subscription. The CLIs use different subscriptions, system prompts and start-up steps. A Claude-vs-GPT row compares route + model pairs, not the models alone.
- Both CLIs disabled tools. Every prompt asks for the answer only. The runners differ in one way: the Codex runner adds a developer instruction not to call tools, and the Claude Code runner adds none. 11 of 40 Claude Code calls tried a tool anyway. By model: Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12. None of them passed. None of the 16 Codex CLI calls did. We did not test the instruction, so we cannot say how much of the gap it explains. With tools on, the Claude Code models might have run code and passed on these tasks. This study does not test that either.
- The runner stopped 10 of 56 calls at the 300 s timeout and counted them as no answer. Counts: GPT-6.1 Sol (medium) 3, Opus 5.5 1, Sonnet 5.5 4 and Haiku 4.5 2. A longer limit might turn some into answers, passes or fails. Timing medians cover completed calls only. They include wrong answers and format misses. They omit only timeouts and tool-call parse errors in this run.
- Claude Code ran with an output-token cap setting of 16,000. Reported totals exceeded 16,000 in 11 of 40 Claude Code calls. The setting did not bound reported totals. We cannot tell whether it cut any reply. The tool-call parse errors came at 16,746 and 16,859 reported output tokens. The Codex CLI calls had no such cap.
- The kept tasks are mostly long exact-search or computation tasks. In the pilot Claude Sonnet 5.5 passed 10 of 10 code, SQL, spec, numeric and simulation candidates 2 of 2. This set says little about everyday coding.
- Claude Haiku 4.5 · Claude Code passed none: a floor on this set, not a general rating.
- Strict format rules decide part of the result: a reply with extra text fails. The chart shows the lenient reading next to the strict score.
- Arena servers shared the Mac during part of the run. Host load was not controlled. CLI timings include start-up and each CLI’s system prompt. These times do not isolate model speed.
- The protocol first said to drop tasks with validator defects. An amendment instead repaired the JavaScript validators and re-scored stored pilot replies. None of those code-writing tasks entered the counted set.
- No configuration passed every call across the full set. Some per-task cells hit a ceiling: Sonnet and Opus on the nonogram, and Sol on Skyscrapers. Their perfect cells do not prove equal ability.
- Wilson intervals treat calls as independent. The same four tasks repeat, so these are descriptive call-level intervals, not population intervals for coding tasks. The study ran no task-level paired significance test.
- Default effort means the effort flag was not passed; the CLI chose. Median reasoning tokens per completed call, where the CLI reports them (n and ranges are in the token chart): GPT-6.1 Sol (medium) 4,971, Opus 5.5 8,352, Sonnet 5.5 6,557, Haiku 4.5 12,483. They are part of the output tokens.
- List-price costs are calculations; the calls used flat subscriptions. Opus 5.5 figures are provisional: its cache-read price is under re-check. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. Those cells show a lower bound on cost per pass.
Sources
Harder tasks head-to-head
Frozen tasks selected with a Sonnet pilot. Fresh counted calls retain failures, format misses and timeouts. Replies and expected answers are not published.
Repricing calculation
Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Token prices as listed by the vendor on 2026-10-03.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks”, updated October 7, 2026, https://agent.sasid.ai/benchmarks/harder-tasks-head-to-head.
Models and comparisons in this study
Write-ups on this study
GPT-6.1 Sol vs Claude Opus 5.5 and Sonnet 5.5 on harder tasks: only one strict pair separates
GPT-6.1 Sol vs Claude Opus and Sonnet on four selected harder tasks. Pass intervals overlap for these models; only Sol vs Haiku separates on strict passes.
More studies
All benchmarksHaiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.