• Head to head
  • Harder Tasks
  • GPT-6.1 Sol
  • Codex CLI
  • Claude Sonnet
  • Claude Opus
  • Claude Haiku
  • Reasoning
  • Selection Effect
  • Format Misses
  • Latency

GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

On a task set built so that Claude Sonnet 5.5 did not pass it every time, does pass rate separate GPT-6.1 Sol (Codex CLI) from Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 (Claude Code)?

Published · 8 charts · Download the data or a carousel

39%

95% CI 28%–52% · n = 56

22/56 · Counted calls that passed strictly (harder set)

The answer

22 of 56 counted calls passed strictly (39%, 95% Wilson interval 28% to 52%) on 4 tasks. The Sonnet pilot did not pass these twice. GPT-6.1 Sol (medium) passed 11/16 strictly (69%, 95% interval 44% to 86%). Opus 5.5 passed 5/12 strictly (42%, 95% interval 19% to 68%). Sonnet 5.5 passed 6/16 strictly (38%, 95% interval 18% to 61%). Haiku 4.5 passed 0/12 strictly (0%, 95% interval 0% to 24%). GPT-6.1 Sol (medium) is ahead of Haiku 4.5 (the 95% intervals do not overlap). The other 5 of 6 pairs overlap, so this set cannot rank them. The lenient reading counts format misses. Opus 5.5 is also ahead of Haiku 4.5 on the lenient reading. Opus 5.5 6/12 (n = 12, 95% Wilson interval 25.4% to 74.6%); Haiku 4.5 0/12 (n = 12, 95% Wilson interval 0.0% to 24.2%). The intervals miss by 1.1 points (calculation), so this is fragile. Haiku 4.5 passed none (a floor for that configuration on this set). No configuration passed every call across the full set. Some per-task cells still hit a ceiling; see the task chart. Of 34 non-passes, 1 was a format miss with the right answer. Another 21 were wrong answers. 12 gave no answer; 10 hit the 300 s timeout. 11 of 56 calls tried a tool although tools were off. Per configuration: GPT-6.1 Sol (medium) 0 of 16, Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12. None of them passed. The Codex runner also tells the model not to call tools; the Claude Code runner does not. Median total time per completed call (wrong answers and format misses included; timeouts and tool-call parse errors excluded; ranges are not intervals): GPT-6.1 Sol (medium) 120.2 s (n = 13, range 46.2 s to 273.5 s). Opus 5.5 80.3 s (n = 9, range 3.8 s to 279.5 s). Sonnet 5.5 70.4 s (n = 12, range 4.3 s to 210.1 s). Haiku 4.5 109.0 s (n = 10, range 25.7 s to 223.9 s). Every configuration’s fastest-to-slowest range overlaps every other, so the medians describe this run and are not a tested ranking. Cost is a list-price calculation; the calls ran on subscriptions. The lowest recorded lower bound was GPT-6.1 Sol (medium) · Codex CLI: $0.083 (a lower bound: 3 timed-out calls report no tokens; $0.100 if each had cost a median call, an assumption). Selection effect: the study picked tasks with mixed or failed Sonnet pilot results. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls. The counted calls are new calls. Selection can produce this pattern; the run does not establish its cause.

Key numbers

41% (23/56)

Counted calls with a correct answer, format misses included (lenient reading)

95% CI 29%–54% · n = 56

69% (11/16)

GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks

95% CI 44%–86% · n = 16

42% (5/12)

Opus 5.5 (Claude Code): strict pass rate on the harder tasks

95% CI 19%–68% · n = 12

38% (6/16)

Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks

95% CI 18%–61% · n = 16

0% (0/12)

Haiku 4.5 (Claude Code): strict pass rate on the harder tasks

95% CI 0%–24% · n = 12

1

Non-passes that were format misses, not wrong answers

of 34 non-passes (21 wrong answers, 12 no answer) · n = 34

20% (11/56)

Counted calls that tried a tool although tools were off (a behaviour, not a quality score)

95% CI 11%–32% · n = 56

18% (10/56)

Counted calls that ran past the 300 s limit and gave no answer

95% CI 10%–30% · n = 56

81% (26/32)

Pilot (not scored): strict passes of Claude Sonnet 5.5 on all candidate tasks

95% CI 65%–91% · n = 32

25% (2/8)

Pilot (not scored): strict passes of Claude Sonnet 5.5 on the tasks that were kept

95% CI 7.1%–59% · n = 8

4

Candidate tasks kept for the counted set

of 16 (Sonnet passed 12 candidates 2 of 2, including 10 of 10 code, SQL, spec, numeric and simulation tasks) · n = 16

$0.0829

GPT-6.1 Sol (medium) · Codex CLI

Lowest recorded cost lower bound per strict pass (calculation)

(lower bound) · n = 16

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

  • Strict pass
  • Lenient (format misses counted)
GPT-6.1 Sol (medium)
Claude Opus 5.5
Claude Sonnet 5.5
Claude Haiku 4.5

4 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest GPT-6.1 Sol (medium) · Codex CLI 69% (95% interval 44%–86%, n 16). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–24%, n 12). Not all intervals overlap. Lenient (format misses counted): highest GPT-6.1 Sol (medium) · Codex CLI 69% (95% interval 44%–86%, n 16). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–24%, n 12). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 12–16 per row

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row. Counted calls are new calls.

Source: Harder tasks head-to-head

Share card (PNG)
  • Strict pass
  • Format miss (correct answer, wrong format)
  • Wrong answer
  • No answer (timeout or error)
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

One square per call; counts at the right are exact and in legend order.

4 rows, 4 series: Strict pass, Format miss (correct answer, wrong format), Wrong answer, No answer (timeout or error). Strict pass: highest GPT-6.1 Sol (medium) · Codex CLI 11 (n 16). Lowest Claude Haiku 4.5 · Claude Code 0 (n 12). Format miss (correct answer, wrong format): highest Claude Opus 5.5 · Claude Code 1 (n 12). Lowest Claude Haiku 4.5 · Claude Code 0 (n 12).

Notesn 12–16 per row

Counts per configuration: strict passes, format misses, wrong answers and calls with no answer

A format miss fails strictly but has an extracted answer that passes the same validator. Extra working or grids can cause this outcome. It is not a pass. A call with no answer is a timeout or an error; it counts as a non-pass.

Source: Harder tasks head-to-head

Share card (PNG)
GPT-6.1 Sol (medium)
Claude Opus 5.5
Claude Sonnet 5.5
Claude Haiku 4.5

Every interval overlaps every other: this chart does not order these rows.

4 rows. Highest Claude Opus 5.5 · Claude Code 42% (95% interval 19%–68%, n 12). Lowest GPT-6.1 Sol (medium) · Codex CLI 0% (95% interval 0%–19%, n 16). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 12–16 per row

Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse

Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator. It is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.

Source: Harder tasks head-to-head

Share card (PNG)

GPT-6.1 Sol (medium) · Codex CLI

10x10 nonogram
Sudoku, 22 givens
6x6 Skyscrapers
Seeded shuffle output

Claude Opus 5.5 · Claude Code

10x10 nonogram
Sudoku, 22 givens
6x6 Skyscrapers
Seeded shuffle output

Claude Sonnet 5.5 · Claude Code

10x10 nonogram
Sudoku, 22 givens
6x6 Skyscrapers
Seeded shuffle output

Claude Haiku 4.5 · Claude Code

10x10 nonogram
Sudoku, 22 givens
6x6 Skyscrapers
Seeded shuffle output

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

4 rows, 4 series: GPT-6.1 Sol (medium) · Codex CLI, Claude Opus 5.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Haiku 4.5 · Claude Code. GPT-6.1 Sol (medium) · Codex CLI: highest 6x6 Skyscrapers 100% (95% interval 51%–100%, n 4). Lowest Sudoku, 22 givens 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Claude Opus 5.5 · Claude Code: highest 10x10 nonogram 100% (95% interval 44%–100%, n 3). Lowest Sudoku, 22 givens 0% (95% interval 0%–56%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 3–4 per row

One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide

Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.

Source: Harder tasks head-to-head

Share card (PNG)
Entrance: medians race at 86× real time
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

4 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 120 s (range 46.2 s–273 s, n 13). Fastest Claude Sonnet 5.5 · Claude Code 70.4 s (range 4.3 s–210 s, n 12). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 9–13 per row

Median per configuration; whiskers = fastest and slowest call

Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session. Arena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.

Source: Harder tasks head-to-head

Share card (PNG)
  • Output tokens
  • of which reasoning tokens (inner bar)
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code

4 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 12,508 (range 2,965–26,532, n 10). Lowest GPT-6.1 Sol (medium) · Codex CLI 4,994 (range 2,099–13,413, n 13). All run ranges overlap. Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 12,483 (range 2,924–26,510, n 10). Lowest GPT-6.1 Sol (medium) · Codex CLI 4,971 (range 2,070–13,372, n 13). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 9–13 per row

Median per configuration; reasoning tokens as the CLI reports them

Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. Claude Code used an output-token cap setting of 16,000. Some reported totals exceeded it. Codex CLI had no cap. More tokens is not better or worse by itself.

Source: Harder tasks head-to-head

Share card (PNG)
Calculation
GPT-6.1 Sol (medium) · Codex CLI
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code

Hover or focus a bar for its ratio to GPT-6.1 Sol (medium) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Opus 5.5 · Claude Code $0.59 (n 12). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.083 (n 16).

Notesn 12–16 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. These cells show a lower bound. Assume each unpriced call cost its cell’s median priced call. This sensitivity calculation gives GPT-6.1 Sol (medium) $0.100, Opus 5.5 $0.633 and Sonnet 5.5 $0.303. Opus 5.5 figures are provisional: its cache-read price is under re-check. Highlights mark the observed frontier of these lower-bound costs. Unknown timeout costs can change it; this is not a cost ranking. Claude Haiku 4.5 · Claude Code had no strict pass, so it has no cost per pass.

Sources: Harder tasks head-to-head, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
  • Codex CLI
  • Claude Code
Better: upper left

Haloed: on the frontier (1 of 3). A point in the shaded area is no better on either axis than a haloed point.

List-price calculation, not a run. 3 points: Strict pass rate against USD per strict pass (list-price calculation). USD per strict pass (list-price calculation) runs from $0.083 to $0.59; Strict pass rate from 38% to 69%. Highlighted: GPT-6.1 Sol (medium) · Codex CLI.

NotesWhiskers: 95% Wilson intervaln 12–16 per point

Strict pass rate against list-price cost per strict pass

Upper-left has a higher observed pass rate and lower recorded cost per pass. Highlights mark the observed frontier of lower-bound costs. Unknown timeout costs can change it. This is not a tested ranking. Frontier: GPT-6.1 Sol (medium) · Codex CLI. Costs are calculations from tokens; calls that timed out are not priced (GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16), so those cells are lower bounds. The 95% Wilson intervals are listed below; the plot shows point estimates. Strict rates: GPT-6.1 Sol (medium) 11/16 (n = 16, 95% interval 44.4% to 85.8%); Opus 5.5 5/12 (n = 12, 95% interval 19.3% to 68.0%); Sonnet 5.5 6/16 (n = 16, 95% interval 18.5% to 61.4%).

Sources: Harder tasks head-to-head, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)

Tables

Cells: configuration, calls and outcomes

ConfigurationCallsStrict passes95% Wilson intervalFormat missesWrong answersNo answerCalls with a token reportCompleted calls (timing and token n)Median completed-call total (s)Completed-call time range (s, not an interval)p95 total (s)Median output tokens
GPT-6.1 Sol (medium) · Codex CLI1611/1644% to 86%0231313120 s46.24 to 273.46241 s4,994
Claude Opus 5.5 · Claude Code125/1219% to 68%13311980.3 s3.82 to 279.5209 s8,420
Claude Sonnet 5.5 · Claude Code166/1618% to 61%064121270.4 s4.32 to 210.08209 s9,287
Claude Haiku 4.5 · Claude Code120/120% to 24%01021010109 s25.73 to 223.95202 s12,508

Candidate tasks, pilot result and selection (prompts not shown)

TaskKindValidator checksPilot roundSonnet 5.5 pilot (strict passes)In the counted setControls (wrong answers rejected, wrapped reference flagged)
Write a TTL and LRU cache class (random operation sequences)code13first 12 candidates2/2no5/5; yes
Write a glob matcher (braces, **, character sets, hidden files)code68first 12 candidates2/2no4/4; yes
Write a unified diff (minimal script, fixed tie-break, hunk headers)code20first 12 candidates2/2no4/4; yes
Evaluate Python-style integer expressions (precedence, chained comparisons)code64first 12 candidates2/2no5/5; yes
SQLite sessions report (gaps and islands, median, logout rule)SQL7first 12 candidates2/2no5/5; yes
SQLite as-of price and currency report (missing days, rounding)SQL14first 12 candidates2/2no5/5; yes
Solve a 6x6 Skyscrapers puzzle (14 clues, one solution)reasoning1first 12 candidates0/2 (+2 format misses)yes3/3; yes
Pick the most profitable jobs for two machines (one optimum)reasoning1first 12 candidates2/2no3/3; yes
Decode UTF-8 with one U+FFFD per maximal subpartspec35first 12 candidates2/2no4/4; yes
Next run of a cron expression in UTC (day-of-month or day-of-week rule)spec37first 12 candidates2/2no4/4; yes
Simulate a retry queue (priorities, timeouts, backoff, ties)simulation1first 12 candidates2/2no4/4; yes
Correctly rounded sum of doubles (ties, overflow, negative zero)numeric32first 12 candidates2/2no4/4; yes
Solve a 10x10 nonogram (one solution)reasoning1replacement candidates1/2yes4/4; yes
Pick the most profitable jobs for three machines (30 jobs, one optimum)reasoning1replacement candidates2/2no5/5; yes
Predict the output of a seeded shuffle (32-bit integer arithmetic)code reading1replacement candidates0/2yes4/4; yes
Solve a 9x9 Sudoku with 22 givens (one solution)reasoning1replacement candidates1/2yes3/3; yes

Method

  1. The original protocol file predates the first probe and counted call. This review checked file birth times. Later changes are amendments. This follows the hard head-to-head (/benchmarks/hard-model-head-to-head), which hit a pass-rate ceiling for the strong models.
  2. The study started with 16 candidate tasks, written in two rounds (12 first, then 4 replacements; code, SQL, reasoning, spec, simulation, numeric). Each is one call with no tools, an exact output format and a deterministic validator that runs in a sandbox without network. Claude Sonnet 5.5 at default effort ran each candidate twice (the pilot, 32 calls). It passed 12 of the 16 candidates 2 of 2. The study dropped them.
  3. The dropped candidates include 10 of 10 code, SQL, spec, numeric and simulation tasks.
  4. The protocol declared the selection rule before the pilot. Keep at most 8 tasks: first those Sonnet passed 1 of 2, then 0 of 2. Drop tasks with 2 of 2 passes. Only 1 of the first 12 qualified, below the threshold of 6. The study wrote 4 replacement candidates once and piloted them twice; 3 qualified.
  5. The counted set is 4 tasks: 10x10 nonogram; Sudoku, 22 givens; 6x6 Skyscrapers; Seeded shuffle output. The study ran no second round of replacements.
  6. Controls ran in stages. The first 12 candidates had controls before the first probe. After a module-export validator defect, controls ran again and stored pilot replies were re-scored without new calls. Final controls covered all 16 candidates before the replacement pilot and all counted calls. 16/16 references pass. 66/66 wrong answers fail. 16/16 wrapped references are format misses. Independent solvers checked unique reasoning answers. Python integers confirmed the code-reading answer.
  7. Counted cells: GPT-6.1 Sol (medium) · Codex CLI (16 calls, 4 per task). Claude Opus 5.5 · Claude Code (12 calls, 3 per task). Claude Sonnet 5.5 · Claude Code (16 calls, 4 per task). Claude Haiku 4.5 · Claude Code (12 calls, 3 per task). Rep-major, round-robin order, one call at a time per account, 300 s timeout.
  8. Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no other text. Format miss: the strict check fails, but an extracted answer passes the same validator. The extractor reads fenced blocks, answer lines or grid rows. For exact tasks it reads only the last answer-shaped candidate, not an earlier guess. Reported apart from wrong answers, never as a pass.
  9. An error or timeout counts as a non-pass. Tool use is a separate flag, not a score. It marks tool-call markup or a CLI tool-call parse error.
  10. Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn. The Codex runner adds a developer instruction to answer directly and not to call tools. The Claude Code runner adds no such instruction. Claude Code ran with an output-token cap setting of 16,000; the Codex CLI had none.
  11. Default effort means the effort flag was not passed. GPT-6.1 Sol ran at medium effort.
  12. Counted calls: Claude Code: 40 counted calls, 40 reached a model; Codex CLI: 16 counted calls, 16 reached a model. The study trimmed no calls. The study repeated no counted key. No run reported a usage or rate limit. The Claude CLI reported its own failed retry on tool-call parse errors; those receipts remain failures. The Claude batch stopped 1 time on a tool-call parse error. A later amendment kept that failure but allowed the lane to continue. The lane repeated no counted key.
  13. Uncounted probes: 2, kept apart from 32 pilot calls and 56 counted attempts. Call caps, including probes and pilots: Claude Code: 73/105 calls; Codex CLI: 17/33 calls.
  14. Cost per strict pass is a calculation: total reported-token cost divided by strict passes. It includes failed calls with tokens and prices cache reads and writes. Timeouts report no tokens, so the cost is a lower bound.

Caveats

  • Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.
  • Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
  • Small samples: GPT-6.1 Sol (medium) n = 16, Opus 5.5 n = 12, Sonnet 5.5 n = 16, Haiku 4.5 n = 12. Per-task cells have 3 to 4 calls. The intervals are wide, and the calls on one task are not independent, so the intervals are likely narrower than the truth.
  • Each row pairs a CLI with a model on one subscription. The CLIs use different subscriptions, system prompts and start-up steps. A Claude-vs-GPT row compares route + model pairs, not the models alone.
  • Both CLIs disabled tools. Every prompt asks for the answer only. The runners differ in one way: the Codex runner adds a developer instruction not to call tools, and the Claude Code runner adds none. 11 of 40 Claude Code calls tried a tool anyway. By model: Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12. None of them passed. None of the 16 Codex CLI calls did. We did not test the instruction, so we cannot say how much of the gap it explains. With tools on, the Claude Code models might have run code and passed on these tasks. This study does not test that either.
  • The runner stopped 10 of 56 calls at the 300 s timeout and counted them as no answer. Counts: GPT-6.1 Sol (medium) 3, Opus 5.5 1, Sonnet 5.5 4 and Haiku 4.5 2. A longer limit might turn some into answers, passes or fails. Timing medians cover completed calls only. They include wrong answers and format misses. They omit only timeouts and tool-call parse errors in this run.
  • Claude Code ran with an output-token cap setting of 16,000. Reported totals exceeded 16,000 in 11 of 40 Claude Code calls. The setting did not bound reported totals. We cannot tell whether it cut any reply. The tool-call parse errors came at 16,746 and 16,859 reported output tokens. The Codex CLI calls had no such cap.
  • The kept tasks are mostly long exact-search or computation tasks. In the pilot Claude Sonnet 5.5 passed 10 of 10 code, SQL, spec, numeric and simulation candidates 2 of 2. This set says little about everyday coding.
  • Claude Haiku 4.5 · Claude Code passed none: a floor on this set, not a general rating.
  • Strict format rules decide part of the result: a reply with extra text fails. The chart shows the lenient reading next to the strict score.
  • Arena servers shared the Mac during part of the run. Host load was not controlled. CLI timings include start-up and each CLI’s system prompt. These times do not isolate model speed.
  • The protocol first said to drop tasks with validator defects. An amendment instead repaired the JavaScript validators and re-scored stored pilot replies. None of those code-writing tasks entered the counted set.
  • No configuration passed every call across the full set. Some per-task cells hit a ceiling: Sonnet and Opus on the nonogram, and Sol on Skyscrapers. Their perfect cells do not prove equal ability.
  • Wilson intervals treat calls as independent. The same four tasks repeat, so these are descriptive call-level intervals, not population intervals for coding tasks. The study ran no task-level paired significance test.
  • Default effort means the effort flag was not passed; the CLI chose. Median reasoning tokens per completed call, where the CLI reports them (n and ranges are in the token chart): GPT-6.1 Sol (medium) 4,971, Opus 5.5 8,352, Sonnet 5.5 6,557, Haiku 4.5 12,483. They are part of the output tokens.
  • List-price costs are calculations; the calls used flat subscriptions. Opus 5.5 figures are provisional: its cache-read price is under re-check. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. Those cells show a lower bound on cost per pass.

Sources

  • Harder tasks head-to-head

    Our recorded runs ·

    Frozen tasks selected with a Sonnet pilot. Fresh counted calls retain failures, format misses and timeouts. Replies and expected answers are not published.

    Raw data: harder-tasks/receipts.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks”, updated October 7, 2026, https://agent.sasid.ai/benchmarks/harder-tasks-head-to-head.

Models and comparisons in this study

More studies

All benchmarks
Live story
  • Head to head
  • Hard Tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.

91% (139/152)Calls that passed strictly (hard set) · n = 152

7 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.