GPT-6.1 Sol vs Claude Opus 5.5 and Sonnet 5.5 on harder tasks: only one strict pair separates
GPT-6.1 Sol vs Claude Opus and Sonnet on four selected harder tasks. Pass intervals overlap for these models; only Sol vs Haiku separates on strict passes.
TL;DR
- Sol vs Opus: no winner on this set. On four hard tasks, GPT-6.1 Sol (Codex CLI, medium effort) passed 11 of 16 calls (69%, 95% interval 44% to 86%). Claude Opus 5.5 (Claude Code) passed 5 of 12 (42%, 19% to 68%). The intervals overlap, so four tasks cannot rank them.
- One strict pair does separate. Sol is ahead of Claude Haiku 4.5, which passed 0 of 12 (0% to 24%). Claude Sonnet 5.5 passed 6 of 16 (38%, 18% to 61%). All other pairs overlap.
- Many Claude Code calls tried a tool that was off. 11 of 40 Claude Code calls tried a tool (95% interval 16% to 43%), although tools were off and each prompt asked for the answer only. None of the 11 passed. 0 of 16 Codex CLI calls did (95% interval 0% to 19%). One difference: the Codex runner also tells the model not to call tools, and the Claude Code runner does not. We tested neither tools on nor that instruction.
- The time limit decided some calls. 10 of 56 calls ran past 300 s (95% interval 10% to 30%) and count as no answer: Sonnet 4, Sol 3, Haiku 2, Opus 1.
- Cost per strict pass (calculation): Sol $0.083, Sonnet $0.238, Opus $0.593 (provisional). Timed-out calls report no tokens, so each figure is a lower bound. In a sensitivity calculation, if each timed-out call cost a median priced call, they become $0.100, $0.303 and $0.633. Haiku had no pass.
- We chose the tasks against Sonnet. On the kept tasks Sonnet passed 2 of 8 pilot calls (95% interval 7% to 59%). It passed 6 of 16 counted calls (18% to 61%). Selection can produce this pattern. The intervals overlap, and this run cannot establish its cause.
- Most task types did not get hard. Of 16 candidate tasks, Sonnet passed 12 twice in the pilot. That includes all 10 code, SQL, spec, numeric and simulation candidates.
Study, every call and every interval: harder tasks head-to-head. It follows the hard head-to-head, where pass rate could not separate the configurations that passed every call.
The short answer
Sol had the highest pass rate here. That is a point estimate, not a ranking. With 12 to 16 calls per model, only one gap is large enough: Sol against Haiku. Sol against Opus, Sol against Sonnet and Sonnet against Opus all overlap. The comparison pages show the same: the pass-rate rows from this study say "tie" on Sonnet vs Sol, Opus vs Sol and Sonnet vs Opus, and name Sol as ahead on Haiku vs Sol.
One more reading is fragile. If you count Opus's one format miss, Opus passed 6 of 12 (25% to 75%) and is also ahead of Haiku (0 of 12, 0% to 24%). The two intervals miss each other by 1.1 percentage points (calculation), so do not lean on it.
| Configuration | Calls | Strict passes | 95% interval | Format misses | Wrong | No answer | Tried a tool |
|---|---|---|---|---|---|---|---|
| GPT-6.1 Sol (medium), Codex CLI | 16 | 11 | 44% to 86% | 0 | 2 | 3 | 0 |
| Claude Opus 5.5, Claude Code | 12 | 5 | 19% to 68% | 1 | 3 | 3 | 5 |
| Claude Sonnet 5.5, Claude Code | 16 | 6 | 18% to 61% | 0 | 6 | 4 | 5 |
| Claude Haiku 4.5, Claude Code | 12 | 0 | 0% to 24% | 0 | 10 | 2 | 1 |
A strict pass needs the whole reply to match the validator. A format miss has the right answer in the wrong wrapper. It never counts as a pass. The "tried a tool" column overlaps the other columns: a tool attempt is a wrong answer or no answer.
Why we needed new tasks
The earlier hard set hit a ceiling. Sonnet, Opus and Fable each passed 24 of 24 (95% interval 86% to 100%). Sol passed 16 of 16 (81% to 100%). A ceiling cannot separate models.
So we wrote 12 new candidate tasks: code, SQL, reasoning, spec, simulation and numeric. Each has a validator that runs in a sandbox. Initial controls covered these 12 candidates before the first probe. A module-export defect appeared in the first pilot call. We repaired the JavaScript validators, reran controls and re-scored stored pilot replies without new calls. Final controls covered all 16 candidates before the replacement pilot and counted calls.
Then we piloted Sonnet twice on each candidate. We kept a task only if Sonnet passed it 0 or 1 of 2 times. Only 1 of the 12 qualified. Sonnet passed every code, SQL, spec, numeric and simulation candidate twice. So we wrote four replacements once, with long exact search or computation, and piloted them the same way. Three qualified.
The counted set is four tasks:
| Task | Kind | Sonnet pilot (strict passes) | 95% interval |
|---|---|---|---|
| 10x10 nonogram, one solution | reasoning | 1/2 | 9% to 91% |
| Sudoku with 22 givens, one solution | reasoning | 1/2 | 9% to 91% |
| 6x6 Skyscrapers, 14 clues, one solution | reasoning | 0/2 (+2 format misses) | 0% to 66% |
| Output of a seeded shuffle (32-bit integer arithmetic) | code reading | 0/2 | 0% to 66% |
Each task is one call, no tools, an exact output format and a 300 s limit. Counted calls are new calls: nothing from the pilot is in the counts.
What went wrong for Claude Code: tool attempts
We read the failures before we wrote this post. Many Claude Code replies were a tool call, not an answer. A typical one is a Bash call that writes and runs a script. Tools were off in both CLIs and every prompt asked for the answer only.
Sonnet tried a tool on 5 of 16 calls (95% interval 14% to 56%). Opus did so on 5 of 12 (19% to 68%). Haiku did so on 1 of 12 (1% to 35%). Sol did so on 0 of 16 (0% to 19%). In two Opus calls, the CLI reported that it could not parse the tool call, and the call ended without an answer. None of the 11 calls with a tool attempt passed.
This changes how to read the gap. It may come in part from this one-call, no-tools setup. A Claude model that may run code could do better. We did not test that, so we do not claim it.
There is a second difference between the two runners. The Codex runner adds a developer instruction to every call: answer directly and do not call tools. The Claude Code runner turns the tools off but adds no such instruction. So Sol's 0 of 16 and Claude's 11 of 40 may differ in part because of that instruction. We did not test it, so we cannot say how much it explains.
By task
Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third, and the intervals overlap almost everywhere.
- Nonogram. Sonnet 4/4 (95% interval 51% to 100%) and Opus 3/3 (44% to 100%) passed every call. Sol passed 3/4 (30% to 95%) and Haiku 0/3 (0% to 56%). These are point estimates, with overlapping intervals. Sonnet and Opus hit a ceiling on this task; that does not prove equal ability.
- Skyscrapers. Sol passed 4/4 (95% interval 51% to 100%), a ceiling in this cell. Opus passed 1/3 (6% to 79%), plus one format miss. Sonnet passed 0/4 (0% to 49%); three timed out. Haiku passed 0/3 (0% to 56%).
- Sudoku. Only Sol passed, 1 of 4 (95% interval 5% to 70%). Its other three calls timed out. The Claude Code models passed 0 of 10 (0% to 28%): Sonnet and Opus mostly tried a tool, and Haiku gave wrong grids.
- Seeded shuffle. Sol 3/4 (95% interval 30% to 95%), Sonnet 2/4 (15% to 85%), Opus 1/3 (6% to 79%), Haiku 0/3 (0% to 56%). The Sonnet and Opus misses were all tool attempts.
Speed and tokens
Median total time per completed call, with the fastest and slowest call (ranges, not intervals): Sonnet 70.4 s (n = 12, 4.3 s to 210.1 s), Opus 80.3 s (n = 9, 3.8 s to 279.5 s), Haiku 109.0 s (n = 10, 25.7 s to 223.9 s), Sol 120.2 s (n = 13, 46.2 s to 273.5 s). All ranges overlap, so the medians describe this run and are not a ranking. A range is not a confidence interval.
The medians exclude 10 timeouts and two Opus tool-call parse errors. They include wrong answers and format misses. Arena servers shared the host during part of the run. Host load was not controlled, so these times do not isolate model speed. A longer limit might turn some timeouts into answers. Sol's only Sudoku pass took 273.5 s of the 300 s.
Median tokens per completed call, with minimum-to-maximum ranges:
| Configuration | Completed n | Output median (range) | Reported reasoning median (range) |
|---|---|---|---|
| Sol | 13 | 4,994 (2,099 to 13,413) | 4,971 (2,070 to 13,372) |
| Opus | 9 | 8,420 (279 to 40,044) | 8,352 (21 to 9,897) |
| Sonnet | 12 | 9,287 (407 to 27,921) | 6,557 (63 to 27,902) |
| Haiku | 10 | 12,508 (2,965 to 26,532) | 12,483 (2,924 to 26,510) |
Reasoning tokens are part of output where the CLI reports them. Ranges are not confidence intervals. More tokens is not better or worse by itself.
Price per pass (a calculation)
Each figure is reported tokens times list price, for every call that reported tokens, divided by the strict passes. The calls ran on subscriptions, so this is not a bill. A failed call still costs, so a low pass rate raises the cost per pass.
Sol cost $0.083 per strict pass, Sonnet $0.238 and Opus $0.593. Opus figures are provisional: its cache-read price is under re-check. Haiku had no pass, so it has no cost per pass.
These are lower bounds. A call that timed out reports no tokens, so we could not price it. That affects Sol (3 of 16 calls), Sonnet (4 of 16) and Opus (1 of 12). In a sensitivity calculation, we price each like its configuration's median priced call (an assumption). Then the figures become Sol $0.100, Sonnet $0.303 and Opus $0.633. The point-estimate order stays the same under that assumption. Unknown timeout costs can change it; this is not a tested cost ranking. The cost rows have no interval, so the comparison pages mark them "unclear", not as a winner.
Which one to pick
This is our reading, not a ranking.
- You need exact, multi-step answers in one call, with no tools. Sol had the highest observed pass rate: 11 of 16 (95% interval 44% to 86%). Its calculated cost per pass had the lowest recorded lower bound. Its interval overlaps Opus and Sonnet, so test your own tasks before you commit. Allow for long calls: 3 of its 16 calls hit the limit.
- You can give the model tools. This study does not apply. Claude Code tried to use tools on 11 of 40 calls (95% interval 16% to 43%). Test with tools on.
- You do everyday coding, SQL or parsing. This study says little. Sonnet passed all 10 such candidates twice in our pilot, so they did not separate the models.
- You consider Haiku 4.5 at default effort. It passed 0 of 12 here (95% interval 0% to 24%). Test your own tasks before choosing it for work like this.
How to build a task set that separates models
What we would repeat:
- Run controls first. Every reference passes, every wrong answer fails, and a wrapped reference is flagged as a format miss. The initial controls missed a module-export defect; the first pilot call exposed it. Later controls passed 16 references, rejected 66 wrong answers and flagged 16 wrapped references.
- Pilot on one model, count on fresh calls. Declare the selection rule before the pilot. Our rule allowed 0 or 1 strict pass in 2 calls, including tasks never passed in the pilot.
- Say that you selected. A set picked against one model can understate that model. Opus and Haiku share Sonnet's family, so the set may be hard for them for the same reasons. Sol was not selected against.
- Decide about tools before you run. Tools on or off changes what a model tries to do.
- Keep every call. Failures, timeouts and format misses stay in the counts. Show n and the 95% interval (how Wilson intervals work).
How we measured
- Calls: Claude Sonnet 5.5 n = 16 (4 per task), Opus 5.5 n = 12, Haiku 4.5 n = 12 (3 per task), all in Claude Code at default effort. GPT-6.1 Sol n = 16 in the Codex CLI at medium effort. One call at a time per account, 300 s limit.
- Isolation: fresh empty folder, tools off, no MCP servers, one turn, no repeated counted calls. The Claude CLI reported failed internal retries on tool-call parse errors. The Codex runner adds an instruction not to call tools; the Claude Code runner adds none. Claude Code ran with an output-token setting of 16,000; the Codex CLI had none.
- Protocol: the original protocol file predates the first probe and counted call, as checked from file birth times. Later changes are amendments.
- Scoring: strict pass, format miss, wrong answer, or no answer (timeout or error). Intervals are 95% Wilson intervals, a calculation. A side is "ahead" only when the intervals do not overlap.
- Pilot: 32 Sonnet calls on 16 candidates. It chose the tasks and is excluded from counted results.
- Deviation: we repaired defective JavaScript validators instead of dropping their tasks, then re-scored stored pilot replies. No repaired code-writing task entered the counted set. The Claude batch stopped once, after a tool-call parse error on one Opus call. We then treated that error like a timeout, a scored non-pass that does not stop the batch, and resumed without repeating any call. The protocol records the change with its time.
- Prices: reported tokens times list price (Anthropic as of 2026-09-21, OpenAI as of 2026-10-03).
Caveats
- Selection effect. We chose the tasks because Sonnet's pilot missed them. A fixed set picked before any run would be fairer.
- Small samples. 12 to 16 calls per configuration, 3 to 4 per task. Repeated calls on four tasks are not independent population samples. Wilson intervals describe calls, not coding-task populations. We ran no task-level paired significance test.
- Route plus model. Each row pairs a CLI with a model, on two subscriptions. The CLIs add their own prompts and start-up time.
- Four tasks, mostly puzzles. This says little about everyday coding.
- Tools off, and one runner instruction. See above. The result is for one call with no tools, and only the Codex runner told the model not to call tools.
- Time limit. 10 of 56 calls ran past 300 s (95% interval 10% to 30%). A longer limit could change some results.
- Output cap. Reported output tokens passed 16,000 in 11 of 40 Claude Code calls (95% interval 16% to 43%), so the 16,000 setting did not bound the total. We cannot tell whether it cut any reply.
- Floor. Haiku passed none. That is a floor on this set, not a general rating.
- Disclosure. I build Agent, a product that routes work between models. The public study links the tasks, controls, protocol summary and sanitized receipts. Replies are omitted. You can check each published number from those files.
What to read next
- Claude vs Codex on hard tasks: GPT-6.1 Sol joins the hard set
- Claude Sonnet 5.5 vs Opus 5.5: when is Opus worth the price?
Run your own comparison
Agent records the model, the tokens and the result of every step. Try Agent to test your own tasks.