GPT-6.1 Sol vs Claude Sonnet 5.5 vs Opus 5.5: every row we measured
Of 70 comparison rows for GPT-6.1 Sol, Claude Sonnet 5.5 and Opus 5.5, only 2 have a winner (speed). All 15 pass-rate rows tie. Tokens, price and route differ.
TL;DR
- Quality ties on our tasks. Sonnet 5.5 and Opus 5.5 each passed 24 of 24 hard calls (95% interval 86% to 100%). GPT-6.1 Sol passed 16 of 16 at medium and at high effort (81% to 100%). All 15 pass-rate rows tie. The hard set is at its ceiling.
- Only 2 of 70 rows have a winner. Both measure speed. Sonnet was ahead of Sol on two prompts, each run 10 times: 6.89 s vs 13.4 s and 2.67 s vs 11.3 s per call.
- Route changes the result. Every Sol row ran in the Codex CLI. Sol repaired a scheduler in a median 17.3 s through the API and 61.2 s through the Codex CLI (3 runs each).
- Tokens: Sol's median output was the lowest, 335 per hard call at medium effort against 1,050 (Sonnet) and 945 (Opus). On the easy set, the Codex CLI sent about 5.8 times Sonnet's input per call (calculation). No token row has a winner.
- Price (calculation): Sol and Sonnet list at the same input and output price; Opus at twice that. Per hard pass: Sonnet $0.0143, Sol (high) $0.0151, Sol (medium) $0.0256, Opus $0.0282. Opus figures are provisional: its cache-read price is under re-check.
Every row: Sonnet vs Sol, Opus vs Sol and Sonnet vs Opus. Model page: GPT-6.1 Sol (Codex CLI).
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks
139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.
Transcript
- Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
- 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
- Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
- Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
- Open benchmarks: intervals, sources and every failure kept.
The three models side by side
Sonnet and Opus ran in Claude Code at default effort. Sol ran in the Codex CLI at medium and high.
| Measure | Sonnet 5.5 | Opus 5.5 | GPT-6.1 Sol |
|---|---|---|---|
| Hard set, strict pass | 24/24 (86% to 100%) | 24/24 (86% to 100%) | 16/16 at medium and high (81% to 100%) |
| Easy set, pass | 12/15 (55% to 93%) | 15/15 (80% to 100%) | 15/15 at medium and high (80% to 100%) |
| Median time per hard call | 7.75 s | 9.18 s | 13.11 s (medium), 18.12 s (high) |
| Median output tokens, hard call | 1,050 | 945 | 335 (medium), 436 (high) |
| Cost per hard pass (calculation) | $0.0143 | $0.0282 (provisional) | $0.0256 (medium), $0.0151 (high) |
How the 70 rows came out
The three comparison pages hold 70 rows, and one measurement can appear on more than one page. A row has a winner only when the 95% intervals or run ranges do not overlap. Cost rows carry no interval. The Sonnet vs Sol page holds both winners.
| Dimension | Rows | Ties | Winner | Unclear |
|---|---|---|---|---|
| Pass rate | 15 | 15 | 0 | 0 |
| Answer variety (same prompt, 10 times) | 3 | 2 | 0 | 1 |
| Time | 22 | 0 | 2 | 20 |
| Tokens | 16 | 1 | 0 | 15 |
| List-price cost (calculation) | 14 | 0 | 0 | 14 |
| All rows | 70 | 18 | 2 | 50 |
Quality: all three at the ceiling
- Effort ladder: all 11 cells passed 16/16 (176 of 176), across Sonnet, Opus and Sol.
- Same prompt, 10 times: Sonnet and Sol (medium) each passed 10/10 on all three prompts (72% to 100%). Opus did not run here.
- Easy set: Sonnet's 3 misses were format misses: the right number plus extra working lines, which the exact-text validator rejects.
- A tie is not equality. At the ceiling, our tasks cannot tell the models apart.
Speed: Sonnet's median is lowest, and it is ahead on two rows
The single-call ranges overlap: Sonnet 2.26 s to 34.79 s, Opus 4.24 s to 27.21 s, Sol (medium) 8.54 s to 61.6 s.
A range is not a confidence interval, so the medians describe this run and do not rank the models. In all 22 time rows on the three pages, the medians run in one order: Sonnet lowest, then Opus, then Sol.
We sent each of three prompts 10 times. For two, Sonnet's runs and Sol's (medium) runs did not overlap:
- Exact number: Sonnet 6.89 s (5.81 s to 7.81 s), Sol 13.4 s (12.3 s to 18.0 s).
- Code fix: Sonnet 2.67 s (2.32 s to 4.34 s), Sol 11.3 s (9.08 s to 14.8 s).
- JSON object: Sonnet 2.89 s, Sol 6.42 s. The ranges overlap, so no winner.
The route changes the result
All 9 runs of a scheduler repair passed 296 checks. Medians at medium effort, 3 runs each: Sonnet in Claude Code 15.0 s, Sol through the OpenAI API 17.3 s, Sol through the Codex CLI 61.2 s. The Sonnet and Codex CLI ranges do not overlap: 13.9 s to 15.9 s against 59.9 s to 69.5 s. With 3 runs, the comparison page still calls the row unclear.
For a one-line answer (5 runs), Sol's median was 1.02 s (low) and 1.52 s (high) through the API. Through the Codex CLI it was 4.18 s and 4.19 s. The ranges do not overlap at either effort. See Sol CLI vs API.
Part of Sol's gap is the route, and we cannot say how much. These studies ran no Claude model through an API, so no row compares the three over one route.
Tokens: Sol writes fewer output tokens, the Codex CLI sends more input
Median output per hard call: Sol 335 (medium) and 436 (high), Opus 945, Sonnet 1,050. The comparison pages mark these rows as behaviour, not a winner, because fewer tokens is not better or worse by itself.
On the easy set, Sol at medium sent a mean 12,123 input tokens per call (5,180 cache reads, 6,943 other). Sonnet sent 2,086 and Opus 2,081. Sol's figure is about 5.8 times Sonnet's (calculation). The study finds most of it is the CLI's own system prompt and tool context.
Price: same rate for Sol and Sonnet, double for Opus (calculation)
List price per million tokens:
| Model | Input | Cache read | Output |
|---|---|---|---|
| GPT-6.1 Sol | $2 | $0.10 | $10 |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 |
| Claude Opus 5.5 | $4 | $0.20 (under re-check) | $20 |
Cost per hard pass, all calls counted (calculation, no interval): Sonnet $0.01435, Sol (high) $0.01514, Sol (medium) $0.02564, Opus $0.02824 (provisional).
Sol (medium) cost more per pass than Sol (high), although its median output was lower (335 against 436 tokens). We did not test the cause.
Our agent used 162.9M input tokens (94.0% cache reads) and 1.8M output tokens in 33 SWE-bench runs, and resolved 25 instances. We priced these Sonnet tokens at each list price. Per resolved instance (calculation): Sol $2.10, Sonnet $3.49, Opus $5.75 (provisional).
This is not a run: Sol or Opus would use other tokens. Sol's lower figure comes from cache reads, which cost half as much, and from cache writes. OpenAI writes list at plain input; Anthropic one-hour writes list at twice input.
Which to pick for what
Our reading, not a ranking:
- You need the best answers. We cannot pick: no pass-rate row has a winner. Run your own tasks.
- You need speed. Start with Sonnet in Claude Code: its median is lowest in all 17 time rows where it appears. For Sol, also test the API route.
- You pay for many output tokens. Sol's median output is lowest, and so is its repriced figure (calculation). Check input tokens first: the Codex CLI sends more.
- You consider Opus. No row shows what its price buys. Read when Opus is worth it.
- You want the lowest cost per pass. Sonnet at default: $0.0143 (hard) and $0.0062 (easy) per pass (calculation). Sol at high effort is 1.06 times Sonnet on the hard set (calculation).
How we measured
- Sets: hard, 8 tasks with sandboxed validators (Claude Code n = 24, Codex CLI n = 16); easy, 5 tasks (n = 15). Each call: fresh folder, tools off, one turn, no retries.
- Studies: hard, easy, effort ladder, consistency, CLI and API latency and cost.
- Costs: reported tokens times list price. The calls ran on subscriptions. Prices as of 2026-09-21 (Anthropic) and 2026-10-03 (OpenAI).
Caveats
- Ceiling. Six of the seven hard-study configurations passed every call, so pass rate cannot separate them.
- Route plus model. The Claude and Codex batches ran on different days on one host. Effort is not one scale across vendors.
- Small samples. Scheduler rows have 3 runs, one-line rows 5.
- Blocked tries. 30 earlier Codex CLI tries stopped before any model call (no signed-in account). We do not score them.
- Opus price. We are re-checking the Opus 5.5 cache-read price ($0.20 per million here), so Opus cost figures stay provisional.
What to read next
- Claude vs Codex on hard tasks: GPT-6.1 Sol joins the hard set
- Claude Sonnet 5.5 vs Opus 5.5: when is Opus worth the price?
Compare the three on your own work
Agent records the model, the tokens and the result of every step. Try Agent to run your own comparison.