• Head to head
  • GPT-6.1 Sol
  • Claude Sonnet
  • Claude Opus

GPT-6.1 Sol vs Claude Sonnet 5.5 vs Opus 5.5: every row we measured

Of 70 comparison rows for GPT-6.1 Sol, Claude Sonnet 5.5 and Opus 5.5, only 2 have a winner (speed). All 15 pass-rate rows tie. Tokens, price and route differ.

TL;DR

  • Quality ties on our tasks. Sonnet 5.5 and Opus 5.5 each passed 24 of 24 hard calls (95% interval 86% to 100%). GPT-6.1 Sol passed 16 of 16 at medium and at high effort (81% to 100%). All 15 pass-rate rows tie. The hard set is at its ceiling.
  • Only 2 of 70 rows have a winner. Both measure speed. Sonnet was ahead of Sol on two prompts, each run 10 times: 6.89 s vs 13.4 s and 2.67 s vs 11.3 s per call.
  • Route changes the result. Every Sol row ran in the Codex CLI. Sol repaired a scheduler in a median 17.3 s through the API and 61.2 s through the Codex CLI (3 runs each).
  • Tokens: Sol's median output was the lowest, 335 per hard call at medium effort against 1,050 (Sonnet) and 945 (Opus). On the easy set, the Codex CLI sent about 5.8 times Sonnet's input per call (calculation). No token row has a winner.
  • Price (calculation): Sol and Sonnet list at the same input and output price; Opus at twice that. Per hard pass: Sonnet $0.0143, Sol (high) $0.0151, Sol (medium) $0.0256, Opus $0.0282. Opus figures are provisional: its cache-read price is under re-check.

Every row: Sonnet vs Sol, Opus vs Sol and Sonnet vs Opus. Model page: GPT-6.1 Sol (Codex CLI).

Live story · 44 sHaiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.

Transcript
  1. Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
  2. 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  3. Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  4. Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  5. Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  6. Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  7. Open benchmarks: intervals, sources and every failure kept.

The three models side by side

Sonnet and Opus ran in Claude Code at default effort. Sol ran in the Codex CLI at medium and high.

MeasureSonnet 5.5Opus 5.5GPT-6.1 Sol
Hard set, strict pass24/24 (86% to 100%)24/24 (86% to 100%)16/16 at medium and high (81% to 100%)
Easy set, pass12/15 (55% to 93%)15/15 (80% to 100%)15/15 at medium and high (80% to 100%)
Median time per hard call7.75 s9.18 s13.11 s (medium), 18.12 s (high)
Median output tokens, hard call1,050945335 (medium), 436 (high)
Cost per hard pass (calculation)$0.0143$0.0282 (provisional)$0.0256 (medium), $0.0151 (high)

How the 70 rows came out

The three comparison pages hold 70 rows, and one measurement can appear on more than one page. A row has a winner only when the 95% intervals or run ranges do not overlap. Cost rows carry no interval. The Sonnet vs Sol page holds both winners.

DimensionRowsTiesWinnerUnclear
Pass rate151500
Answer variety (same prompt, 10 times)3201
Time220220
Tokens161015
List-price cost (calculation)140014
All rows7018250

Quality: all three at the ceiling

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

  • Effort ladder: all 11 cells passed 16/16 (176 of 176), across Sonnet, Opus and Sol.
  • Same prompt, 10 times: Sonnet and Sol (medium) each passed 10/10 on all three prompts (72% to 100%). Opus did not run here.
  • Easy set: Sonnet's 3 misses were format misses: the right number plus extra working lines, which the exact-text validator rejects.
  • A tie is not equality. At the ceiling, our tasks cannot tell the models apart.

Speed: Sonnet's median is lowest, and it is ahead on two rows

Entrance: medians race at 28× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

7 rows. Slowest Claude Haiku 4.5 · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 16–24 per row

Median per configuration; whiskers = fastest and slowest call

One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

The single-call ranges overlap: Sonnet 2.26 s to 34.79 s, Opus 4.24 s to 27.21 s, Sol (medium) 8.54 s to 61.6 s.

A range is not a confidence interval, so the medians describe this run and do not rank the models. In all 22 time rows on the three pages, the medians run in one order: Sonnet lowest, then Opus, then Sol.

  • Exact number
  • JSON object
  • Code fix
Entrance: medians race at 9.6× real timeMotion reduced: press Replay to animateThe slowest median is 13.4 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: slowest GPT-6.1 Sol (medium) · Codex CLI 13.4 s (range 12.3 s–18 s, n 10). Fastest Claude Haiku 4.5 · Claude Code 5.1 s (range 4.4 s–6.2 s, n 10). Not all run ranges overlap. JSON object: slowest Claude Haiku 4.5 · Claude Code 7 s (range 5.3 s–12.3 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 2.9 s (range 2.7 s–5.3 s, n 10). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 10 per row

Median; whiskers = fastest and slowest of 10 calls

Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

We sent each of three prompts 10 times. For two, Sonnet's runs and Sol's (medium) runs did not overlap:

  • Exact number: Sonnet 6.89 s (5.81 s to 7.81 s), Sol 13.4 s (12.3 s to 18.0 s).
  • Code fix: Sonnet 2.67 s (2.32 s to 4.34 s), Sol 11.3 s (9.08 s to 14.8 s).
  • JSON object: Sonnet 2.89 s, Sol 6.42 s. The ranges overlap, so no winner.

The route changes the result

  • Total time
  • First useful output
Entrance: medians race at 44× real time
Claude Code CLI · Sonnet 5.5 · medium
Codex CLI · GPT-6.1 Sol · medium
OpenAI API · GPT-6.1 Sol · medium

3 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · medium 61.2 s (range 59.9 s–69.5 s, n 3). Fastest Claude Code CLI · Sonnet 5.5 · medium 15 s (range 13.9 s–15.9 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · medium 15.6 s (range 13.7 s–23 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · medium 7.5 s (range 6.7 s–9.1 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Same prompt, medium effort, 296 behavioral checks, 3 runs each

All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.

Source: Provider explorer receipts: CLI vs API

All 9 runs of a scheduler repair passed 296 checks. Medians at medium effort, 3 runs each: Sonnet in Claude Code 15.0 s, Sol through the OpenAI API 17.3 s, Sol through the Codex CLI 61.2 s. The Sonnet and Codex CLI ranges do not overlap: 13.9 s to 15.9 s against 59.9 s to 69.5 s. With 3 runs, the comparison page still calls the row unclear.

For a one-line answer (5 runs), Sol's median was 1.02 s (low) and 1.52 s (high) through the API. Through the Codex CLI it was 4.18 s and 4.19 s. The ranges do not overlap at either effort. See Sol CLI vs API.

Part of Sol's gap is the route, and we cannot say how much. These studies ran no Claude model through an API, so no row compares the three over one route.

Tokens: Sol writes fewer output tokens, the Codex CLI sends more input

Median output per hard call: Sol 335 (medium) and 436 (high), Opus 945, Sonnet 1,050. The comparison pages mark these rows as behaviour, not a winner, because fewer tokens is not better or worse by itself.

On the easy set, Sol at medium sent a mean 12,123 input tokens per call (5,180 cache reads, 6,943 other). Sonnet sent 2,086 and Opus 2,081. Sol's figure is about 5.8 times Sonnet's (calculation). The study finds most of it is the CLI's own system prompt and tool context.

Price: same rate for Sol and Sonnet, double for Opus (calculation)

List price per million tokens:

ModelInputCache readOutput
GPT-6.1 Sol$2$0.10$10
Claude Sonnet 5.5$2$0.20$10
Claude Opus 5.5$4$0.20 (under re-check)$20
Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Cost per hard pass, all calls counted (calculation, no interval): Sonnet $0.01435, Sol (high) $0.01514, Sol (medium) $0.02564, Opus $0.02824 (provisional).

Sol (medium) cost more per pass than Sol (high), although its median output was lower (335 against 436 tokens). We did not test the cause.

Our agent used 162.9M input tokens (94.0% cache reads) and 1.8M output tokens in 33 SWE-bench runs, and resolved 25 instances. We priced these Sonnet tokens at each list price. Per resolved instance (calculation): Sol $2.10, Sonnet $3.49, Opus $5.75 (provisional).

This is not a run: Sol or Opus would use other tokens. Sol's lower figure comes from cache reads, which cost half as much, and from cache writes. OpenAI writes list at plain input; Anthropic one-hour writes list at twice input.

Which to pick for what

Our reading, not a ranking:

  • You need the best answers. We cannot pick: no pass-rate row has a winner. Run your own tasks.
  • You need speed. Start with Sonnet in Claude Code: its median is lowest in all 17 time rows where it appears. For Sol, also test the API route.
  • You pay for many output tokens. Sol's median output is lowest, and so is its repriced figure (calculation). Check input tokens first: the Codex CLI sends more.
  • You consider Opus. No row shows what its price buys. Read when Opus is worth it.
  • You want the lowest cost per pass. Sonnet at default: $0.0143 (hard) and $0.0062 (easy) per pass (calculation). Sol at high effort is 1.06 times Sonnet on the hard set (calculation).

How we measured

  • Sets: hard, 8 tasks with sandboxed validators (Claude Code n = 24, Codex CLI n = 16); easy, 5 tasks (n = 15). Each call: fresh folder, tools off, one turn, no retries.
  • Studies: hard, easy, effort ladder, consistency, CLI and API latency and cost.
  • Costs: reported tokens times list price. The calls ran on subscriptions. Prices as of 2026-09-21 (Anthropic) and 2026-10-03 (OpenAI).

Caveats

  • Ceiling. Six of the seven hard-study configurations passed every call, so pass rate cannot separate them.
  • Route plus model. The Claude and Codex batches ran on different days on one host. Effort is not one scale across vendors.
  • Small samples. Scheduler rows have 3 runs, one-line rows 5.
  • Blocked tries. 30 earlier Codex CLI tries stopped before any model call (no signed-in account). We do not score them.
  • Opus price. We are re-checking the Opus 5.5 cache-read price ($0.20 per million here), so Opus cost figures stay provisional.

Compare the three on your own work

Agent records the model, the tokens and the result of every step. Try Agent to run your own comparison.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.