31 measured metrics · 19 calculated · 7 studies

Claude Opus 5.5vsGPT-6.1 Sol (Codex CLI)

No row separates them: 12 ties, 38 unclear.

The verdict

Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) share 31 measured metrics and 19 list-price calculations from 7 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 12 ties and 38 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Input tokens per call: what the CLI sends (Cache read), 4.6x (GPT-6.1 Sol (Codex CLI) larger).

Watch it build

A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.

Live story · 41 sClaude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say

Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say

50 comparison rows from 7 studies: 0 rows favour Opus 5.5, 0 favour GPT-6.1 Sol (Codex CLI), 50 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 50 rows · 7 studies. Opus 5.5 vs GPT-6.1 Sol (Codex CLI). A winner only where the 95% intervals or run ranges do not overlap.
  2. 50 comparison rows from 7 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Opus 5.5 is ahead: 0 (of 50). Rows where GPT-6.1 Sol (Codex CLI) is ahead: 0 (of 50). Ties or unclear: 50 (12 ties · 38 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. Five short tasks: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 8 rows separates them. Table: Five short tasks · 5 of 8 rows · Claude Code vs Codex CLI · n = 15 per side. Source study: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head. Rows shown: Pass rate on five validated tasks; Total time per call; Time to first useful output; Input tokens per call: what the CLI sends (Cache read); Input tokens per call: what the CLI sends (Other input). Recorded settings: Claude Code · effort high · five short validated tasks; Codex CLI · effort high · five short validated tasks. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  4. Eight hard tasks: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 6 rows separates them. Table: Eight hard tasks · 5 of 6 rows · Claude Code vs Codex CLI · n = 24 vs 16. Source study: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks. Rows shown: Pass rate on eight hard tasks (Strict pass); Pass rate on eight hard tasks (Lenient (format misses counted)); Total time per call on hard tasks (separate batches); Time to first useful output on hard tasks; Output tokens per call on hard tasks (Output tokens). Recorded settings: Claude Code · effort high · eight hard validated tasks; Codex CLI · effort high · eight hard validated tasks. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  5. Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code vs Codex CLI · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Code · six small repository tasks with hidden tests; Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  6. Effort ladder, default effort: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code vs Codex CLI · n = 16 per side. Source study: Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks. Rows shown: Strict pass rate by effort on eight hard tasks; Total time per call by effort on hard tasks; Output tokens per call by effort on hard tasks (Output tokens); List-price cost per strict pass by effort (calculation). Recorded settings: Claude Code · effort medium · eight hard validated tasks, effort ladder; Codex CLI · effort medium · eight hard validated tasks, effort ladder. Includes a calculation, not a bill or a new run. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  7. No winner where the data shows none. Showing 18 of 50 rows; every row and its reason online.

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Opus 5.5
  • GPT-6.1 Sol (Codex CLI)
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

1 tie · 7 unclear
  • Pass rate on five validated tasks: Claude Opus 5.5 100% (15/15) (n 15, 95% interval 80%–100%); GPT-6.1 Sol (Codex CLI) 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
  • Total time per call: Claude Opus 5.5 2.71 s (n 15, run range 2.5 s–11.8 s); GPT-6.1 Sol (Codex CLI) 5.60 s (n 15, run range 4.1 s–19.5 s). Unclear.
  • Time to first useful output: Claude Opus 5.5 2.04 s (n 15, run range 1.4 s–9.9 s); GPT-6.1 Sol (Codex CLI) 5.32 s (n 15, run range 3.6 s–16.4 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): Claude Opus 5.5 1,463 (n 15); GPT-6.1 Sol (Codex CLI) 6,716 (n 15). Unclear.
  • Input tokens per call: what the CLI sends (Other input): Claude Opus 5.5 619 (n 15); GPT-6.1 Sol (Codex CLI) 5,406 (n 15). Unclear.
  • Output tokens per call (Output tokens): Claude Opus 5.5 78 (n 15); GPT-6.1 Sol (Codex CLI) 42 (n 15). Unclear.
  • List-price cost per call (calculation), calculation: Claude Opus 5.5 $0.0069 (n 15, run range $0.0059–$0.027); GPT-6.1 Sol (Codex CLI) $0.010 (n 15, run range $0.0066–$0.028). Unclear.
  • List-price cost per passing answer (calculation), calculation: Claude Opus 5.5 $0.010 (n 15); GPT-6.1 Sol (Codex CLI) $0.013 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

2 ties · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call on hard tasks (separate batches): Claude Opus 5.5 11.0 s (n 24, run range 3.6 s–63 s); GPT-6.1 Sol (Codex CLI) 18.1 s (n 16, run range 11.7 s–92.2 s). Unclear.
  • Time to first useful output on hard tasks: Claude Opus 5.5 7.13 s (n 24, run range 2.2 s–56.2 s); GPT-6.1 Sol (Codex CLI) 12.7 s (n 16, run range 8.9 s–75.9 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): Claude Opus 5.5 1,052 (n 24); GPT-6.1 Sol (Codex CLI) 436 (n 16). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: Claude Opus 5.5 $0.033 (n 24); GPT-6.1 Sol (Codex CLI) $0.015 (n 16). Unclear.

Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks

1 tie · 3 unclear
  • Coding sessions that passed every hidden check: Claude Opus 5.5 100% (12/12) (n 12, 95% interval 76%–100%); GPT-6.1 Sol (Codex CLI) 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • Time per coding session: Claude Opus 5.5 56.9 s (n 12, run range 29.8 s–186 s); GPT-6.1 Sol (Codex CLI) 113.4 s (n 12, run range 78.5 s–222 s). Unclear.
  • Tool calls per coding session: Claude Opus 5.5 7.5 (n 12, run range 5–14); GPT-6.1 Sol (Codex CLI) 12.5 (n 12, run range 8–18). Unclear.
  • List-price cost per passing coding session (calculation), calculation: Claude Opus 5.5 $0.22 (n 12); GPT-6.1 Sol (Codex CLI) $0.098 (n 12). Unclear.

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks

1 tie · 3 unclear
  • Strict pass rate by effort on eight hard tasks: Claude Opus 5.5 100% (16/16) (n 16, 95% interval 81%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call by effort on hard tasks: Claude Opus 5.5 9.72 s (n 16, run range 4.8 s–31.4 s); GPT-6.1 Sol (Codex CLI) 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
  • Output tokens per call by effort on hard tasks (Output tokens): Claude Opus 5.5 853 (n 16); GPT-6.1 Sol (Codex CLI) 335 (n 16). Unclear.
  • List-price cost per strict pass by effort (calculation), calculation: Claude Opus 5.5 $0.029 (n 16); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

8 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Opus 5.5 54.4% (n 24, run range 36%–96%); GPT-6.1 Sol (Codex CLI) 57% (n 16, run range 29%–91%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Opus 5.5 $0.018 (n 24); GPT-6.1 Sol (Codex CLI) $0.0041 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Opus 5.5 $0.0080 (n 24); GPT-6.1 Sol (Codex CLI) $0.0029 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Opus 5.5 $0.0074 (n 24); GPT-6.1 Sol (Codex CLI) $0.0081 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Claude Opus 5.5 $0.013 (n 16); GPT-6.1 Sol (Codex CLI) $0.0023 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Claude Opus 5.5 $0.029 (n 16); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Opus 5.5 54.4% (n 24, run range 36%–96%); GPT-6.1 Sol (Codex CLI) 57% (n 16, run range 29%–91%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Opus 5.5 43.6% (n 15, run range 0%–93%); GPT-6.1 Sol (Codex CLI) 58.1% (n 15, run range 0%–76%). Unclear.

Where the seconds go: first text, output speed and prompt size for 6 LLMs

1 tie · 9 unclear
  • Time to first text: a 250-line answer, six models: Claude Opus 5.5 1.97 s (n 4, run range 1.7 s–2.4 s); GPT-6.1 Sol (Codex CLI) 3.52 s (n 4, run range 2.8 s–4.4 s). Unclear.
  • Output speed after the first text: visible tokens per second (calculation), calculation: Claude Opus 5.5 156 (n 4, run range 155–156); GPT-6.1 Sol (Codex CLI) 80 (n 4, run range 72–81). Unclear.
  • Output speed in characters per second after the first text (calculation), calculation: Claude Opus 5.5 347 (n 4, run range 345–349); GPT-6.1 Sol (Codex CLI) 323 (n 4, run range 291–327). Unclear.
  • Time to first text as the prompt grows: 1k, calculation: Claude Opus 5.5 1.51 s (n 3, run range 1.5 s–2 s); GPT-6.1 Sol (Codex CLI) 3.36 s (n 3, run range 3.4 s–4.8 s). Unclear.
  • Time to first text as the prompt grows: 16k, calculation: Claude Opus 5.5 1.74 s (n 3, run range 1.7 s–3 s); GPT-6.1 Sol (Codex CLI) 4.02 s (n 3, run range 3.3 s–4.3 s). Unclear.
  • Time to first text as the prompt grows: 64k, calculation: Claude Opus 5.5 1.79 s (n 3, run range 1.7 s–3.7 s); GPT-6.1 Sol (Codex CLI) 3.93 s (n 3, run range 3.4 s–4.4 s). Unclear.
  • Total time per call by prompt size (1k prompt): Claude Opus 5.5 1.83 s (n 3, run range 1.8 s–2.4 s); GPT-6.1 Sol (Codex CLI) 3.43 s (n 3, run range 3.4 s–4.9 s). Unclear.
  • Total time per call by prompt size (16k prompt): Claude Opus 5.5 2.36 s (n 3, run range 2.1 s–3.4 s); GPT-6.1 Sol (Codex CLI) 4.14 s (n 3, run range 4 s–4.7 s). Unclear.
  • Total time per call by prompt size (64k prompt): Claude Opus 5.5 2.35 s (n 3, run range 2.3 s–4.3 s); GPT-6.1 Sol (Codex CLI) 3.96 s (n 3, run range 3.5 s–4.4 s). Unclear.
  • Exact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Opus 5.5 56% (5/9) (n 9, 95% interval 27%–81%); GPT-6.1 Sol (Codex CLI) 100% (9/9) (n 9, 95% interval 70%–100%). Tie.

    Calculation: at these rates, about 13 runs per side would separate them.

GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

6 ties · 4 unclear
  • Pass rate on 4 harder tasks (Strict pass): Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%); GPT-6.1 Sol (Codex CLI) 69% (11/16) (n 16, 95% interval 44%–86%). Tie.

    Calculation: at these rates, about 49 runs per side would separate them.

  • Pass rate on 4 harder tasks (Lenient (format misses counted)): Claude Opus 5.5 50% (6/12) (n 12, 95% interval 25%–75%); GPT-6.1 Sol (Codex CLI) 69% (11/16) (n 16, 95% interval 44%–86%). Tie.

    Calculation: at these rates, about 104 runs per side would separate them.

  • Calls that tried a tool although tools were off: Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%); GPT-6.1 Sol (Codex CLI) 0% (0/16) (n 16, 95% interval 0%–19%). Unclear.
  • Strict pass rate by task: 10x10 nonogram: Claude Opus 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6.1 Sol (Codex CLI) 75% (3/4) (n 4, 95% interval 30%–95%). Tie.

    Calculation: at these rates, about 27 runs per side would separate them.

  • Strict pass rate by task: Sudoku, 22 givens: Claude Opus 5.5 0% (0/3) (n 3, 95% interval 0%–56%); GPT-6.1 Sol (Codex CLI) 25% (1/4) (n 4, 95% interval 4.6%–70%). Tie.

    Calculation: at these rates, about 26 runs per side would separate them.

  • Strict pass rate by task: 6x6 Skyscrapers: Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%); GPT-6.1 Sol (Codex CLI) 100% (4/4) (n 4, 95% interval 51%–100%). Tie.

    Calculation: at these rates, about 7 runs per side would separate them.

  • Strict pass rate by task: Seeded shuffle output: Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%); GPT-6.1 Sol (Codex CLI) 75% (3/4) (n 4, 95% interval 30%–95%). Tie.

    Calculation: at these rates, about 21 runs per side would separate them.

  • Total time per call on harder tasks: Claude Opus 5.5 80.3 s (n 9, run range 3.8 s–280 s); GPT-6.1 Sol (Codex CLI) 120.2 s (n 13, run range 46.2 s–273 s). Unclear.
  • Output tokens per call on harder tasks (Output tokens): Claude Opus 5.5 8,420 (n 9, run range 279–40,044); GPT-6.1 Sol (Codex CLI) 4,994 (n 13, run range 2,099–13,413). Unclear.
  • List-price cost per strict pass on harder tasks (calculation), calculation: Claude Opus 5.5 $0.59 (n 12); GPT-6.1 Sol (Codex CLI) $0.083 (n 16). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

50 rows from 7 studies. No row separates them: 12 ties, 38 unclear.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Opus 5.5

  • List-price cost per call (calculation): $0.0069 vs $0.010. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 1k: 1.51 s vs 3.36 s. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 16k: 1.74 s vs 4.02 s. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 64k: 1.79 s vs 3.93 s. A list-price calculation, not a measured difference. Calculation

When to pick GPT-6.1 Sol (Codex CLI)

  • List-price cost per strict pass on hard tasks (calculation): $0.015 vs $0.033. A list-price calculation, not a measured difference. Calculation
  • List-price cost per passing coding session (calculation): $0.098 vs $0.22. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0041 vs $0.018. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0029 vs $0.0080. A list-price calculation, not a measured difference. Calculation
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0023 vs $0.013. A list-price calculation, not a measured difference. Calculation
  • List-price cost per strict pass on harder tasks (calculation): $0.083 vs $0.59. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only these two. Open a study for every configuration.

Every rate is 95% or more
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (high) · Codex CLI

2 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 15 per row

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 4× real timeMotion reduced: press Replay to animateThe slowest median is 5.6 s. The clock runs at the recorded speed.
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (high) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 5.6 s (range 4.1 s–19.5 s, n 15). Fastest Claude Opus 5.5 (high) · Claude Code 2.7 s (range 2.5 s–11.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 3.8× real timeMotion reduced: press Replay to animateThe slowest median is 5.3 s. The clock runs at the recorded speed.
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (high) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 5.3 s (range 3.6 s–16.4 s, n 15). Fastest Claude Opus 5.5 (high) · Claude Code 2 s (range 1.4 s–9.9 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (high) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

2 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (high) · Codex CLI 6,716 (n 15). Lowest Claude Opus 5.5 (high) · Claude Code 1,463 (n 15). Other input: highest GPT-6.1 Sol (high) · Codex CLI 5,406 (n 15). Lowest Claude Opus 5.5 (high) · Claude Code 619 (n 15).

Notesn = 15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Opus 5.5 or GPT-6.1 Sol (Codex CLI)?
Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) share 31 measured metrics and 19 list-price calculations from 7 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 12 ties and 38 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
How were Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) measured?
They share 31 measured metrics and 19 list-price calculations from 7 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on pass rate on five validated tasks?
Claude Opus 5.5: 100% (15/15) (Claude Code · effort high · five short validated tasks; n = 15; 95% interval 80% to 100%). GPT-6.1 Sol (Codex CLI): 100% (15/15) (Codex CLI · effort high · five short validated tasks; n = 15; 95% interval 80% to 100%). The 95% intervals overlap (Claude Opus 5.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.
How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on total time per call?
Claude Opus 5.5: 2.71 s (Claude Code · effort high · five short validated tasks; n = 15; run range 2.5 s to 11.8 s). GPT-6.1 Sol (Codex CLI): 5.60 s (Codex CLI · effort high · five short validated tasks; n = 15; run range 4.1 s to 19.5 s). The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.45 s to 11.8 s; GPT-6.1 Sol (Codex CLI) 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on time to first useful output?
Claude Opus 5.5: 2.04 s (Claude Code · effort high · five short validated tasks; n = 15; run range 1.4 s to 9.9 s). GPT-6.1 Sol (Codex CLI): 5.32 s (Codex CLI · effort high · five short validated tasks; n = 15; run range 3.6 s to 16.4 s). The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.40 s to 9.94 s; GPT-6.1 Sol (Codex CLI) 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on pass rate on eight hard tasks (Strict pass)?
Claude Opus 5.5: 100% (24/24) (Claude Code · effort high · eight hard validated tasks; n = 24; 95% interval 86% to 100%). GPT-6.1 Sol (Codex CLI): 100% (16/16) (Codex CLI · effort high · eight hard validated tasks; n = 16; 95% interval 81% to 100%). The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.
How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on pass rate on eight hard tasks (Lenient (format misses counted))?
Claude Opus 5.5: 100% (24/24) (Claude Code · effort high · eight hard validated tasks; n = 24; 95% interval 86% to 100%). GPT-6.1 Sol (Codex CLI): 100% (16/16) (Codex CLI · effort high · eight hard validated tasks; n = 16; 95% interval 81% to 100%). The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.