14 measured metrics · 3 calculated · 2 studies

Claude Sonnet 5.5vsGPT-6 Luna (Codex CLI)

Claude Sonnet 5.5 ahead on 1; 9 ties, 7 unclear. A side is ahead only where the intervals or ranges do not overlap.

The verdict

Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. Claude Sonnet 5.5 leads on 1 row: Strict pass rate: single call vs agent loop on eight hard tasks, 100% (24/24) vs 63% (10/16). On those rows the 95% intervals do not overlap. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

5 headline metrics as the ratio of the two values. 1 of them separate the sides in the data. Widest ratio: List-price cost per strict pass: single call vs agent loop (calculation), 12x (Claude Sonnet 5.5 larger).

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Sonnet 5.5
  • GPT-6 Luna (Codex CLI)
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks

Claude Sonnet 5.5 1 · 9 ties · 4 unclear
  • Strict pass rate: single call vs agent loop on eight hard tasks: Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6 Luna (Codex CLI) 63% (10/16) (n 16, 95% interval 39%–82%). Claude Sonnet 5.5 ahead.
  • Strict passes per task: single call vs agent loop: Interval merge fix: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
  • Strict passes per task: single call vs agent loop: DST day-length fix: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
  • Strict passes per task: single call vs agent loop: CSV parser: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
  • Strict passes per task: single call vs agent loop: Event-loop order: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 0% (0/2) (n 2, 95% interval 0%–66%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: Room schedule: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.

    Calculation: at these rates, about 12 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: SemVer regex: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.

    Calculation: at these rates, about 12 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: Money refactor: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 0% (0/2) (n 2, 95% interval 0%–66%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: SQL report: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
  • Total time per attempt: single call vs agent loop: Claude Sonnet 5.5 7.75 s (n 24, run range 2.3 s–34.8 s); GPT-6 Luna (Codex CLI) 5.16 s (n 16, run range 3.6 s–11.3 s). Unclear.
  • Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)): Claude Sonnet 5.5 2,281 (n 24, run range 2,234–2,669); GPT-6 Luna (Codex CLI) 11,582 (n 16, run range 11,526–11,818). Unclear.
  • Tokens per attempt: single call vs agent loop (Output tokens): Claude Sonnet 5.5 1,050 (n 24, run range 176–3,895); GPT-6 Luna (Codex CLI) 345 (n 16, run range 36–634). Unclear.
  • Tool calls per agent-loop attempt: Claude Sonnet 5.5 0 (n 16, run range 0–3); GPT-6 Luna (Codex CLI) 0 (n 14, run range 0–1). Tie.
  • List-price cost per strict pass: single call vs agent loop (calculation), calculation: Claude Sonnet 5.5 $0.014 (n 24); GPT-6 Luna (Codex CLI) $0.0012 (n 16). Unclear.

Where the seconds go: first text, output speed and prompt size for 6 LLMs

3 unclear
  • Time to first text: a 250-line answer, six models: Claude Sonnet 5.5 1.96 s (n 4, run range 0.9 s–4.1 s); GPT-6 Luna (Codex CLI) 3.30 s (n 4, run range 3.2 s–3.5 s). Unclear.
  • Output speed after the first text: visible tokens per second (calculation), calculation: Claude Sonnet 5.5 232 (n 4, run range 230–233); GPT-6 Luna (Codex CLI) 129 (n 4, run range 56–259). Unclear.
  • Output speed in characters per second after the first text (calculation), calculation: Claude Sonnet 5.5 517 (n 4, run range 513–519); GPT-6 Luna (Codex CLI) 524 (n 4, run range 225–1,052). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

17 rows from 2 studies. Claude Sonnet 5.5 ahead on 1; 9 ties, 7 unclear. A side is ahead only where the intervals or ranges do not overlap.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Sonnet 5.5

  • Strict pass rate: single call vs agent loop on eight hard tasks: 100% (24/24) vs 63% (10/16). The 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%).

When to pick GPT-6 Luna (Codex CLI)

  • List-price cost per strict pass: single call vs agent loop (calculation): $0.0012 vs $0.014. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only these two. Open a study for every configuration.

Claude Sonnet 5.5 (single call) · Claude Code
GPT-6 Luna (single call) · Codex CLI

2 rows. Highest Claude Sonnet 5.5 (single call) · Claude Code 100% (95% interval 86%–100%, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 63% (95% interval 39%–82%, n 16). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row

Same tasks and validators. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Claude Haiku 4.5 (single call) · Claude Code

Interval merge fix
Event-loop order
Room schedule

Claude Haiku 4.5 (agent loop) · Claude Code

Interval merge fix
Event-loop order
Room schedule

Claude Sonnet 5.5 (single call) · Claude Code

Interval merge fix
Event-loop order
Room schedule

Claude Sonnet 5.5 (agent loop) · Claude Code

Interval merge fix
Event-loop order
Room schedule

GPT-6 Luna (single call) · Codex CLI

Interval merge fix
Event-loop order
Room schedule

GPT-6 Luna (agent loop) · Codex CLI

Interval merge fix
Event-loop order
Room schedule

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

3 rows, 6 series: Claude Haiku 4.5 (single call) · Claude Code, Claude Haiku 4.5 (agent loop) · Claude Code, Claude Sonnet 5.5 (single call) · Claude Code, Claude Sonnet 5.5 (agent loop) · Claude Code, GPT-6 Luna (single call) · Codex CLI, GPT-6 Luna (agent loop) · Codex CLI. Claude Haiku 4.5 (single call) · Claude Code: highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest Room schedule 0% (95% interval 0%–56%, n 3). All intervals overlap. Claude Haiku 4.5 (agent loop) · Claude Code: highest Event-loop order 100% (95% interval 44%–100%, n 3). Lowest Room schedule 67% (95% interval 21%–94%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 2–3 per row

Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration

Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Entrance: medians race at 5.5× real timeMotion reduced: press Replay to animateThe slowest median is 7.8 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (single call) · Claude Code
GPT-6 Luna (single call) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest Claude Sonnet 5.5 (single call) · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). Fastest GPT-6 Luna (single call) · Codex CLI 5.2 s (range 3.6 s–11.3 s, n 16). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 16–24 per row

Median per configuration; whiskers = fastest and slowest attempt

Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

  • Input tokens (cache reads included)
  • Output tokens
Claude Sonnet 5.5 (single call) · Claude Code
GPT-6 Luna (single call) · Codex CLI

2 rows, 2 series: Input tokens (cache reads included), Output tokens. Input tokens (cache reads included): highest GPT-6 Luna (single call) · Codex CLI 11,582 (range 11,526–11,818, n 16). Lowest Claude Sonnet 5.5 (single call) · Claude Code 2,281 (range 2,234–2,669, n 24). Not all run ranges overlap. Output tokens: highest Claude Sonnet 5.5 (single call) · Claude Code 1,050 (range 176–3,895, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 345 (range 36–634, n 16). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 16–24 per row

Median per configuration; whiskers = fewest and most

Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Sonnet 5.5 or GPT-6 Luna (Codex CLI)?
Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. Claude Sonnet 5.5 leads on 1 row: Strict pass rate: single call vs agent loop on eight hard tasks, 100% (24/24) vs 63% (10/16). On those rows the 95% intervals do not overlap. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).
How were Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) measured?
They share 14 measured metrics and 3 list-price calculations from 2 public studies: Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks; Where the seconds go: first text, output speed and prompt size for 6 LLMs. Every row names its configuration, its sample size and its interval or range.
How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict pass rate: single call vs agent loop on eight hard tasks?
Claude Sonnet 5.5: 100% (24/24) (Claude Code · single call; n = 24; 95% interval 86% to 100%). GPT-6 Luna (Codex CLI): 63% (10/16) (Codex CLI · single call; n = 16; 95% interval 39% to 82%). The 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%).
How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: Interval merge fix?
Claude Sonnet 5.5: 100% (3/3) (Claude Code · single call; n = 3; 95% interval 44% to 100%). GPT-6 Luna (Codex CLI): 100% (2/2) (Codex CLI · single call; n = 2; 95% interval 34% to 100%). The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.
How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: DST day-length fix?
Claude Sonnet 5.5: 100% (3/3) (Claude Code · single call; n = 3; 95% interval 44% to 100%). GPT-6 Luna (Codex CLI): 100% (2/2) (Codex CLI · single call; n = 2; 95% interval 34% to 100%). The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.
How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: CSV parser?
Claude Sonnet 5.5: 100% (3/3) (Claude Code · single call; n = 3; 95% interval 44% to 100%). GPT-6 Luna (Codex CLI): 100% (2/2) (Codex CLI · single call; n = 2; 95% interval 34% to 100%). The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.
How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: Event-loop order?
Claude Sonnet 5.5: 100% (3/3) (Claude Code · single call; n = 3; 95% interval 44% to 100%). GPT-6 Luna (Codex CLI): 0% (0/2) (Codex CLI · single call; n = 2; 95% interval 0% to 66%). The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.

The studies behind this page

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.