14 measured metrics · 3 calculated · 2 studies
Claude Sonnet 5.5vsGPT-6 Luna (Codex CLI)
Claude Sonnet 5.5 ahead on 1; 9 ties, 7 unclear. A side is ahead only where the intervals or ranges do not overlap.
The verdict
Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. Claude Sonnet 5.5 leads on 1 row: Strict pass rate: single call vs agent loop on eight hard tasks, 100% (24/24) vs 63% (10/16). On those rows the 95% intervals do not overlap. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Claude Sonnet 5.5 | GPT-6 Luna (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Strict pass rate: single call vs agent loop on eight hard tasks | 100% (24/24)Claude Code · single call | 63% (10/16)Codex CLI · single call | 24 / 16 | 95% CI: 86%–100% vs 39%–82% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%). | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Total time per attempt: single call vs agent loop | 7.75 sClaude Code · single call | 5.16 sCodex CLI · single call | 24 / 16 | range: 2.3 s–34.8 s vs 3.6 s–11.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)) | 2,281Claude Code · single call | 11,582Codex CLI · single call | 24 / 16 | range: 2,234–2,669 vs 11,526–11,818 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Output tokens) | 1,050Claude Code · single call | 345Codex CLI · single call | 24 / 16 | range: 176–3,895 vs 36–634 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| List-price cost per strict pass: single call vs agent loop (calculation)Calculation | $0.014Claude Code · single call | $0.0012Codex CLI · single call | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
5 headline metrics as the ratio of the two values. 1 of them separate the sides in the data. Widest ratio: List-price cost per strict pass: single call vs agent loop (calculation), 12x (Claude Sonnet 5.5 larger).
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Claude Sonnet 5.5
- GPT-6 Luna (Codex CLI)
- 95% interval
- fastest–slowest run (not an interval)
- where the two overlap
- hollow: list-price calculation
These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
- Strict pass rate: single call vs agent loop on eight hard tasks100% (24/24)n 2463% (10/16)n 16Claude Sonnet 5.5 aheadStrict pass rate: single call vs agent loop on eight hard tasks: Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6 Luna (Codex CLI) 63% (10/16) (n 16, 95% interval 39%–82%). Claude Sonnet 5.5 ahead.
- Strict passes per task: single call vs agent loop: Interval merge fix100% (3/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: Interval merge fix: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
- Strict passes per task: single call vs agent loop: DST day-length fix100% (3/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: DST day-length fix: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
- Strict passes per task: single call vs agent loop: CSV parser100% (3/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: CSV parser: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
- Strict passes per task: single call vs agent loop: Event-loop order100% (3/3)n 30% (0/2)n 2TieStrict passes per task: single call vs agent loop: Event-loop order: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 0% (0/2) (n 2, 95% interval 0%–66%). Tie.
Calculation: at these rates, about 4 runs per side would separate them.
- Strict passes per task: single call vs agent loop: Room schedule100% (3/3)n 350% (1/2)n 2TieStrict passes per task: single call vs agent loop: Room schedule: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.
Calculation: at these rates, about 12 runs per side would separate them.
- Strict passes per task: single call vs agent loop: SemVer regex100% (3/3)n 350% (1/2)n 2TieStrict passes per task: single call vs agent loop: SemVer regex: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.
Calculation: at these rates, about 12 runs per side would separate them.
- Strict passes per task: single call vs agent loop: Money refactor100% (3/3)n 30% (0/2)n 2TieStrict passes per task: single call vs agent loop: Money refactor: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 0% (0/2) (n 2, 95% interval 0%–66%). Tie.
Calculation: at these rates, about 4 runs per side would separate them.
- Strict passes per task: single call vs agent loop: SQL report100% (3/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: SQL report: Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
- Total time per attempt: single call vs agent loop7.75 sn 245.16 sn 16UnclearTotal time per attempt: single call vs agent loop: Claude Sonnet 5.5 7.75 s (n 24, run range 2.3 s–34.8 s); GPT-6 Luna (Codex CLI) 5.16 s (n 16, run range 3.6 s–11.3 s). Unclear.
- Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))2,281n 2411,582n 16UnclearTokens per attempt: single call vs agent loop (Input tokens (cache reads included)): Claude Sonnet 5.5 2,281 (n 24, run range 2,234–2,669); GPT-6 Luna (Codex CLI) 11,582 (n 16, run range 11,526–11,818). Unclear.
- Tokens per attempt: single call vs agent loop (Output tokens)1,050n 24345n 16UnclearTokens per attempt: single call vs agent loop (Output tokens): Claude Sonnet 5.5 1,050 (n 24, run range 176–3,895); GPT-6 Luna (Codex CLI) 345 (n 16, run range 36–634). Unclear.
- Tool calls per agent-loop attempt0n 160n 14TieTool calls per agent-loop attempt: Claude Sonnet 5.5 0 (n 16, run range 0–3); GPT-6 Luna (Codex CLI) 0 (n 14, run range 0–1). Tie.
- List-price cost per strict pass: single call vs agent loop (calculation)Calculation$0.014n 24$0.0012n 16UnclearList-price cost per strict pass: single call vs agent loop (calculation), calculation: Claude Sonnet 5.5 $0.014 (n 24); GPT-6 Luna (Codex CLI) $0.0012 (n 16). Unclear.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
- Time to first text: a 250-line answer, six models1.96 sn 43.30 sn 4UnclearTime to first text: a 250-line answer, six models: Claude Sonnet 5.5 1.96 s (n 4, run range 0.9 s–4.1 s); GPT-6 Luna (Codex CLI) 3.30 s (n 4, run range 3.2 s–3.5 s). Unclear.
- Output speed after the first text: visible tokens per second (calculation)Calculation232n 4129n 4UnclearOutput speed after the first text: visible tokens per second (calculation), calculation: Claude Sonnet 5.5 232 (n 4, run range 230–233); GPT-6 Luna (Codex CLI) 129 (n 4, run range 56–259). Unclear.
- Output speed in characters per second after the first text (calculation)Calculation517n 4524n 4UnclearOutput speed in characters per second after the first text (calculation), calculation: Claude Sonnet 5.5 517 (n 4, run range 513–519); GPT-6 Luna (Codex CLI) 524 (n 4, run range 225–1,052). Unclear.
| Metric | Claude Sonnet 5.5 | GPT-6 Luna (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Strict pass rate: single call vs agent loop on eight hard tasks | 100% (24/24)Claude Code · single call | 63% (10/16)Codex CLI · single call | 24 / 16 | 95% CI: 86%–100% vs 39%–82% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%). | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Interval merge fix | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: DST day-length fix | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: CSV parser | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Event-loop order | 100% (3/3)Claude Code · single call | 0% (0/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 0%–66% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Room schedule | 100% (3/3)Claude Code · single call | 50% (1/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SemVer regex | 100% (3/3)Claude Code · single call | 50% (1/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Money refactor | 100% (3/3)Claude Code · single call | 0% (0/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 0%–66% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SQL report | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Total time per attempt: single call vs agent loop | 7.75 sClaude Code · single call | 5.16 sCodex CLI · single call | 24 / 16 | range: 2.3 s–34.8 s vs 3.6 s–11.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)) | 2,281Claude Code · single call | 11,582Codex CLI · single call | 24 / 16 | range: 2,234–2,669 vs 11,526–11,818 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Output tokens) | 1,050Claude Code · single call | 345Codex CLI · single call | 24 / 16 | range: 176–3,895 vs 36–634 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tool calls per agent-loop attempt | 0Claude Code · agent loop | 0Codex CLI · agent loop | 16 / 14 | range: 0–3 vs 0–1 | Tie | Same value. More or fewer is not better by itself for this metric. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| List-price cost per strict pass: single call vs agent loop (calculation)Calculation | $0.014Claude Code · single call | $0.0012Codex CLI · single call | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Time to first text: a 250-line answer, six models | 1.96 sClaude Code | 3.30 sCodex CLI · effort low | 4 | range: 0.9 s–4.1 s vs 3.2 s–3.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 232Claude Code | 129Codex CLI · effort low | 4 | range: 230–233 vs 56–259 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 517Claude Code | 524Codex CLI · effort low | 4 | range: 513–519 vs 225–1,052 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
17 rows from 2 studies. Claude Sonnet 5.5 ahead on 1; 9 ties, 7 unclear. A side is ahead only where the intervals or ranges do not overlap.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Claude Sonnet 5.5
- Strict pass rate: single call vs agent loop on eight hard tasks: 100% (24/24) vs 63% (10/16). The 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%).
When to pick GPT-6 Luna (Codex CLI)
- List-price cost per strict pass: single call vs agent loop (calculation): $0.0012 vs $0.014. A list-price calculation, not a measured difference. Calculation
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 (single call) · Claude Code | 100% | 86%–100% | 24 |
| GPT-6 Luna (single call) · Codex CLI | 63% | 39%–82% | 16 |
2 rows. Highest Claude Sonnet 5.5 (single call) · Claude Code 100% (95% interval 86%–100%, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 63% (95% interval 39%–82%, n 16). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 16–24 per row
Same tasks and validators. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Claude Haiku 4.5 (single call) · Claude Code | Claude Haiku 4.5 (agent loop) · Claude Code | Claude Sonnet 5.5 (single call) · Claude Code | Claude Sonnet 5.5 (agent loop) · Claude Code | GPT-6 Luna (single call) · Codex CLI | GPT-6 Luna (agent loop) · Codex CLI | 95% interval | n |
|---|---|---|---|---|---|---|---|---|
| Interval merge fix | 100% | 67% | 100% | 100% | 100% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 44%–100%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 34%–100%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
| Event-loop order | 0% | 100% | 100% | 100% | 0% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 0%–56%; Claude Haiku 4.5 (agent loop) · Claude Code: 44%–100%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 0%–66%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
| Room schedule | 0% | 67% | 100% | 100% | 50% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 0%–56%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 9.4%–91%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
3 rows, 6 series: Claude Haiku 4.5 (single call) · Claude Code, Claude Haiku 4.5 (agent loop) · Claude Code, Claude Sonnet 5.5 (single call) · Claude Code, Claude Sonnet 5.5 (agent loop) · Claude Code, GPT-6 Luna (single call) · Codex CLI, GPT-6 Luna (agent loop) · Codex CLI. Claude Haiku 4.5 (single call) · Claude Code: highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest Room schedule 0% (95% interval 0%–56%, n 3). All intervals overlap. Claude Haiku 4.5 (agent loop) · Claude Code: highest Event-loop order 100% (95% interval 44%–100%, n 3). Lowest Room schedule 67% (95% interval 21%–94%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 2–3 per row
Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration
Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 (single call) · Claude Code | 7.8 s | 2.3 s–34.8 s | 24 |
| GPT-6 Luna (single call) · Codex CLI | 5.2 s | 3.6 s–11.3 s | 16 |
2 rows. Slowest Claude Sonnet 5.5 (single call) · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). Fastest GPT-6 Luna (single call) · Codex CLI 5.2 s (range 3.6 s–11.3 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest attempt
Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
- Input tokens (cache reads included)
- Output tokens
| Item | Input tokens (cache reads included) | Output tokens | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 (single call) · Claude Code | 2,281 | 1,050 | Input tokens (cache reads included): 2,234–2,669; Output tokens: 176–3,895 | 24 |
| GPT-6 Luna (single call) · Codex CLI | 11,582 | 345 | Input tokens (cache reads included): 11,526–11,818; Output tokens: 36–634 | 16 |
2 rows, 2 series: Input tokens (cache reads included), Output tokens. Input tokens (cache reads included): highest GPT-6 Luna (single call) · Codex CLI 11,582 (range 11,526–11,818, n 16). Lowest Claude Sonnet 5.5 (single call) · Claude Code 2,281 (range 2,234–2,669, n 24). Not all run ranges overlap. Output tokens: highest Claude Sonnet 5.5 (single call) · Claude Code 1,050 (range 176–3,895, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 345 (range 36–634, n 16). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fewest and most
Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Tool calls per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 (agent loop) · Claude Code | 0 | 0–3 | 16 |
| GPT-6 Luna (agent loop) · Codex CLI | 0 | 0–1 | 14 |
2 rows. All at 0.
NotesLines: lowest–highest run (not an interval)n 14–16 per row
Median per configuration; whiskers = fewest and most. A single call makes none
Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.
Source: Single call vs agent loop
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 (single call) · Claude Code | $0.014 | 24 |
| GPT-6 Luna (single call) · Codex CLI | $0.0012 | 16 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 (single call) · Claude Code $0.014 (n 24). Lowest GPT-6 Luna (single call) · Codex CLI $0.0012 (n 16).
Notesn 16–24 per row
All attempts in a configuration divided by its strict passes
Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 2 s | 0.9 s–4.1 s | 4 |
| GPT-6 Luna (low) · Codex CLI | 3.3 s | 3.2 s–3.5 s | 4 |
2 rows. Slowest GPT-6 Luna (low) · Codex CLI 3.3 s (range 3.2 s–3.5 s, n 4). Fastest Claude Sonnet 5.5 · Claude Code 2 s (range 0.9 s–4.1 s, n 4). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
| Item | Visible tokens per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 232 | 230–233 | 4 |
| GPT-6 Luna (low) · Codex CLI | 129 | 56–259 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 · Claude Code 232 (range 230–233, n 4). Lowest GPT-6 Luna (low) · Codex CLI 129 (range 56–259, n 4). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = slowest and fastest call
Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.
Source: LLM speed anatomy
| Item | Characters per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 517 | 513–519 | 4 |
| GPT-6 Luna (low) · Codex CLI | 524 | 225–1,052 | 4 |
List-price calculation, not a run. 2 rows. Highest GPT-6 Luna (low) · Codex CLI 524 (range 225–1,052, n 4). Lowest Claude Sonnet 5.5 · Claude Code 517 (range 513–519, n 4). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call
Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Claude Sonnet 5.5 or GPT-6 Luna (Codex CLI)?
- Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. Claude Sonnet 5.5 leads on 1 row: Strict pass rate: single call vs agent loop on eight hard tasks, 100% (24/24) vs 63% (10/16). On those rows the 95% intervals do not overlap. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).
- How were Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) measured?
- They share 14 measured metrics and 3 list-price calculations from 2 public studies: Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks; Where the seconds go: first text, output speed and prompt size for 6 LLMs. Every row names its configuration, its sample size and its interval or range.
- How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict pass rate: single call vs agent loop on eight hard tasks?
- Claude Sonnet 5.5: 100% (24/24) (Claude Code · single call; n = 24; 95% interval 86% to 100%). GPT-6 Luna (Codex CLI): 63% (10/16) (Codex CLI · single call; n = 16; 95% interval 39% to 82%). The 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%).
- How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: Interval merge fix?
- Claude Sonnet 5.5: 100% (3/3) (Claude Code · single call; n = 3; 95% interval 44% to 100%). GPT-6 Luna (Codex CLI): 100% (2/2) (Codex CLI · single call; n = 2; 95% interval 34% to 100%). The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.
- How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: DST day-length fix?
- Claude Sonnet 5.5: 100% (3/3) (Claude Code · single call; n = 3; 95% interval 44% to 100%). GPT-6 Luna (Codex CLI): 100% (2/2) (Codex CLI · single call; n = 2; 95% interval 34% to 100%). The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.
- How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: CSV parser?
- Claude Sonnet 5.5: 100% (3/3) (Claude Code · single call; n = 3; 95% interval 44% to 100%). GPT-6 Luna (Codex CLI): 100% (2/2) (Codex CLI · single call; n = 2; 95% interval 34% to 100%). The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.
- How do Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: Event-loop order?
- Claude Sonnet 5.5: 100% (3/3) (Claude Code · single call; n = 3; 95% interval 44% to 100%). GPT-6 Luna (Codex CLI): 0% (0/2) (Codex CLI · single call; n = 2; 95% interval 0% to 66%). The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.
The studies behind this page
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
118 attempts on 8 hard tasks: one call vs an agent loop that runs code in a sandbox. Pass rate, time, tokens and cost.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.