14 measured metrics · 3 calculated · 2 studies
Claude Haiku 4.5vsGPT-6 Luna (Codex CLI)
GPT-6 Luna (Codex CLI) ahead on 1; 9 ties, 7 unclear. A side is ahead only where the intervals or ranges do not overlap.
The verdict
Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. GPT-6 Luna (Codex CLI) leads on 1 row: Total time per attempt: single call vs agent loop, 5.16 s vs 39.0 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Claude Haiku 4.5 | GPT-6 Luna (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Strict pass rate: single call vs agent loop on eight hard tasks | 46% (11/24)Claude Code · single call | 63% (10/16)Codex CLI · single call | 24 / 16 | 95% CI: 28%–65% vs 39%–82% | Tie | The 95% intervals overlap (Claude Haiku 4.5 28% to 65%; GPT-6 Luna (Codex CLI) 39% to 82%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Total time per attempt: single call vs agent loop | 39.0 sClaude Code · single call | 5.16 sCodex CLI · single call | 24 / 16 | range: 15.3 s–75.1 s vs 3.6 s–11.3 s | GPT-6 Luna (Codex CLI) ahead | The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s). A range is not a confidence interval. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)) | 3,941Claude Code · single call | 11,582Codex CLI · single call | 24 / 16 | range: 3,879–4,221 vs 11,526–11,818 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Output tokens) | 5,064Claude Code · single call | 345Codex CLI · single call | 24 / 16 | range: 1,899–9,321 vs 36–634 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| List-price cost per strict pass: single call vs agent loop (calculation)Calculation | $0.067Claude Code · single call | $0.0012Codex CLI · single call | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.067 vs $0.0012, 58x) is not tested against run-to-run variation. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
5 headline metrics as the ratio of the two values. 1 of them separate the sides in the data. Widest ratio: List-price cost per strict pass: single call vs agent loop (calculation), 58x (Claude Haiku 4.5 larger).
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Claude Haiku 4.5
- GPT-6 Luna (Codex CLI)
- 95% interval
- fastest–slowest run (not an interval)
- where the two overlap
- hollow: list-price calculation
These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
- Strict pass rate: single call vs agent loop on eight hard tasks46% (11/24)n 2463% (10/16)n 16TieStrict pass rate: single call vs agent loop on eight hard tasks: Claude Haiku 4.5 46% (11/24) (n 24, 95% interval 28%–65%); GPT-6 Luna (Codex CLI) 63% (10/16) (n 16, 95% interval 39%–82%). Tie.
Calculation: at these rates, about 132 runs per side would separate them.
- Strict passes per task: single call vs agent loop: Interval merge fix100% (3/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: Interval merge fix: Claude Haiku 4.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
- Strict passes per task: single call vs agent loop: DST day-length fix33% (1/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: DST day-length fix: Claude Haiku 4.5 33% (1/3) (n 3, 95% interval 6.2%–79%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
Calculation: at these rates, about 7 runs per side would separate them.
- Strict passes per task: single call vs agent loop: CSV parser67% (2/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: CSV parser: Claude Haiku 4.5 67% (2/3) (n 3, 95% interval 21%–94%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
Calculation: at these rates, about 20 runs per side would separate them.
- Strict passes per task: single call vs agent loop: Event-loop order0% (0/3)n 30% (0/2)n 2TieStrict passes per task: single call vs agent loop: Event-loop order: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); GPT-6 Luna (Codex CLI) 0% (0/2) (n 2, 95% interval 0%–66%). Tie.
- Strict passes per task: single call vs agent loop: Room schedule0% (0/3)n 350% (1/2)n 2TieStrict passes per task: single call vs agent loop: Room schedule: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); GPT-6 Luna (Codex CLI) 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.
Calculation: at these rates, about 11 runs per side would separate them.
- Strict passes per task: single call vs agent loop: SemVer regex100% (3/3)n 350% (1/2)n 2TieStrict passes per task: single call vs agent loop: SemVer regex: Claude Haiku 4.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6 Luna (Codex CLI) 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.
Calculation: at these rates, about 12 runs per side would separate them.
- Strict passes per task: single call vs agent loop: Money refactor67% (2/3)n 30% (0/2)n 2TieStrict passes per task: single call vs agent loop: Money refactor: Claude Haiku 4.5 67% (2/3) (n 3, 95% interval 21%–94%); GPT-6 Luna (Codex CLI) 0% (0/2) (n 2, 95% interval 0%–66%). Tie.
Calculation: at these rates, about 7 runs per side would separate them.
- Strict passes per task: single call vs agent loop: SQL report0% (0/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: SQL report: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); GPT-6 Luna (Codex CLI) 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
Calculation: at these rates, about 4 runs per side would separate them.
- Total time per attempt: single call vs agent loop39.0 sn 245.16 sn 16GPT-6 Luna (Codex CLI) aheadTotal time per attempt: single call vs agent loop: Claude Haiku 4.5 39.0 s (n 24, run range 15.3 s–75.1 s); GPT-6 Luna (Codex CLI) 5.16 s (n 16, run range 3.6 s–11.3 s). GPT-6 Luna (Codex CLI) ahead.
- Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))3,941n 2411,582n 16UnclearTokens per attempt: single call vs agent loop (Input tokens (cache reads included)): Claude Haiku 4.5 3,941 (n 24, run range 3,879–4,221); GPT-6 Luna (Codex CLI) 11,582 (n 16, run range 11,526–11,818). Unclear.
- Tokens per attempt: single call vs agent loop (Output tokens)5,064n 24345n 16UnclearTokens per attempt: single call vs agent loop (Output tokens): Claude Haiku 4.5 5,064 (n 24, run range 1,899–9,321); GPT-6 Luna (Codex CLI) 345 (n 16, run range 36–634). Unclear.
- Tool calls per agent-loop attempt3n 240n 14UnclearTool calls per agent-loop attempt: Claude Haiku 4.5 3 (n 24, run range 2–18); GPT-6 Luna (Codex CLI) 0 (n 14, run range 0–1). Unclear.
- List-price cost per strict pass: single call vs agent loop (calculation)Calculation$0.067n 24$0.0012n 16UnclearList-price cost per strict pass: single call vs agent loop (calculation), calculation: Claude Haiku 4.5 $0.067 (n 24); GPT-6 Luna (Codex CLI) $0.0012 (n 16). Unclear.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
- Time to first text: a 250-line answer, six models4.00 sn 43.30 sn 4UnclearTime to first text: a 250-line answer, six models: Claude Haiku 4.5 4.00 s (n 4, run range 2.8 s–6.4 s); GPT-6 Luna (Codex CLI) 3.30 s (n 4, run range 3.2 s–3.5 s). Unclear.
- Output speed after the first text: visible tokens per second (calculation)Calculation153n 4129n 4UnclearOutput speed after the first text: visible tokens per second (calculation), calculation: Claude Haiku 4.5 153 (n 4, run range 153–216); GPT-6 Luna (Codex CLI) 129 (n 4, run range 56–259). Unclear.
- Output speed in characters per second after the first text (calculation)Calculation547n 3524n 4UnclearOutput speed in characters per second after the first text (calculation), calculation: Claude Haiku 4.5 547 (n 3, run range 546–548); GPT-6 Luna (Codex CLI) 524 (n 4, run range 225–1,052). Unclear.
| Metric | Claude Haiku 4.5 | GPT-6 Luna (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Strict pass rate: single call vs agent loop on eight hard tasks | 46% (11/24)Claude Code · single call | 63% (10/16)Codex CLI · single call | 24 / 16 | 95% CI: 28%–65% vs 39%–82% | Tie | The 95% intervals overlap (Claude Haiku 4.5 28% to 65%; GPT-6 Luna (Codex CLI) 39% to 82%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Interval merge fix | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: DST day-length fix | 33% (1/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 6.2%–79% vs 34%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 6% to 79%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: CSV parser | 67% (2/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 21%–94% vs 34%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Event-loop order | 0% (0/3)Claude Code · single call | 0% (0/2)Codex CLI · single call | 3 / 2 | 95% CI: 0%–56% vs 0%–66% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Room schedule | 0% (0/3)Claude Code · single call | 50% (1/2)Codex CLI · single call | 3 / 2 | 95% CI: 0%–56% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SemVer regex | 100% (3/3)Claude Code · single call | 50% (1/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Money refactor | 67% (2/3)Claude Code · single call | 0% (0/2)Codex CLI · single call | 3 / 2 | 95% CI: 21%–94% vs 0%–66% | Tie | The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SQL report | 0% (0/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 0%–56% vs 34%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Total time per attempt: single call vs agent loop | 39.0 sClaude Code · single call | 5.16 sCodex CLI · single call | 24 / 16 | range: 15.3 s–75.1 s vs 3.6 s–11.3 s | GPT-6 Luna (Codex CLI) ahead | The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s). A range is not a confidence interval. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)) | 3,941Claude Code · single call | 11,582Codex CLI · single call | 24 / 16 | range: 3,879–4,221 vs 11,526–11,818 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Output tokens) | 5,064Claude Code · single call | 345Codex CLI · single call | 24 / 16 | range: 1,899–9,321 vs 36–634 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tool calls per agent-loop attempt | 3Claude Code · agent loop | 0Codex CLI · agent loop | 24 / 14 | range: 2–18 vs 0–1 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| List-price cost per strict pass: single call vs agent loop (calculation)Calculation | $0.067Claude Code · single call | $0.0012Codex CLI · single call | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.067 vs $0.0012, 58x) is not tested against run-to-run variation. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Time to first text: a 250-line answer, six models | 4.00 sClaude Code | 3.30 sCodex CLI · effort low | 4 | range: 2.8 s–6.4 s vs 3.2 s–3.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 153Claude Code | 129Codex CLI · effort low | 4 | range: 153–216 vs 56–259 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 547Claude Code | 524Codex CLI · effort low | 3 / 4 | range: 546–548 vs 225–1,052 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
17 rows from 2 studies. GPT-6 Luna (Codex CLI) ahead on 1; 9 ties, 7 unclear. A side is ahead only where the intervals or ranges do not overlap.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Claude Haiku 4.5
No row in this data puts Claude Haiku 4.5 ahead of GPT-6 Luna (Codex CLI). Pick on other grounds (price, access, the tasks you run), or measure your own workload.
When to pick GPT-6 Luna (Codex CLI)
- Total time per attempt: single call vs agent loop: 5.16 s vs 39.0 s. The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s). A range is not a confidence interval.
- List-price cost per strict pass: single call vs agent loop (calculation): $0.0012 vs $0.067. A list-price calculation, not a measured difference. Calculation
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| Claude Haiku 4.5 (single call) · Claude Code | 46% | 28%–65% | 24 |
| GPT-6 Luna (single call) · Codex CLI | 63% | 39%–82% | 16 |
2 rows. Highest GPT-6 Luna (single call) · Codex CLI 63% (95% interval 39%–82%, n 16). Lowest Claude Haiku 4.5 (single call) · Claude Code 46% (95% interval 28%–65%, n 24). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 16–24 per row
Same tasks and validators. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Claude Haiku 4.5 (single call) · Claude Code | Claude Haiku 4.5 (agent loop) · Claude Code | Claude Sonnet 5.5 (single call) · Claude Code | Claude Sonnet 5.5 (agent loop) · Claude Code | GPT-6 Luna (single call) · Codex CLI | GPT-6 Luna (agent loop) · Codex CLI | 95% interval | n |
|---|---|---|---|---|---|---|---|---|
| Interval merge fix | 100% | 67% | 100% | 100% | 100% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 44%–100%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 34%–100%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
| DST day-length fix | 33% | 67% | 100% | 100% | 100% | — | Claude Haiku 4.5 (single call) · Claude Code: 6.2%–79%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 34%–100% | 3 |
| CSV parser | 67% | 67% | 100% | 100% | 100% | 50% | Claude Haiku 4.5 (single call) · Claude Code: 21%–94%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 34%–100%; GPT-6 Luna (agent loop) · Codex CLI: 9.4%–91% | 3 |
| Event-loop order | 0% | 100% | 100% | 100% | 0% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 0%–56%; Claude Haiku 4.5 (agent loop) · Claude Code: 44%–100%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 0%–66%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
| Room schedule | 0% | 67% | 100% | 100% | 50% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 0%–56%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 9.4%–91%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
5 rows, 6 series: Claude Haiku 4.5 (single call) · Claude Code, Claude Haiku 4.5 (agent loop) · Claude Code, Claude Sonnet 5.5 (single call) · Claude Code, Claude Sonnet 5.5 (agent loop) · Claude Code, GPT-6 Luna (single call) · Codex CLI, GPT-6 Luna (agent loop) · Codex CLI. Claude Haiku 4.5 (single call) · Claude Code: highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest Room schedule 0% (95% interval 0%–56%, n 3). All intervals overlap. Claude Haiku 4.5 (agent loop) · Claude Code: highest Event-loop order 100% (95% interval 44%–100%, n 3). Lowest Room schedule 67% (95% interval 21%–94%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 2–3 per row
Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration
Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
| Item | Total time per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 (single call) · Claude Code | 39 s | 15.3 s–75.1 s | 24 |
| GPT-6 Luna (single call) · Codex CLI | 5.2 s | 3.6 s–11.3 s | 16 |
2 rows. Slowest Claude Haiku 4.5 (single call) · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest GPT-6 Luna (single call) · Codex CLI 5.2 s (range 3.6 s–11.3 s, n 16). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest attempt
Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
- Input tokens (cache reads included)
- Output tokens
| Item | Input tokens (cache reads included) | Output tokens | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Haiku 4.5 (single call) · Claude Code | 3,941 | 5,064 | Input tokens (cache reads included): 3,879–4,221; Output tokens: 1,899–9,321 | 24 |
| GPT-6 Luna (single call) · Codex CLI | 11,582 | 345 | Input tokens (cache reads included): 11,526–11,818; Output tokens: 36–634 | 16 |
2 rows, 2 series: Input tokens (cache reads included), Output tokens. Input tokens (cache reads included): highest GPT-6 Luna (single call) · Codex CLI 11,582 (range 11,526–11,818, n 16). Lowest Claude Haiku 4.5 (single call) · Claude Code 3,941 (range 3,879–4,221, n 24). Not all run ranges overlap. Output tokens: highest Claude Haiku 4.5 (single call) · Claude Code 5,064 (range 1,899–9,321, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 345 (range 36–634, n 16). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fewest and most
Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Tool calls per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 (agent loop) · Claude Code | 3 | 2–18 | 24 |
| GPT-6 Luna (agent loop) · Codex CLI | 0 | 0–1 | 14 |
2 rows. Highest Claude Haiku 4.5 (agent loop) · Claude Code 3 (range 2–18, n 24). Lowest GPT-6 Luna (agent loop) · Codex CLI 0 (range 0–1, n 14). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 14–24 per row
Median per configuration; whiskers = fewest and most. A single call makes none
Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.
Source: Single call vs agent loop
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Haiku 4.5 (single call) · Claude Code | $0.067 | 24 |
| GPT-6 Luna (single call) · Codex CLI | $0.0012 | 16 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 (single call) · Claude Code $0.067 (n 24). Lowest GPT-6 Luna (single call) · Codex CLI $0.0012 (n 16).
Notesn 16–24 per row
All attempts in a configuration divided by its strict passes
Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 4 s | 2.8 s–6.4 s | 4 |
| GPT-6 Luna (low) · Codex CLI | 3.3 s | 3.2 s–3.5 s | 4 |
2 rows. Slowest Claude Haiku 4.5 · Claude Code 4 s (range 2.8 s–6.4 s, n 4). Fastest GPT-6 Luna (low) · Codex CLI 3.3 s (range 3.2 s–3.5 s, n 4). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
| Item | Visible tokens per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 153 | 153–216 | 4 |
| GPT-6 Luna (low) · Codex CLI | 129 | 56–259 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 · Claude Code 153 (range 153–216, n 4). Lowest GPT-6 Luna (low) · Codex CLI 129 (range 56–259, n 4). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = slowest and fastest call
Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.
Source: LLM speed anatomy
| Item | Characters per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 547 | 546–548 | 3 |
| GPT-6 Luna (low) · Codex CLI | 524 | 225–1,052 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 · Claude Code 547 (range 546–548, n 3). Lowest GPT-6 Luna (low) · Codex CLI 524 (range 225–1,052, n 4). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 3–4 per row
Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call
Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Claude Haiku 4.5 or GPT-6 Luna (Codex CLI)?
- Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. GPT-6 Luna (Codex CLI) leads on 1 row: Total time per attempt: single call vs agent loop, 5.16 s vs 39.0 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).
- How were Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) measured?
- They share 14 measured metrics and 3 list-price calculations from 2 public studies: Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks; Where the seconds go: first text, output speed and prompt size for 6 LLMs. Every row names its configuration, its sample size and its interval or range.
- How do Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) compare on total time per attempt: single call vs agent loop?
- Claude Haiku 4.5: 39.0 s (Claude Code · single call; n = 24; run range 15.3 s to 75.1 s). GPT-6 Luna (Codex CLI): 5.16 s (Codex CLI · single call; n = 16; run range 3.6 s to 11.3 s). The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s). A range is not a confidence interval.
- How do Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) compare on strict pass rate: single call vs agent loop on eight hard tasks?
- Claude Haiku 4.5: 46% (11/24) (Claude Code · single call; n = 24; 95% interval 28% to 65%). GPT-6 Luna (Codex CLI): 63% (10/16) (Codex CLI · single call; n = 16; 95% interval 39% to 82%). The 95% intervals overlap (Claude Haiku 4.5 28% to 65%; GPT-6 Luna (Codex CLI) 39% to 82%), so this sample cannot separate them.
- How do Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: Interval merge fix?
- Claude Haiku 4.5: 100% (3/3) (Claude Code · single call; n = 3; 95% interval 44% to 100%). GPT-6 Luna (Codex CLI): 100% (2/2) (Codex CLI · single call; n = 2; 95% interval 34% to 100%). The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.
- How do Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: DST day-length fix?
- Claude Haiku 4.5: 33% (1/3) (Claude Code · single call; n = 3; 95% interval 6.2% to 79%). GPT-6 Luna (Codex CLI): 100% (2/2) (Codex CLI · single call; n = 2; 95% interval 34% to 100%). The 95% intervals overlap (Claude Haiku 4.5 6% to 79%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.
- How do Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) compare on strict passes per task: single call vs agent loop: CSV parser?
- Claude Haiku 4.5: 67% (2/3) (Claude Code · single call; n = 3; 95% interval 21% to 94%). GPT-6 Luna (Codex CLI): 100% (2/2) (Codex CLI · single call; n = 2; 95% interval 34% to 100%). The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.
The studies behind this page
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
118 attempts on 8 hard tasks: one call vs an agent loop that runs code in a sandbox. Pass rate, time, tokens and cost.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.