40 measured metrics · 16 calculated · 7 studies

Claude Haiku 4.5vsGPT-6.1 Sol (Codex CLI)

Claude Haiku 4.5 ahead on 2, GPT-6.1 Sol (Codex CLI) ahead on 6; 13 ties, 35 unclear. A side is ahead only where the intervals or ranges do not overlap.

The verdict

Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) share 40 measured metrics and 16 list-price calculations from 7 studies. Claude Haiku 4.5 leads on 2 rows: Same prompt, 10 times: time per call (Exact number), 5.06 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 5.95 s vs 11.3 s. GPT-6.1 Sol (Codex CLI) leads on 6 rows: Pass rate on eight hard tasks (Strict pass), 100% (16/16) vs 46% (11/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); Same prompt, 10 times: strict pass rate (JSON object), 100% (10/10) vs 10% (1/10); and 3 more. On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 13 ties and 35 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Output tokens per call (Output tokens), 8.7x (Claude Haiku 4.5 larger).

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Haiku 4.5
  • GPT-6.1 Sol (Codex CLI)
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

1 tie · 7 unclear
  • Pass rate on five validated tasks: Claude Haiku 4.5 100% (15/15) (n 15, 95% interval 80%–100%); GPT-6.1 Sol (Codex CLI) 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
  • Total time per call: Claude Haiku 4.5 4.43 s (n 15, run range 3.2 s–23.6 s); GPT-6.1 Sol (Codex CLI) 5.65 s (n 15, run range 4.1 s–25.5 s). Unclear.
  • Time to first useful output: Claude Haiku 4.5 3.63 s (n 15, run range 2.8 s–22.3 s); GPT-6.1 Sol (Codex CLI) 5.05 s (n 15, run range 3.4 s–17.8 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): Claude Haiku 4.5 0 (n 15); GPT-6.1 Sol (Codex CLI) 5,180 (n 15). Unclear.
  • Input tokens per call: what the CLI sends (Other input): Claude Haiku 4.5 3,790 (n 15); GPT-6.1 Sol (Codex CLI) 6,943 (n 15). Unclear.
  • Output tokens per call (Output tokens): Claude Haiku 4.5 367 (n 15); GPT-6.1 Sol (Codex CLI) 42 (n 15). Unclear.
  • List-price cost per call (calculation), calculation: Claude Haiku 4.5 $0.0057 (n 15, run range $0.0051–$0.018); GPT-6.1 Sol (Codex CLI) $0.010 (n 15, run range $0.0054–$0.027). Unclear.
  • List-price cost per passing answer (calculation), calculation: Claude Haiku 4.5 $0.0084 (n 15); GPT-6.1 Sol (Codex CLI) $0.016 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

GPT-6.1 Sol (Codex CLI) 1 · 1 tie · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): Claude Haiku 4.5 46% (11/24) (n 24, 95% interval 28%–65%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). GPT-6.1 Sol (Codex CLI) ahead.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): Claude Haiku 4.5 67% (16/24) (n 24, 95% interval 47%–82%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.

    Calculation: at these rates, about 20 runs per side would separate them.

  • Total time per call on hard tasks (separate batches): Claude Haiku 4.5 39.0 s (n 24, run range 15.3 s–75.1 s); GPT-6.1 Sol (Codex CLI) 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
  • Time to first useful output on hard tasks: Claude Haiku 4.5 35.5 s (n 24, run range 12.9 s–70.3 s); GPT-6.1 Sol (Codex CLI) 10.2 s (n 16, run range 6.1 s–40.4 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): Claude Haiku 4.5 5,064 (n 24); GPT-6.1 Sol (Codex CLI) 335 (n 16). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: Claude Haiku 4.5 $0.067 (n 24); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

Claude Haiku 4.5 2 · GPT-6.1 Sol (Codex CLI) 2 · 4 ties · 1 unclear
  • Same prompt, 10 times: strict pass rate (Exact number): Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); GPT-6.1 Sol (Codex CLI) 100% (10/10) (n 10, 95% interval 72%–100%). GPT-6.1 Sol (Codex CLI) ahead.
  • Same prompt, 10 times: strict pass rate (JSON object): Claude Haiku 4.5 10% (1/10) (n 10, 95% interval 1.8%–40%); GPT-6.1 Sol (Codex CLI) 100% (10/10) (n 10, 95% interval 72%–100%). GPT-6.1 Sol (Codex CLI) ahead.
  • Same prompt, 10 times: strict pass rate (Code fix): Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); GPT-6.1 Sol (Codex CLI) 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
  • Same prompt, 10 times: how many different answers (Exact number): Claude Haiku 4.5 1 (n 10); GPT-6.1 Sol (Codex CLI) 1 (n 10). Tie.
  • Same prompt, 10 times: how many different answers (JSON object): Claude Haiku 4.5 1 (n 10); GPT-6.1 Sol (Codex CLI) 1 (n 10). Tie.
  • Same prompt, 10 times: how many different answers (Code fix): Claude Haiku 4.5 6 (n 10); GPT-6.1 Sol (Codex CLI) 6 (n 10). Tie.
  • Same prompt, 10 times: time per call (Exact number): Claude Haiku 4.5 5.06 s (n 10, run range 4.4 s–6.2 s); GPT-6.1 Sol (Codex CLI) 13.4 s (n 10, run range 12.3 s–18 s). Claude Haiku 4.5 ahead.
  • Same prompt, 10 times: time per call (JSON object): Claude Haiku 4.5 7.03 s (n 10, run range 5.3 s–12.3 s); GPT-6.1 Sol (Codex CLI) 6.42 s (n 10, run range 5.3 s–8.3 s). Unclear.
  • Same prompt, 10 times: time per call (Code fix): Claude Haiku 4.5 5.95 s (n 10, run range 4.9 s–7.3 s); GPT-6.1 Sol (Codex CLI) 11.3 s (n 10, run range 9.1 s–14.9 s). Claude Haiku 4.5 ahead.

Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI

GPT-6.1 Sol (Codex CLI) 1 · 2 ties · 5 unclear
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON), calculation: Claude Haiku 4.5 0% (0/24) (n 24, 95% interval 0%–14%); GPT-6.1 Sol (Codex CLI) 100% (12/12) (n 12, 95% interval 76%–100%). GPT-6.1 Sol (Codex CLI) ahead.
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss)), calculation: Claude Haiku 4.5 71% (17/24) (n 24, 95% interval 51%–85%); GPT-6.1 Sol (Codex CLI) 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Strict pass): Claude Haiku 4.5 0 (n 24); GPT-6.1 Sol (Codex CLI) 12 (n 12). Unclear.
  • What each call produced: strict pass, format miss, wrong values or error (Format miss): Claude Haiku 4.5 17 (n 24); GPT-6.1 Sol (Codex CLI) 0 (n 12). Unclear.
  • What each call produced: strict pass, format miss, wrong values or error (Wrong values): Claude Haiku 4.5 7 (n 24); GPT-6.1 Sol (Codex CLI) 0 (n 12). Unclear.
  • What each call produced: strict pass, format miss, wrong values or error (Error): Claude Haiku 4.5 0 (n 24); GPT-6.1 Sol (Codex CLI) 0 (n 12). Tie.
  • Time per call, instructions vs schema mode: Claude Haiku 4.5 9.52 s (n 24, run range 5.7 s–17 s); GPT-6.1 Sol (Codex CLI) 6.21 s (n 12, run range 4.2 s–12.3 s). Unclear.
  • Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call): Claude Haiku 4.5 1,128 (n 24); GPT-6.1 Sol (Codex CLI) 117 (n 12). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

6 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Haiku 4.5 91.7% (n 24, run range 76%–99%); GPT-6.1 Sol (Codex CLI) 46.3% (n 16, run range 11%–87%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Haiku 4.5 $0.024 (n 24); GPT-6.1 Sol (Codex CLI) $0.0023 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Haiku 4.5 $0.0018 (n 24); GPT-6.1 Sol (Codex CLI) $0.0030 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Haiku 4.5 $0.0045 (n 24); GPT-6.1 Sol (Codex CLI) $0.020 (n 16). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Haiku 4.5 91.7% (n 24, run range 76%–99%); GPT-6.1 Sol (Codex CLI) 46.3% (n 16, run range 11%–87%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Haiku 4.5 90.2% (n 15, run range 73%–98%); GPT-6.1 Sol (Codex CLI) 41% (n 15, run range 0%–71%). Unclear.

Where the seconds go: first text, output speed and prompt size for 6 LLMs

1 tie · 9 unclear
  • Time to first text: a 250-line answer, six models: Claude Haiku 4.5 4.00 s (n 4, run range 2.8 s–6.4 s); GPT-6.1 Sol (Codex CLI) 3.52 s (n 4, run range 2.8 s–4.4 s). Unclear.
  • Output speed after the first text: visible tokens per second (calculation), calculation: Claude Haiku 4.5 153 (n 4, run range 153–216); GPT-6.1 Sol (Codex CLI) 80 (n 4, run range 72–81). Unclear.
  • Output speed in characters per second after the first text (calculation), calculation: Claude Haiku 4.5 547 (n 3, run range 546–548); GPT-6.1 Sol (Codex CLI) 323 (n 4, run range 291–327). Unclear.
  • Time to first text as the prompt grows: 1k, calculation: Claude Haiku 4.5 1.93 s (n 3, run range 1.9 s–2 s); GPT-6.1 Sol (Codex CLI) 3.36 s (n 3, run range 3.4 s–4.8 s). Unclear.
  • Time to first text as the prompt grows: 16k, calculation: Claude Haiku 4.5 2.27 s (n 3, run range 2.2 s–2.5 s); GPT-6.1 Sol (Codex CLI) 4.02 s (n 3, run range 3.3 s–4.3 s). Unclear.
  • Time to first text as the prompt grows: 64k, calculation: Claude Haiku 4.5 2.78 s (n 3, run range 2.5 s–2.9 s); GPT-6.1 Sol (Codex CLI) 3.93 s (n 3, run range 3.4 s–4.4 s). Unclear.
  • Total time per call by prompt size (1k prompt): Claude Haiku 4.5 2.34 s (n 3, run range 2.2 s–2.5 s); GPT-6.1 Sol (Codex CLI) 3.43 s (n 3, run range 3.4 s–4.9 s). Unclear.
  • Total time per call by prompt size (16k prompt): Claude Haiku 4.5 2.79 s (n 3, run range 2.6 s–2.8 s); GPT-6.1 Sol (Codex CLI) 4.14 s (n 3, run range 4 s–4.7 s). Unclear.
  • Total time per call by prompt size (64k prompt): Claude Haiku 4.5 3.13 s (n 3, run range 2.8 s–3.3 s); GPT-6.1 Sol (Codex CLI) 3.96 s (n 3, run range 3.5 s–4.4 s). Unclear.
  • Exact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Haiku 4.5 100% (9/9) (n 9, 95% interval 70%–100%); GPT-6.1 Sol (Codex CLI) 100% (9/9) (n 9, 95% interval 70%–100%). Tie.

GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

GPT-6.1 Sol (Codex CLI) 2 · 4 ties · 3 unclear
  • Pass rate on 4 harder tasks (Strict pass): Claude Haiku 4.5 0% (0/12) (n 12, 95% interval 0%–24%); GPT-6.1 Sol (Codex CLI) 69% (11/16) (n 16, 95% interval 44%–86%). GPT-6.1 Sol (Codex CLI) ahead.
  • Pass rate on 4 harder tasks (Lenient (format misses counted)): Claude Haiku 4.5 0% (0/12) (n 12, 95% interval 0%–24%); GPT-6.1 Sol (Codex CLI) 69% (11/16) (n 16, 95% interval 44%–86%). GPT-6.1 Sol (Codex CLI) ahead.
  • Calls that tried a tool although tools were off: Claude Haiku 4.5 8% (1/12) (n 12, 95% interval 1.5%–35%); GPT-6.1 Sol (Codex CLI) 0% (0/16) (n 16, 95% interval 0%–19%). Unclear.
  • Strict pass rate by task: 10x10 nonogram: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); GPT-6.1 Sol (Codex CLI) 75% (3/4) (n 4, 95% interval 30%–95%). Tie.

    Calculation: at these rates, about 6 runs per side would separate them.

  • Strict pass rate by task: Sudoku, 22 givens: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); GPT-6.1 Sol (Codex CLI) 25% (1/4) (n 4, 95% interval 4.6%–70%). Tie.

    Calculation: at these rates, about 26 runs per side would separate them.

  • Strict pass rate by task: 6x6 Skyscrapers: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); GPT-6.1 Sol (Codex CLI) 100% (4/4) (n 4, 95% interval 51%–100%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Strict pass rate by task: Seeded shuffle output: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); GPT-6.1 Sol (Codex CLI) 75% (3/4) (n 4, 95% interval 30%–95%). Tie.

    Calculation: at these rates, about 6 runs per side would separate them.

  • Total time per call on harder tasks: Claude Haiku 4.5 109.0 s (n 10, run range 25.7 s–224 s); GPT-6.1 Sol (Codex CLI) 120.2 s (n 13, run range 46.2 s–273 s). Unclear.
  • Output tokens per call on harder tasks (Output tokens): Claude Haiku 4.5 12,508 (n 10, run range 2,965–26,532); GPT-6.1 Sol (Codex CLI) 4,994 (n 13, run range 2,099–13,413). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

56 rows from 7 studies. Claude Haiku 4.5 ahead on 2, GPT-6.1 Sol (Codex CLI) ahead on 6; 13 ties, 35 unclear. A side is ahead only where the intervals or ranges do not overlap.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Haiku 4.5

  • List-price cost per call (calculation): $0.0057 vs $0.010. A list-price calculation, not a measured difference. Calculation
  • List-price cost per passing answer (calculation): $0.0084 vs $0.016. A list-price calculation, not a measured difference. Calculation
  • Same prompt, 10 times: time per call (Exact number): 5.06 s vs 13.4 s. The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.42 s to 6.20 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.
  • Same prompt, 10 times: time per call (Code fix): 5.95 s vs 11.3 s. The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0018 vs $0.0030. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)): $0.0045 vs $0.020. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 1k: 1.93 s vs 3.36 s. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 16k: 2.27 s vs 4.02 s. A list-price calculation, not a measured difference. Calculation

When to pick GPT-6.1 Sol (Codex CLI)

  • Pass rate on eight hard tasks (Strict pass): 100% (16/16) vs 46% (11/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; GPT-6.1 Sol (Codex CLI) 81% to 100%).
  • List-price cost per strict pass on hard tasks (calculation): $0.026 vs $0.067. A list-price calculation, not a measured difference. Calculation
  • Same prompt, 10 times: strict pass rate (Exact number): 100% (10/10) vs 0% (0/10). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; GPT-6.1 Sol (Codex CLI) 72% to 100%).
  • Same prompt, 10 times: strict pass rate (JSON object): 100% (10/10) vs 10% (1/10). The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; GPT-6.1 Sol (Codex CLI) 72% to 100%).
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON): 100% (12/12) vs 0% (0/24). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; GPT-6.1 Sol (Codex CLI) 76% to 100%).
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0023 vs $0.024. A list-price calculation, not a measured difference. Calculation
  • Pass rate on 4 harder tasks (Strict pass): 69% (11/16) vs 0% (0/12). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%).
  • Pass rate on 4 harder tasks (Lenient (format misses counted)): 69% (11/16) vs 0% (0/12). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%).

Side by side

The study charts, showing only these two. Open a study for every configuration.

Every rate is 95% or more
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI

2 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 15 per row

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 4× real timeMotion reduced: press Replay to animateThe slowest median is 5.7 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.7 s (range 4.1 s–25.5 s, n 15). Fastest Claude Haiku 4.5 · Claude Code 4.4 s (range 3.2 s–23.6 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 3.6× real timeMotion reduced: press Replay to animateThe slowest median is 5.1 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.1 s (range 3.4 s–17.8 s, n 15). Fastest Claude Haiku 4.5 · Claude Code 3.6 s (range 2.8 s–22.3 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

2 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (medium) · Codex CLI 5,180 (n 15). Lowest Claude Haiku 4.5 · Claude Code 0 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Haiku 4.5 · Claude Code 3,790 (n 15).

Notesn = 15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Haiku 4.5 or GPT-6.1 Sol (Codex CLI)?
Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) share 40 measured metrics and 16 list-price calculations from 7 studies. Claude Haiku 4.5 leads on 2 rows: Same prompt, 10 times: time per call (Exact number), 5.06 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 5.95 s vs 11.3 s. GPT-6.1 Sol (Codex CLI) leads on 6 rows: Pass rate on eight hard tasks (Strict pass), 100% (16/16) vs 46% (11/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); Same prompt, 10 times: strict pass rate (JSON object), 100% (10/10) vs 10% (1/10); and 3 more. On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 13 ties and 35 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
How were Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) measured?
They share 40 measured metrics and 16 list-price calculations from 7 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Prompt caching and run-to-run consistency in Claude Code and Codex CLI; Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
How do Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) compare on pass rate on eight hard tasks (Strict pass)?
Claude Haiku 4.5: 46% (11/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 28% to 65%). GPT-6.1 Sol (Codex CLI): 100% (16/16) (Codex CLI · effort medium · eight hard validated tasks; n = 16; 95% interval 81% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; GPT-6.1 Sol (Codex CLI) 81% to 100%).
How do Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) compare on same prompt, 10 times: strict pass rate (Exact number)?
Claude Haiku 4.5: 0% (0/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 0% to 28%). GPT-6.1 Sol (Codex CLI): 100% (10/10) (Codex CLI · effort medium · same prompt repeated 10 times; n = 10; 95% interval 72% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; GPT-6.1 Sol (Codex CLI) 72% to 100%).
How do Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) compare on same prompt, 10 times: strict pass rate (JSON object)?
Claude Haiku 4.5: 10% (1/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 1.8% to 40%). GPT-6.1 Sol (Codex CLI): 100% (10/10) (Codex CLI · effort medium · same prompt repeated 10 times; n = 10; 95% interval 72% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; GPT-6.1 Sol (Codex CLI) 72% to 100%).
How do Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) compare on same prompt, 10 times: time per call (Exact number)?
Claude Haiku 4.5: 5.06 s (Claude Code · same prompt repeated 10 times; n = 10; run range 4.4 s to 6.2 s). GPT-6.1 Sol (Codex CLI): 13.4 s (Codex CLI · effort medium · same prompt repeated 10 times; n = 10; run range 12.3 s to 18 s). The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.42 s to 6.20 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.
How do Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) compare on same prompt, 10 times: time per call (Code fix)?
Claude Haiku 4.5: 5.95 s (Claude Code · same prompt repeated 10 times; n = 10; run range 4.9 s to 7.3 s). GPT-6.1 Sol (Codex CLI): 11.3 s (Codex CLI · effort medium · same prompt repeated 10 times; n = 10; run range 9.1 s to 14.9 s). The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.