• Head to head
  • Claude Haiku
  • Claude Sonnet
  • Claude Opus
  • Claude Fable
  • Codex
  • Latency

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

On short tasks with strict validators, how do the Claude Code models and efforts compare with Codex on pass rate, speed and tokens?

Published · 9 charts · Download the data or a carousel

98%

95% CI 93%–99% · n = 130

127/130 · Calls that passed their validator

The answer

127 of 130 calls passed (98%, 95% interval 93% to 99%), so on these short tasks pass rate barely separates the configurations (non-passes: Claude Sonnet 5.5 · Claude Code on Multi-step shift arithmetic ×3 ("Output does not exactly match the expected text")). Every non-pass ended on the expected answer (2292) but added working lines, which the exact-text validator rejects by design: a format miss, not a wrong answer. Speed separates them more: the fastest configuration was Claude Fable 5.1 · Claude Code at a median 1.9 s per call, the slowest GPT-6.1 Sol (low) · Codex CLI at 6.3 s. Claude Code medians ran from 1.9 s to 4.4 s. Codex CLI medians ran from 5.6 s to 6.3 s. Single calls vary a lot (see the ranges), so neighbouring configurations are not separated. Codex CLI sends a median 12,124 input tokens per call against 2,130 for Claude Code, mostly the CLI’s own context. At list price (a calculation; the calls ran on subscriptions), the cheapest passing answer came from Claude Sonnet 5.5 · Claude Code at $0.0062. On speed, cost per pass and pass rate together, the frontier is Claude Fable 5.1 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (low) · Claude Code; with 10 to 15 calls per configuration, small gaps inside it are within the spread of single calls. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.

Where each call’s input tokens come from. Pick up to three configurations to compare.

Claude Opus 5.5 (low) · Claude Code2,081 tokens per call

1,463 + 618 = 2,081

GPT-6.1 Sol (medium) · Codex CLI12,123 tokens per call

5,180 + 6,943 = 12,123

Ribbon thickness is tokens per call, on one scale for every configuration. Only parts that sum to the total are drawn. First shown: the largest and the smallest total.

9 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (low) · Codex CLI 8,064 (n 10). Lowest Claude Haiku 4.5 · Claude Code 0 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Fable 5.1 · Claude Code 473 (n 15).

Notesn 10–15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Share card (PNG)

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 35 sHaiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

127 of 130 calls passed, so speed and tokens separate the models: Fable 5.1 was fastest at 1.9 s median.

Transcript
  1. Head-to-head · 130 timed calls · 9 configurations. Haiku vs Sonnet vs Opus vs Fable vs Codex. Five short tasks with strict validators. Every call kept, nothing retried.
  2. 127 of 130 calls passed. Pass rate barely separates them; speed and tokens do. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Cheapest passing answer (list-price calculation): Sonnet 5.5: $0.0062 (n = 15). Caveat: The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.
  3. Fable 5.1 finishes first at 1.9 s. The Codex CLI needs 5.6–6.3 s. Chart: Median total time per call · real time (n = 10–15 each). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
  4. What the CLI sends: a median 12,124 input tokens per call on the Codex CLI, 2,130 on Claude Code. Chart: Input tokens per call: what the CLI sends (n = 10–15 each). Caveat: The prompt cache stayed at the provider default, so cache counters differ by route and by call order.
  5. Speed vs cost per passing answer. Ringed: no other setup is faster, cheaper per pass and as accurate. Chart: Speed, cost and quality frontier (n = 10–15 each). Calculation, not a run. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  6. Open benchmarks: intervals, sources and every failure kept.

Key numbers

9

Configurations compared

1.9s

Claude Fable 5.1 · Claude Code

Fastest configuration (median total time)

n = 15

6.3s

GPT-6.1 Sol (low) · Codex CLI

Slowest configuration (median total time)

n = 10

$0.0062

Claude Sonnet 5.5 · Claude Code

Lowest list-price cost per passing answer (calculation)

n = 15

12,124

Median input tokens per call, Codex CLI vs Claude Code

vs 2,130 · n = 130

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Claude Fable 5.1
Claude Sonnet 5.5
Claude Opus 5.5 (high)
Claude Opus 5.5
Claude Opus 5.5 (low)
Claude Haiku 4.5
GPT-6.1 Sol (high)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (low)

Every interval overlaps every other: this chart does not order these rows.

9 rows. Highest Claude Fable 5.1 · Claude Code 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row8 of 9 at 100%: this task set cannot separate them.

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Share card (PNG)
Entrance: medians race at 4.5× real timeMotion reduced: press Replay to animateThe slowest median is 6.3 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

9 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 6.3 s (range 4.7 s–10.5 s, n 10). Fastest Claude Fable 5.1 · Claude Code 1.9 s (range 1.4 s–9.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Share card (PNG)
Entrance: medians race at 3.8× real timeMotion reduced: press Replay to animateThe slowest median is 5.3 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

9 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 5.3 s (range 3.6 s–16.4 s, n 15). Fastest Claude Fable 5.1 · Claude Code 1.2 s (range 1 s–7.9 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Share card (PNG)
  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

9 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 367 (n 15). Lowest GPT-6.1 Sol (low) · Codex CLI 42 (n 10). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 297 (n 15). Lowest Claude Opus 5.5 (low) · Claude Code 0 (n 15).

Notesn 10–15 per row

Median per configuration; reasoning tokens where the CLI reports them

Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Share card (PNG)
Calculation
Claude Fable 5.1
Claude Sonnet 5.5
Claude Opus 5.5 (high)
Claude Opus 5.5
Claude Opus 5.5 (low)
Claude Haiku 4.5
GPT-6.1 Sol (high)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (low)

Every interval overlaps every other: this chart does not order these rows.

List-price calculation, not a run. 9 rows. Highest GPT-6.1 Sol (high) · Codex CLI $0.01 (range $0.0066–$0.028, n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0036 (range $0.0034–$0.01, n 15). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 10–15 per row

Reported tokens × list price; the calls ran on subscriptions

Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.

Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
Better: lower left

List-price calculation, not a run. 9 points: Median total seconds against USD per call (list-price calculation). USD per call (list-price calculation) runs from $0.0036 to $0.01; Median total seconds from 1.9 s to 6.3 s. Highlighted: Claude Fable 5.1 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (low) · Claude Code, Claude Haiku 4.5 · Claude Code.

Notesn 10–15 per point

Median total time and median list-price cost per call

Lower-left is faster and cheaper. Cost is a calculation from tokens.

Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 9 rows. Highest Claude Fable 5.1 · Claude Code $0.021 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0062 (n 15).

Notesn 10–15 per row

All calls in a configuration, failures included, divided by its passes

Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.

Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
  • Claude Code
  • Codex CLI
Better: lower left

Haloed: on the frontier (5 of 9). The frontier also uses a measure this plot does not show, so no line is drawn.

List-price calculation, not a run. 9 points: Median total seconds against USD per passing answer (list-price calculation). USD per passing answer (list-price calculation) runs from $0.0062 to $0.021; Median total seconds from 1.9 s to 6.3 s. Highlighted: Claude Fable 5.1 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (low) · Claude Code.

Notesn 10–15 per point

Median seconds against list-price cost per passing answer; pass rate in the note

Lower-left is better. Highlighted points are on the frontier: no other configuration is at least as fast, as cheap per pass and as accurate. Frontier: Claude Fable 5.1 · Claude Code; Claude Sonnet 5.5 · Claude Code; Claude Opus 5.5 (high) · Claude Code; Claude Opus 5.5 · Claude Code; Claude Opus 5.5 (low) · Claude Code. Pass rates: Claude Fable 5.1 · Claude Code 15/15; Claude Sonnet 5.5 · Claude Code 12/15; Claude Opus 5.5 (high) · Claude Code 15/15; Claude Opus 5.5 · Claude Code 15/15; Claude Opus 5.5 (low) · Claude Code 15/15; Claude Haiku 4.5 · Claude Code 15/15; GPT-6.1 Sol (high) · Codex CLI 15/15; GPT-6.1 Sol (medium) · Codex CLI 15/15; GPT-6.1 Sol (low) · Codex CLI 10/10. Costs are calculations from tokens; medians come from small samples.

Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)

Tables

Pass matrix: configuration × task

Configuration by column: each cell is k of n from the table; a cell opens its model

ConfigurationFix a buggy median function1Extract invoice fields to JSON2Multi-step shift arithmetic3Refactor recursion to iteration4Classify six support tickets5
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI
  1. Fix a buggy median function
  2. Extract invoice fields to JSON
  3. Multi-step shift arithmetic
  4. Refactor recursion to iteration
  5. Classify six support tickets

Shade is the share k of n in each cell. Totals and other columns are in the Table view.

Counts from the table, not an intervaln = 2–3 per cell

9 rows by 5 columns. 44 of 45 k/n cells are full (n of n) and 1 are zero.

Method

  1. Protocol declared before the first call: 5 cases with deterministic validators (behavioral checks in a sandbox, canonical JSON or exact text).
  2. Claude Code: Haiku, Sonnet, Opus and Fable at default effort, plus Opus at low and high, 3 repetitions. Codex: GPT-6.1 Sol at low, medium and high, 2 repetitions.
  3. Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn. One call at a time per account.
  4. Every attempt is kept. Nothing is retried. A cell that did not run is shown as trimmed.
  5. All 130 planned calls ran. Nothing was trimmed or retried, and no run hit a usage or rate limit.
  6. Cost per passing answer: list price × reported tokens for every call in the configuration (cache reads and writes priced separately), divided by its passes.

Caveats

  • The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.
  • Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  • CLI timings include CLI start-up and the CLI’s own system prompt.
  • One host, one network, one day.
  • List-price costs are calculations; the calls used flat subscriptions.
  • The prompt cache stayed at the provider default, so cache counters differ by route and by call order.
  • Haiku 4.5 reported reasoning tokens on most calls under the CLI default, which explains much of its extra time and output.
  • Haiku and Fable are not in the platform runner catalog, so these calls used the runner’s CLI functions directly; each receipt records the model the CLI reported.

Sources

  • Provider head-to-head: Claude Code models vs Codex efforts

    Our recorded runs ·

    Five short tasks with deterministic validators, declared protocol, every attempt kept.

    Raw data: provider-h2h/receipts.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head”, updated October 5, 2026, https://agent.sasid.ai/benchmarks/model-head-to-head.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

More comparisons based on this study (9)

These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.

Models and comparisons in this study

More studies

All benchmarks
Live story
  • Head to head
  • Hard Tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.

91% (139/152)Calls that passed strictly (hard set) · n = 152

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.