12 measured metrics · 11 calculated · 4 studies
Claude Fable 5.1vsGPT-6.1 Sol (Codex CLI)
No row separates them: 3 ties, 20 unclear.
The verdict
Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 4 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Claude Fable 5.1 | GPT-6.1 Sol (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Claude Code · five short validated tasks | 100% (15/15)Codex CLI · effort medium · five short validated tasks | 15 | 95% CI: 80%–100% vs 80%–100% | Tie | The 95% intervals overlap (Claude Fable 5.1 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 1.94 sClaude Code · five short validated tasks | 5.65 sCodex CLI · effort medium · five short validated tasks | 15 | range: 1.4 s–9.8 s vs 4.1 s–25.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 1.41 s to 9.83 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 1.20 sClaude Code · five short validated tasks | 5.05 sCodex CLI · effort medium · five short validated tasks | 15 | range: 1 s–7.9 s vs 3.4 s–17.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 0.95 s to 7.90 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 2,760Claude Code · five short validated tasks | 5,180Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 64Claude Code · five short validated tasks | 42Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.021Claude Code · five short validated tasks | $0.016Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.021 vs $0.016) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Time to first useful output, 4.2x (GPT-6.1 Sol (Codex CLI) larger).
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Claude Fable 5.1
- GPT-6.1 Sol (Codex CLI)
- 95% interval
- fastest–slowest run (not an interval)
- where the two overlap
- hollow: list-price calculation
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
- Pass rate on five validated tasks100% (15/15)n 15100% (15/15)n 15TiePass rate on five validated tasks: Claude Fable 5.1 100% (15/15) (n 15, 95% interval 80%–100%); GPT-6.1 Sol (Codex CLI) 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
- Total time per call1.94 sn 155.65 sn 15UnclearTotal time per call: Claude Fable 5.1 1.94 s (n 15, run range 1.4 s–9.8 s); GPT-6.1 Sol (Codex CLI) 5.65 s (n 15, run range 4.1 s–25.5 s). Unclear.
- Time to first useful output1.20 sn 155.05 sn 15UnclearTime to first useful output: Claude Fable 5.1 1.20 s (n 15, run range 1 s–7.9 s); GPT-6.1 Sol (Codex CLI) 5.05 s (n 15, run range 3.4 s–17.8 s). Unclear.
- Input tokens per call: what the CLI sends (Cache read)2,760n 155,180n 15UnclearInput tokens per call: what the CLI sends (Cache read): Claude Fable 5.1 2,760 (n 15); GPT-6.1 Sol (Codex CLI) 5,180 (n 15). Unclear.
- Input tokens per call: what the CLI sends (Other input)473n 156,943n 15UnclearInput tokens per call: what the CLI sends (Other input): Claude Fable 5.1 473 (n 15); GPT-6.1 Sol (Codex CLI) 6,943 (n 15). Unclear.
- Output tokens per call (Output tokens)64n 1542n 15UnclearOutput tokens per call (Output tokens): Claude Fable 5.1 64 (n 15); GPT-6.1 Sol (Codex CLI) 42 (n 15). Unclear.
- List-price cost per call (calculation)Calculation$0.0099n 15$0.010n 15UnclearList-price cost per call (calculation), calculation: Claude Fable 5.1 $0.0099 (n 15, run range $0.0049–$0.058); GPT-6.1 Sol (Codex CLI) $0.010 (n 15, run range $0.0054–$0.027). Unclear.
- List-price cost per passing answer (calculation)Calculation$0.021n 15$0.016n 15UnclearList-price cost per passing answer (calculation), calculation: Claude Fable 5.1 $0.021 (n 15); GPT-6.1 Sol (Codex CLI) $0.016 (n 15). Unclear.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
- Pass rate on eight hard tasks (Strict pass)100% (24/24)n 24100% (16/16)n 16TiePass rate on eight hard tasks (Strict pass): Claude Fable 5.1 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Pass rate on eight hard tasks (Lenient (format misses counted))100% (24/24)n 24100% (16/16)n 16TiePass rate on eight hard tasks (Lenient (format misses counted)): Claude Fable 5.1 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Total time per call on hard tasks (separate batches)16.1 sn 2413.1 sn 16UnclearTotal time per call on hard tasks (separate batches): Claude Fable 5.1 16.1 s (n 24, run range 4.5 s–90 s); GPT-6.1 Sol (Codex CLI) 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
- Time to first useful output on hard tasks11.6 sn 2410.2 sn 16UnclearTime to first useful output on hard tasks: Claude Fable 5.1 11.6 s (n 24, run range 2 s–85.3 s); GPT-6.1 Sol (Codex CLI) 10.2 s (n 16, run range 6.1 s–40.4 s). Unclear.
- Output tokens per call on hard tasks (Output tokens)1,366n 24335n 16UnclearOutput tokens per call on hard tasks (Output tokens): Claude Fable 5.1 1,366 (n 24); GPT-6.1 Sol (Codex CLI) 335 (n 16). Unclear.
- List-price cost per strict pass on hard tasks (calculation)Calculation$0.093n 24$0.026n 16UnclearList-price cost per strict pass on hard tasks (calculation), calculation: Claude Fable 5.1 $0.093 (n 24); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.
How much of an AI bill is thinking? Reasoning tokens by model and effort
- Reasoning share of output tokens per call on hard tasks (calculation)Calculation64.2%n 2446.3%n 16UnclearReasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Fable 5.1 64.2% (n 24, run range 23%–97%); GPT-6.1 Sol (Codex CLI) 46.3% (n 16, run range 11%–87%). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.054n 24$0.0023n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Fable 5.1 $0.054 (n 24); GPT-6.1 Sol (Codex CLI) $0.0023 (n 16). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.019n 24$0.0030n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Fable 5.1 $0.019 (n 24); GPT-6.1 Sol (Codex CLI) $0.0030 (n 16). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.021n 24$0.020n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Fable 5.1 $0.021 (n 24); GPT-6.1 Sol (Codex CLI) $0.020 (n 16). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation64.2%n 2446.3%n 16UnclearReasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Fable 5.1 64.2% (n 24, run range 23%–97%); GPT-6.1 Sol (Codex CLI) 46.3% (n 16, run range 11%–87%). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation0%n 1541%n 15UnclearReasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Fable 5.1 0% (n 15, run range 0%–74%); GPT-6.1 Sol (Codex CLI) 41% (n 15, run range 0%–71%). Unclear.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
- Time to first text: a 250-line answer, six models4.43 sn 43.52 sn 4UnclearTime to first text: a 250-line answer, six models: Claude Fable 5.1 4.43 s (n 4, run range 2.3 s–4.6 s); GPT-6.1 Sol (Codex CLI) 3.52 s (n 4, run range 2.8 s–4.4 s). Unclear.
- Output speed after the first text: visible tokens per second (calculation)Calculation123n 480n 4UnclearOutput speed after the first text: visible tokens per second (calculation), calculation: Claude Fable 5.1 123 (n 4, run range 121–131); GPT-6.1 Sol (Codex CLI) 80 (n 4, run range 72–81). Unclear.
- Output speed in characters per second after the first text (calculation)Calculation273n 4323n 4UnclearOutput speed in characters per second after the first text (calculation), calculation: Claude Fable 5.1 273 (n 4, run range 270–293); GPT-6.1 Sol (Codex CLI) 323 (n 4, run range 291–327). Unclear.
| Metric | Claude Fable 5.1 | GPT-6.1 Sol (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Claude Code · five short validated tasks | 100% (15/15)Codex CLI · effort medium · five short validated tasks | 15 | 95% CI: 80%–100% vs 80%–100% | Tie | The 95% intervals overlap (Claude Fable 5.1 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 1.94 sClaude Code · five short validated tasks | 5.65 sCodex CLI · effort medium · five short validated tasks | 15 | range: 1.4 s–9.8 s vs 4.1 s–25.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 1.41 s to 9.83 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 1.20 sClaude Code · five short validated tasks | 5.05 sCodex CLI · effort medium · five short validated tasks | 15 | range: 1 s–7.9 s vs 3.4 s–17.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 0.95 s to 7.90 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 2,760Claude Code · five short validated tasks | 5,180Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 473Claude Code · five short validated tasks | 6,943Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 64Claude Code · five short validated tasks | 42Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0099Claude Code · five short validated tasks | $0.010Codex CLI · effort medium · five short validated tasks | 15 | range: $0.0049–$0.058 vs $0.0054–$0.027 | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 $0.0049 to $0.058; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.021Claude Code · five short validated tasks | $0.016Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.021 vs $0.016) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24)Claude Code · eight hard validated tasks | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24)Claude Code · eight hard validated tasks | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 16.1 sClaude Code · eight hard validated tasks | 13.1 sCodex CLI · effort medium · eight hard validated tasks | 24 / 16 | range: 4.5 s–90 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 4.46 s to 90.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 11.6 sClaude Code · eight hard validated tasks | 10.2 sCodex CLI · effort medium · eight hard validated tasks | 24 / 16 | range: 2 s–85.3 s vs 6.1 s–40.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.00 s to 85.3 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,366Claude Code · eight hard validated tasks | 335Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.093Claude Code · eight hard validated tasks | $0.026Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.093 vs $0.026, 3.6x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 64.2%Claude Code | 46.3%Codex CLI · effort medium | 24 / 16 | range: 23%–97% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.054Claude Code | $0.0023Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.019Claude Code | $0.0030Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.021Claude Code | $0.020Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 64.2%Claude Code | 46.3%Codex CLI · effort medium | 24 / 16 | range: 23%–97% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 0%Claude Code | 41%Codex CLI · effort medium | 15 | range: 0%–74% vs 0%–71% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 4.43 sClaude Code | 3.52 sCodex CLI · effort low | 4 | range: 2.3 s–4.6 s vs 2.8 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.27 s to 4.64 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 123Claude Code | 80Codex CLI · effort low | 4 | range: 121–131 vs 72–81 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 273Claude Code | 323Codex CLI · effort low | 4 | range: 270–293 vs 291–327 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
23 rows from 4 studies. No row separates them: 3 ties, 20 unclear.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Claude Fable 5.1
No row in this data puts Claude Fable 5.1 ahead of GPT-6.1 Sol (Codex CLI). Pick on other grounds (price, access, the tasks you run), or measure your own workload.
When to pick GPT-6.1 Sol (Codex CLI)
- List-price cost per strict pass on hard tasks (calculation): $0.026 vs $0.093. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0023 vs $0.054. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0030 vs $0.019. A list-price calculation, not a measured difference. Calculation
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Pass rate | 95% interval | n |
|---|---|---|---|
| Claude Fable 5.1 · Claude Code | 100% | 80%–100% | 15 |
| GPT-6.1 Sol (high) · Codex CLI | 100% | 80%–100% | 15 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 15 per row
Every call counts; failures and timeouts are non-passes
Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Fable 5.1 · Claude Code | 1.9 s | 1.4 s–9.8 s | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 5.7 s | 4.1 s–25.5 s | 15 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.7 s (range 4.1 s–25.5 s, n 15). Fastest Claude Fable 5.1 · Claude Code 1.9 s (range 1.4 s–9.8 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Fable 5.1 · Claude Code | 1.2 s | 1 s–7.9 s | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 5.1 s | 3.4 s–17.8 s | 15 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.1 s (range 3.4 s–17.8 s, n 15). Fastest Claude Fable 5.1 · Claude Code 1.2 s (range 1 s–7.9 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Cache read
- Other input
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Cache read | Other input | n |
|---|---|---|---|
| Claude Fable 5.1 · Claude Code | 2,760 | 473 | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 5,180 | 6,943 | 15 |
2 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (medium) · Codex CLI 5,180 (n 15). Lowest Claude Fable 5.1 · Claude Code 2,760 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Fable 5.1 · Claude Code 473 (n 15).
Notesn = 15 per row
Mean per call, split into prompt-cache reads and other input
The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Fable 5.1 · Claude Code | 64 | 0 | 15 |
| GPT-6.1 Sol (high) · Codex CLI | 42 | 21 | 15 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Fable 5.1 · Claude Code 64 (n 15). Lowest GPT-6.1 Sol (high) · Codex CLI 42 (n 15). Reasoning tokens: highest GPT-6.1 Sol (high) · Codex CLI 21 (n 15). Lowest Claude Fable 5.1 · Claude Code 0 (n 15).
Notesn = 15 per row
Median per configuration; reasoning tokens where the CLI reports them
Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.
Source: Provider head-to-head: Claude Code models vs Codex efforts
| Item | Cost per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Fable 5.1 · Claude Code | $0.0099 | $0.0049–$0.058 | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.01 | $0.0054–$0.027 | 15 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (medium) · Codex CLI $0.01 (range $0.0054–$0.027, n 15). Lowest Claude Fable 5.1 · Claude Code $0.0099 (range $0.0049–$0.058, n 15). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 15 per row
Reported tokens × list price; the calls ran on subscriptions
Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Cost per pass | n |
|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | $0.016 | 15 |
| Claude Fable 5.1 · Claude Code | $0.021 | 15 |
List-price calculation, not a run. 2 rows. Highest Claude Fable 5.1 · Claude Code $0.021 (n 15). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.016 (n 15).
Notesn = 15 per row
All calls in a configuration, failures included, divided by its passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 100% | Strict pass: 81%–100%; Lenient (format misses counted): 81%–100% | 16 |
| Claude Fable 5.1 · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: all at 100%. Lenient (format misses counted): all at 100%.
NotesWhiskers: 95% Wilson intervaln 16–24 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call on hard tasks (separate batches) | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 8.5 s–61.6 s | 16 |
| Claude Fable 5.1 · Claude Code | 16.1 s | 4.5 s–90 s | 24 |
2 rows. Slowest Claude Fable 5.1 · Claude Code 16.1 s (range 4.5 s–90 s, n 24). Fastest GPT-6.1 Sol (medium) · Codex CLI 13.1 s (range 8.5 s–61.6 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 10.2 s | 6.1 s–40.4 s | 16 |
| Claude Fable 5.1 · Claude Code | 11.6 s | 2 s–85.3 s | 24 |
2 rows. Slowest Claude Fable 5.1 · Claude Code 11.6 s (range 2 s–85.3 s, n 24). Fastest GPT-6.1 Sol (medium) · Codex CLI 10.2 s (range 6.1 s–40.4 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 335 | 150 | 16 |
| Claude Fable 5.1 · Claude Code | 1,366 | 889 | 24 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Fable 5.1 · Claude Code 1,366 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Fable 5.1 · Claude Code 889 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).
Notesn 16–24 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Cost per strict pass | n |
|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | 16 |
| Claude Fable 5.1 · Claude Code | $0.093 | 24 |
List-price calculation, not a run. 2 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.026 (n 16).
Notesn 16–24 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Reasoning (output tokens)
- Remaining output (visible-answer estimate)
- Input (prompt, cache priced)
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Reasoning (output tokens) | Remaining output (visible-answer estimate) | Input (prompt, cache priced) | n |
|---|---|---|---|---|
| Claude Fable 5.1 · Claude Code | $0.054 | $0.019 | $0.021 | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.0023 | $0.003 | $0.02 | 16 |
List-price calculation, not a run. 2 rows, 3 series: Reasoning (output tokens), Remaining output (visible-answer estimate), Input (prompt, cache priced). Reasoning (output tokens): highest Claude Fable 5.1 · Claude Code $0.054 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.0023 (n 16). Remaining output (visible-answer estimate): highest Claude Fable 5.1 · Claude Code $0.019 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.003 (n 16).
Notesn 16–24 per row
Mean per call on the hard tasks; the three parts add up to the call
Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Eight hard tasks
- Five short tasks (square)
Gap labels, Five short tasks vs Eight hard tasks: Five short tasks is x percentage points higher (+) or lower (−) than Eight hard tasks, calculated from the two values shown; lines are the lowest–highest run (not an interval).
| Item | Eight hard tasks | Five short tasks | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Fable 5.1 · Claude Code | 64% | 0% | Eight hard tasks: 23%–97%; Five short tasks: 0%–74% | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 46% | 41% | Eight hard tasks: 11%–87%; Five short tasks: 0%–71% | 16 |
List-price calculation, not a run. 2 rows, 2 series: Eight hard tasks, Five short tasks. Eight hard tasks: highest Claude Fable 5.1 · Claude Code 64% (range 23%–97%, n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 46% (range 11%–87%, n 16). All run ranges overlap. Five short tasks: highest GPT-6.1 Sol (medium) · Codex CLI 41% (range 0%–71%, n 15). Lowest Claude Fable 5.1 · Claude Code 0% (range 0%–74%, n 15). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 15–24 per row
Median call per configuration; five short tasks and eight hard tasks
Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Fable 5.1 · Claude Code | 4.4 s | 2.3 s–4.6 s | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 3.5 s | 2.8 s–4.4 s | 4 |
2 rows. Slowest Claude Fable 5.1 · Claude Code 4.4 s (range 2.3 s–4.6 s, n 4). Fastest GPT-6.1 Sol (low) · Codex CLI 3.5 s (range 2.8 s–4.4 s, n 4). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
| Item | Visible tokens per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Fable 5.1 · Claude Code | 123 | 121–131 | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 80 | 72–81 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Fable 5.1 · Claude Code 123 (range 121–131, n 4). Lowest GPT-6.1 Sol (low) · Codex CLI 80 (range 72–81, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = slowest and fastest call
Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.
Source: LLM speed anatomy
| Item | Characters per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Fable 5.1 · Claude Code | 273 | 270–293 | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 323 | 291–327 | 4 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (low) · Codex CLI 323 (range 291–327, n 4). Lowest Claude Fable 5.1 · Claude Code 273 (range 270–293, n 4). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call
Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Claude Fable 5.1 or GPT-6.1 Sol (Codex CLI)?
- Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 4 at the smallest).
- How were Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) measured?
- They share 12 measured metrics and 11 list-price calculations from 4 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs. Every row names its configuration, its sample size and its interval or range.
- How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on pass rate on five validated tasks?
- Claude Fable 5.1: 100% (15/15) (Claude Code · five short validated tasks; n = 15; 95% interval 80% to 100%). GPT-6.1 Sol (Codex CLI): 100% (15/15) (Codex CLI · effort medium · five short validated tasks; n = 15; 95% interval 80% to 100%). The 95% intervals overlap (Claude Fable 5.1 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.
- How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on total time per call?
- Claude Fable 5.1: 1.94 s (Claude Code · five short validated tasks; n = 15; run range 1.4 s to 9.8 s). GPT-6.1 Sol (Codex CLI): 5.65 s (Codex CLI · effort medium · five short validated tasks; n = 15; run range 4.1 s to 25.5 s). The run ranges (fastest to slowest) overlap (Claude Fable 5.1 1.41 s to 9.83 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on time to first useful output?
- Claude Fable 5.1: 1.20 s (Claude Code · five short validated tasks; n = 15; run range 1 s to 7.9 s). GPT-6.1 Sol (Codex CLI): 5.05 s (Codex CLI · effort medium · five short validated tasks; n = 15; run range 3.4 s to 17.8 s). The run ranges (fastest to slowest) overlap (Claude Fable 5.1 0.95 s to 7.90 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on pass rate on eight hard tasks (Strict pass)?
- Claude Fable 5.1: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). GPT-6.1 Sol (Codex CLI): 100% (16/16) (Codex CLI · effort medium · eight hard validated tasks; n = 16; 95% interval 81% to 100%). The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.
- How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on pass rate on eight hard tasks (Lenient (format misses counted))?
- Claude Fable 5.1: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). GPT-6.1 Sol (Codex CLI): 100% (16/16) (Codex CLI · effort medium · eight hard validated tasks; n = 16; 95% interval 81% to 100%). The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.
The studies behind this page
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.