12 measured metrics · 11 calculated · 4 studies

Claude Fable 5.1vsGPT-6.1 Sol (Codex CLI)

No row separates them: 3 ties, 20 unclear.

The verdict

Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 4 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Time to first useful output, 4.2x (GPT-6.1 Sol (Codex CLI) larger).

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Fable 5.1
  • GPT-6.1 Sol (Codex CLI)
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

1 tie · 7 unclear
  • Pass rate on five validated tasks: Claude Fable 5.1 100% (15/15) (n 15, 95% interval 80%–100%); GPT-6.1 Sol (Codex CLI) 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
  • Total time per call: Claude Fable 5.1 1.94 s (n 15, run range 1.4 s–9.8 s); GPT-6.1 Sol (Codex CLI) 5.65 s (n 15, run range 4.1 s–25.5 s). Unclear.
  • Time to first useful output: Claude Fable 5.1 1.20 s (n 15, run range 1 s–7.9 s); GPT-6.1 Sol (Codex CLI) 5.05 s (n 15, run range 3.4 s–17.8 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): Claude Fable 5.1 2,760 (n 15); GPT-6.1 Sol (Codex CLI) 5,180 (n 15). Unclear.
  • Input tokens per call: what the CLI sends (Other input): Claude Fable 5.1 473 (n 15); GPT-6.1 Sol (Codex CLI) 6,943 (n 15). Unclear.
  • Output tokens per call (Output tokens): Claude Fable 5.1 64 (n 15); GPT-6.1 Sol (Codex CLI) 42 (n 15). Unclear.
  • List-price cost per call (calculation), calculation: Claude Fable 5.1 $0.0099 (n 15, run range $0.0049–$0.058); GPT-6.1 Sol (Codex CLI) $0.010 (n 15, run range $0.0054–$0.027). Unclear.
  • List-price cost per passing answer (calculation), calculation: Claude Fable 5.1 $0.021 (n 15); GPT-6.1 Sol (Codex CLI) $0.016 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

2 ties · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): Claude Fable 5.1 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): Claude Fable 5.1 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call on hard tasks (separate batches): Claude Fable 5.1 16.1 s (n 24, run range 4.5 s–90 s); GPT-6.1 Sol (Codex CLI) 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
  • Time to first useful output on hard tasks: Claude Fable 5.1 11.6 s (n 24, run range 2 s–85.3 s); GPT-6.1 Sol (Codex CLI) 10.2 s (n 16, run range 6.1 s–40.4 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): Claude Fable 5.1 1,366 (n 24); GPT-6.1 Sol (Codex CLI) 335 (n 16). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: Claude Fable 5.1 $0.093 (n 24); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

6 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Fable 5.1 64.2% (n 24, run range 23%–97%); GPT-6.1 Sol (Codex CLI) 46.3% (n 16, run range 11%–87%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Fable 5.1 $0.054 (n 24); GPT-6.1 Sol (Codex CLI) $0.0023 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Fable 5.1 $0.019 (n 24); GPT-6.1 Sol (Codex CLI) $0.0030 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Fable 5.1 $0.021 (n 24); GPT-6.1 Sol (Codex CLI) $0.020 (n 16). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Fable 5.1 64.2% (n 24, run range 23%–97%); GPT-6.1 Sol (Codex CLI) 46.3% (n 16, run range 11%–87%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Fable 5.1 0% (n 15, run range 0%–74%); GPT-6.1 Sol (Codex CLI) 41% (n 15, run range 0%–71%). Unclear.

Where the seconds go: first text, output speed and prompt size for 6 LLMs

3 unclear
  • Time to first text: a 250-line answer, six models: Claude Fable 5.1 4.43 s (n 4, run range 2.3 s–4.6 s); GPT-6.1 Sol (Codex CLI) 3.52 s (n 4, run range 2.8 s–4.4 s). Unclear.
  • Output speed after the first text: visible tokens per second (calculation), calculation: Claude Fable 5.1 123 (n 4, run range 121–131); GPT-6.1 Sol (Codex CLI) 80 (n 4, run range 72–81). Unclear.
  • Output speed in characters per second after the first text (calculation), calculation: Claude Fable 5.1 273 (n 4, run range 270–293); GPT-6.1 Sol (Codex CLI) 323 (n 4, run range 291–327). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

23 rows from 4 studies. No row separates them: 3 ties, 20 unclear.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Fable 5.1

No row in this data puts Claude Fable 5.1 ahead of GPT-6.1 Sol (Codex CLI). Pick on other grounds (price, access, the tasks you run), or measure your own workload.

When to pick GPT-6.1 Sol (Codex CLI)

  • List-price cost per strict pass on hard tasks (calculation): $0.026 vs $0.093. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0023 vs $0.054. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0030 vs $0.019. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only these two. Open a study for every configuration.

Every rate is 95% or more
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI

2 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 15 per row

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 4× real timeMotion reduced: press Replay to animateThe slowest median is 5.7 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.7 s (range 4.1 s–25.5 s, n 15). Fastest Claude Fable 5.1 · Claude Code 1.9 s (range 1.4 s–9.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 3.6× real timeMotion reduced: press Replay to animateThe slowest median is 5.1 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.1 s (range 3.4 s–17.8 s, n 15). Fastest Claude Fable 5.1 · Claude Code 1.2 s (range 1 s–7.9 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

2 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (medium) · Codex CLI 5,180 (n 15). Lowest Claude Fable 5.1 · Claude Code 2,760 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Fable 5.1 · Claude Code 473 (n 15).

Notesn = 15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Fable 5.1 or GPT-6.1 Sol (Codex CLI)?
Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 4 at the smallest).
How were Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) measured?
They share 12 measured metrics and 11 list-price calculations from 4 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs. Every row names its configuration, its sample size and its interval or range.
How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on pass rate on five validated tasks?
Claude Fable 5.1: 100% (15/15) (Claude Code · five short validated tasks; n = 15; 95% interval 80% to 100%). GPT-6.1 Sol (Codex CLI): 100% (15/15) (Codex CLI · effort medium · five short validated tasks; n = 15; 95% interval 80% to 100%). The 95% intervals overlap (Claude Fable 5.1 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.
How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on total time per call?
Claude Fable 5.1: 1.94 s (Claude Code · five short validated tasks; n = 15; run range 1.4 s to 9.8 s). GPT-6.1 Sol (Codex CLI): 5.65 s (Codex CLI · effort medium · five short validated tasks; n = 15; run range 4.1 s to 25.5 s). The run ranges (fastest to slowest) overlap (Claude Fable 5.1 1.41 s to 9.83 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on time to first useful output?
Claude Fable 5.1: 1.20 s (Claude Code · five short validated tasks; n = 15; run range 1 s to 7.9 s). GPT-6.1 Sol (Codex CLI): 5.05 s (Codex CLI · effort medium · five short validated tasks; n = 15; run range 3.4 s to 17.8 s). The run ranges (fastest to slowest) overlap (Claude Fable 5.1 0.95 s to 7.90 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on pass rate on eight hard tasks (Strict pass)?
Claude Fable 5.1: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). GPT-6.1 Sol (Codex CLI): 100% (16/16) (Codex CLI · effort medium · eight hard validated tasks; n = 16; 95% interval 81% to 100%). The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.
How do Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) compare on pass rate on eight hard tasks (Lenient (format misses counted))?
Claude Fable 5.1: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). GPT-6.1 Sol (Codex CLI): 100% (16/16) (Codex CLI · effort medium · eight hard validated tasks; n = 16; 95% interval 81% to 100%). The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.