• Head to head
  • Claude Fable
  • Claude Opus
  • Claude Sonnet

Claude Fable 5.1 vs Opus 5.5 vs Sonnet 5.5: speed, tokens and price tested

24 of 24: Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed every hard task. Fable cost 3.3x Opus and 6.5x Sonnet per pass (list-price calculation).

TL;DR

  • Quality is a tie at the ceiling. On eight hard tasks, Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed 24 of 24 (95% interval 86% to 100%). On five short tasks, Fable and Opus passed 15 of 15 and Sonnet 12 of 15, with intervals that overlap. The two comparison pages hold 28 rows: 7 ties, 21 unclear and no winner.
  • The speed order changes with task length, and every range overlaps. Median time per short call: Fable 1.94 s, Sonnet 2.31 s, Opus 2.75 s. Per hard call: Fable 16.13 s, Sonnet 7.75 s, Opus 9.18 s. Neither order is a tested ranking.
  • Fable read more input and wrote more output. It read 3,233 input tokens per short call, against 2,086 for Sonnet and 2,081 for Opus (a calculation). On hard tasks it wrote a median 1,366 output tokens, against 1,050 and 945.
  • Fable costs the most. It lists at $10 input and $50 output per million tokens: 2.5x Opus and 5x Sonnet. Per hard pass, Fable cost $0.0933, Opus $0.0282 and Sonnet $0.0143 (calculations).
  • When is Fable worth it? Only when your own validator shows a gap. Ours show none.

The side-by-side rows: Opus 5.5 vs Fable 5.1 and Sonnet 5.5 vs Fable 5.1. Model pages: Fable 5.1 and Opus 5.5.

Live story · 35 sHaiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

127 of 130 calls passed, so speed and tokens separate the models: Fable 5.1 was fastest at 1.9 s median.

Transcript
  1. Head-to-head · 130 timed calls · 9 configurations. Haiku vs Sonnet vs Opus vs Fable vs Codex. Five short tasks with strict validators. Every call kept, nothing retried.
  2. 127 of 130 calls passed. Pass rate barely separates them; speed and tokens do. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Cheapest passing answer (list-price calculation): Sonnet 5.5: $0.0062 (n = 15). Caveat: The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.
  3. Fable 5.1 finishes first at 1.9 s. The Codex CLI needs 5.6–6.3 s. Chart: Median total time per call · real time (n = 10–15 each). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
  4. What the CLI sends: a median 12,124 input tokens per call on the Codex CLI, 2,130 on Claude Code. Chart: Input tokens per call: what the CLI sends (n = 10–15 each). Caveat: The prompt cache stayed at the provider default, so cache counters differ by route and by call order.
  5. Speed vs cost per passing answer. Ringed: no other setup is faster, cheaper per pass and as accurate. Chart: Speed, cost and quality frontier (n = 10–15 each). Calculation, not a run. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  6. Open benchmarks: intervals, sources and every failure kept.

The story shows medians. Single-call ranges overlap, so read it as a description of these runs, not a ranking.

Fable 5.1 is the highest-priced Claude model in our studies. This post compares Claude Fable vs Opus and Claude Fable vs Sonnet on speed, tokens and price. We ran all three through Claude Code on two task sets.

Quality: a tie at the ceiling

On the hard set, all three models passed 24 of 24 calls strictly. A perfect 24 of 24 still has a 95% interval of 86% to 100%, so the three sit in one band. On the short set, Fable and Opus each passed 15 of 15 (80% to 100%) and Sonnet 12 of 15 (55% to 93%). Sonnet's three misses were format misses: the right number, plus extra lines. This is a ceiling, not proof of equality: our tasks cannot tell the three apart.

Speed on short tasks: the lowest median, not a proven lead

Entrance: medians race at 4.5× real timeMotion reduced: press Replay to animateThe slowest median is 6.3 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

9 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 6.3 s (range 4.7 s–10.5 s, n 10). Fastest Claude Fable 5.1 · Claude Code 1.9 s (range 1.4 s–9.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

On five short tasks, Fable's median total time per call was 1.94 s (range 1.41 s to 9.83 s, n = 15). Sonnet took 2.31 s (2.17 s to 7.73 s) and Opus 2.75 s (2.47 s to 8.91 s). Time to first useful output ran in the same order: Fable 1.20 s, Sonnet 1.56 s, Opus 1.92 s. The ranges overlap, so no model is ahead: Fable's median is the lowest, and that is all these runs show.

Speed on hard tasks: the medians swap places

Entrance: medians race at 28× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

7 rows. Slowest Claude Haiku 4.5 · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 16–24 per row

Median per configuration; whiskers = fastest and slowest call

One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

On eight hard tasks, Fable's median rose to 16.13 s (range 4.46 s to 90 s, n = 24). Sonnet took 7.75 s (2.26 s to 34.79 s) and Opus 9.18 s (4.24 s to 27.21 s). Fable's median was 2.1 times Sonnet's and 1.8 times Opus's (a calculation). The ranges still overlap, so this is not a tested ranking. The lowest median on short work became the highest of the three on hard work.

Tokens: what Fable reads and writes

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

9 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (low) · Codex CLI 8,064 (n 10). Lowest Claude Haiku 4.5 · Claude Code 0 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Fable 5.1 · Claude Code 473 (n 15).

Notesn 10–15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

On short calls, Fable read a mean 2,760 cache tokens and 473 other input tokens. Sonnet read 1,401 and 685, and Opus 1,401 and 680. Fable's total of 3,233 is about 1.55 times the 2,086 and 2,081 of the others (a calculation). Most input is the CLI's own system prompt and tool context, and we did not check why Fable's is larger. On hard tasks, Fable wrote a median 1,366 output tokens (889 of them reasoning); Sonnet wrote 1,050 (585) and Opus 945 (529).

Price per token: 2.5x to 5x

ModelInput per MCache read per MOutput per M
Claude Fable 5.1$10$0.25$50
Claude Opus 5.5$4under re-check$20
Claude Sonnet 5.5$2$0.20$10
Reported
  1. Inputall 4 at $10.00

  2. Outputall 4 at $50.00

  3. Cache readall 4 at $0.25

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 4 rows, 3 series: Input, Output, Cache read. Input: all at $10.00. Output: all at $50.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Fable's input and output prices are 2.5 times Opus's and 5 times Sonnet's (a calculation). OpenRouter's public API showed one Fable price at all four providers in our snapshot: Amazon Bedrock, Anthropic, Azure and Google Vertex (third-party-reported, 2026-10-06). OpenRouter's per-token price matches Anthropic's list, and it adds a 5.5% fee when you buy credits on its Standard plan.

Price per passing answer

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

We priced every reported token at list price, then divided each model's total by its passes, failures included. These are calculations, not bills: the calls ran on a subscription. On the hard set, Fable cost $0.0933 per strict pass, Opus $0.0282 and Sonnet $0.0143, which is 3.3 times Opus and 6.5 times Sonnet. On the short set, a passing answer cost Fable $0.0205, Opus $0.0101 and Sonnet $0.0062, or 2.0 and 3.3 times. Two effects stack: a higher price per token, and 1.3 times Sonnet's and 1.4 times Opus's output on the hard set.

Price in a long agent loop

We repriced the tokens of our agent's 33 SWE-bench instances (25 resolved): 162.9M input tokens (94.0% cache reads) and 1.8M output tokens. At list price, Fable comes to $321.31, Opus to $143.83 and Sonnet to $87.23, or $12.85, $5.75 and $3.49 per resolved instance. Fable is 3.7 times Sonnet here, not 5 times, because its cache reads cost only 1.25 times as much. Without caching, the ratio is 5.0 (calculations). This is Sonnet's token record repriced, not a run, and the Opus figure is provisional (see Caveats).

When Fable could be worth it

Use a break-even test: Fable costs 3.3 times Opus and 6.5 times Sonnet per hard pass (calculations). So Fable needs a pass rate more than 3.3 times Opus's, or 6.5 times Sonnet's, to cost less per success. Fable can win on cost only where Opus passes under about 30% of tasks, or Sonnet under about 15%, because no pass rate exceeds 100%. Our data cannot test two cases: a wrong answer that costs far more than a call, and work harder than our tasks. Pay for Fable only where your own validator shows a gap; ours show none.

How we measured

  • Setup. Claude Code, one turn, tools off, a fresh empty folder, no MCP servers. Opus means Opus 5.5 at default effort.
  • Sets. Five short tasks, 3 repetitions each: n = 15 per model. Eight hard tasks, 3 repetitions each: n = 24 per model. Validator controls for the hard set ran first: 8 of 8 reference answers passed and 26 of 26 plausible wrong answers failed.
  • Time and tokens. Time is a median with the fastest and slowest call; a range is not a confidence interval. Input tokens are a mean per call; hard-set output tokens are a median. Times include Claude Code start-up.
  • Cost. Reported tokens times list price per call, with cache reads and writes priced apart. Calculations.

Caveats

  • Ceiling. All three models passed every hard call. Harder or longer work could separate them. We did not measure that.
  • Small samples. Fifteen and 24 calls per model, on one host and one network (one day for the short set, one session for the hard set). Every timing range overlaps.
  • Cache state. The prompt cache stayed at the provider default, so cache-read counts depend on call order.
  • Opus cache-read price. We are re-checking the Opus 5.5 cache-read price in our table. The $143.83 Opus figure depends on it most: 94.0% of that workload's input was cache reads. The per-pass Opus costs use the same price. A hard-set call read only about 1,400 cache tokens. A doubled price would move those costs by about 1% (a calculation from the run receipts).
  • Prices. Provider prices are third-party-reported and change often.
  • Disclosure. We build Agent, which runs on Claude models.

Pay for the top model only where it helps

Agent records the model, the tokens and the result of every step, so you can see where a bigger model changes the outcome. Try Agent to test the price on your own work.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.