Claude Fable 5.1 vs Opus 5.5 vs Sonnet 5.5: speed, tokens and price tested
24 of 24: Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed every hard task. Fable cost 3.3x Opus and 6.5x Sonnet per pass (list-price calculation).
TL;DR
- Quality is a tie at the ceiling. On eight hard tasks, Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed 24 of 24 (95% interval 86% to 100%). On five short tasks, Fable and Opus passed 15 of 15 and Sonnet 12 of 15, with intervals that overlap. The two comparison pages hold 28 rows: 7 ties, 21 unclear and no winner.
- The speed order changes with task length, and every range overlaps. Median time per short call: Fable 1.94 s, Sonnet 2.31 s, Opus 2.75 s. Per hard call: Fable 16.13 s, Sonnet 7.75 s, Opus 9.18 s. Neither order is a tested ranking.
- Fable read more input and wrote more output. It read 3,233 input tokens per short call, against 2,086 for Sonnet and 2,081 for Opus (a calculation). On hard tasks it wrote a median 1,366 output tokens, against 1,050 and 945.
- Fable costs the most. It lists at $10 input and $50 output per million tokens: 2.5x Opus and 5x Sonnet. Per hard pass, Fable cost $0.0933, Opus $0.0282 and Sonnet $0.0143 (calculations).
- When is Fable worth it? Only when your own validator shows a gap. Ours show none.
The side-by-side rows: Opus 5.5 vs Fable 5.1 and Sonnet 5.5 vs Fable 5.1. Model pages: Fable 5.1 and Opus 5.5.
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
127 of 130 calls passed, so speed and tokens separate the models: Fable 5.1 was fastest at 1.9 s median.
Transcript
- Head-to-head · 130 timed calls · 9 configurations. Haiku vs Sonnet vs Opus vs Fable vs Codex. Five short tasks with strict validators. Every call kept, nothing retried.
- 127 of 130 calls passed. Pass rate barely separates them; speed and tokens do. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Cheapest passing answer (list-price calculation): Sonnet 5.5: $0.0062 (n = 15). Caveat: The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.
- Fable 5.1 finishes first at 1.9 s. The Codex CLI needs 5.6–6.3 s. Chart: Median total time per call · real time (n = 10–15 each). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
- What the CLI sends: a median 12,124 input tokens per call on the Codex CLI, 2,130 on Claude Code. Chart: Input tokens per call: what the CLI sends (n = 10–15 each). Caveat: The prompt cache stayed at the provider default, so cache counters differ by route and by call order.
- Speed vs cost per passing answer. Ringed: no other setup is faster, cheaper per pass and as accurate. Chart: Speed, cost and quality frontier (n = 10–15 each). Calculation, not a run. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
- Open benchmarks: intervals, sources and every failure kept.
The story shows medians. Single-call ranges overlap, so read it as a description of these runs, not a ranking.
Fable 5.1 is the highest-priced Claude model in our studies. This post compares Claude Fable vs Opus and Claude Fable vs Sonnet on speed, tokens and price. We ran all three through Claude Code on two task sets.
Quality: a tie at the ceiling
On the hard set, all three models passed 24 of 24 calls strictly. A perfect 24 of 24 still has a 95% interval of 86% to 100%, so the three sit in one band. On the short set, Fable and Opus each passed 15 of 15 (80% to 100%) and Sonnet 12 of 15 (55% to 93%). Sonnet's three misses were format misses: the right number, plus extra lines. This is a ceiling, not proof of equality: our tasks cannot tell the three apart.
Speed on short tasks: the lowest median, not a proven lead
On five short tasks, Fable's median total time per call was 1.94 s (range 1.41 s to 9.83 s, n = 15). Sonnet took 2.31 s (2.17 s to 7.73 s) and Opus 2.75 s (2.47 s to 8.91 s). Time to first useful output ran in the same order: Fable 1.20 s, Sonnet 1.56 s, Opus 1.92 s. The ranges overlap, so no model is ahead: Fable's median is the lowest, and that is all these runs show.
Speed on hard tasks: the medians swap places
On eight hard tasks, Fable's median rose to 16.13 s (range 4.46 s to 90 s, n = 24). Sonnet took 7.75 s (2.26 s to 34.79 s) and Opus 9.18 s (4.24 s to 27.21 s). Fable's median was 2.1 times Sonnet's and 1.8 times Opus's (a calculation). The ranges still overlap, so this is not a tested ranking. The lowest median on short work became the highest of the three on hard work.
Tokens: what Fable reads and writes
On short calls, Fable read a mean 2,760 cache tokens and 473 other input tokens. Sonnet read 1,401 and 685, and Opus 1,401 and 680. Fable's total of 3,233 is about 1.55 times the 2,086 and 2,081 of the others (a calculation). Most input is the CLI's own system prompt and tool context, and we did not check why Fable's is larger. On hard tasks, Fable wrote a median 1,366 output tokens (889 of them reasoning); Sonnet wrote 1,050 (585) and Opus 945 (529).
Price per token: 2.5x to 5x
| Model | Input per M | Cache read per M | Output per M |
|---|---|---|---|
| Claude Fable 5.1 | $10 | $0.25 | $50 |
| Claude Opus 5.5 | $4 | under re-check | $20 |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 |
Fable's input and output prices are 2.5 times Opus's and 5 times Sonnet's (a calculation). OpenRouter's public API showed one Fable price at all four providers in our snapshot: Amazon Bedrock, Anthropic, Azure and Google Vertex (third-party-reported, 2026-10-06). OpenRouter's per-token price matches Anthropic's list, and it adds a 5.5% fee when you buy credits on its Standard plan.
Price per passing answer
We priced every reported token at list price, then divided each model's total by its passes, failures included. These are calculations, not bills: the calls ran on a subscription. On the hard set, Fable cost $0.0933 per strict pass, Opus $0.0282 and Sonnet $0.0143, which is 3.3 times Opus and 6.5 times Sonnet. On the short set, a passing answer cost Fable $0.0205, Opus $0.0101 and Sonnet $0.0062, or 2.0 and 3.3 times. Two effects stack: a higher price per token, and 1.3 times Sonnet's and 1.4 times Opus's output on the hard set.
Price in a long agent loop
We repriced the tokens of our agent's 33 SWE-bench instances (25 resolved): 162.9M input tokens (94.0% cache reads) and 1.8M output tokens. At list price, Fable comes to $321.31, Opus to $143.83 and Sonnet to $87.23, or $12.85, $5.75 and $3.49 per resolved instance. Fable is 3.7 times Sonnet here, not 5 times, because its cache reads cost only 1.25 times as much. Without caching, the ratio is 5.0 (calculations). This is Sonnet's token record repriced, not a run, and the Opus figure is provisional (see Caveats).
When Fable could be worth it
Use a break-even test: Fable costs 3.3 times Opus and 6.5 times Sonnet per hard pass (calculations). So Fable needs a pass rate more than 3.3 times Opus's, or 6.5 times Sonnet's, to cost less per success. Fable can win on cost only where Opus passes under about 30% of tasks, or Sonnet under about 15%, because no pass rate exceeds 100%. Our data cannot test two cases: a wrong answer that costs far more than a call, and work harder than our tasks. Pay for Fable only where your own validator shows a gap; ours show none.
How we measured
- Setup. Claude Code, one turn, tools off, a fresh empty folder, no MCP servers. Opus means Opus 5.5 at default effort.
- Sets. Five short tasks, 3 repetitions each: n = 15 per model. Eight hard tasks, 3 repetitions each: n = 24 per model. Validator controls for the hard set ran first: 8 of 8 reference answers passed and 26 of 26 plausible wrong answers failed.
- Time and tokens. Time is a median with the fastest and slowest call; a range is not a confidence interval. Input tokens are a mean per call; hard-set output tokens are a median. Times include Claude Code start-up.
- Cost. Reported tokens times list price per call, with cache reads and writes priced apart. Calculations.
Caveats
- Ceiling. All three models passed every hard call. Harder or longer work could separate them. We did not measure that.
- Small samples. Fifteen and 24 calls per model, on one host and one network (one day for the short set, one session for the hard set). Every timing range overlaps.
- Cache state. The prompt cache stayed at the provider default, so cache-read counts depend on call order.
- Opus cache-read price. We are re-checking the Opus 5.5 cache-read price in our table. The $143.83 Opus figure depends on it most: 94.0% of that workload's input was cache reads. The per-pass Opus costs use the same price. A hard-set call read only about 1,400 cache tokens. A doubled price would move those costs by about 1% (a calculation from the run receipts).
- Prices. Provider prices are third-party-reported and change often.
- Disclosure. We build Agent, which runs on Claude models.
What to read next
- Haiku vs Sonnet vs Opus vs Fable vs Codex: the five-task head-to-head
- When tasks get hard: Haiku, Sonnet, Opus and Fable
- What if every call ran on Opus?
- Sonnet vs Opus: when is Opus worth the price?
- Studies: short tasks, hard tasks, repricing and provider index
Pay for the top model only where it helps
Agent records the model, the tokens and the result of every step, so you can see where a bigger model changes the outcome. Try Agent to test the price on your own work.