Explainer · Cost per correct answer

Cost per correct answer: the LLM price that counts failures

Definition

Cost per correct answer is the total cost of every call, failed calls included, divided by the number of attempts that pass a check. An attempt can use one call or a full agent session. A lower token price can cost more per pass when a model writes more tokens or fails more often.

Agent team · · 5 min read · Every number is from the public studies

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

The formula

cost per correct answer = cost of all calls ÷ number of passes

Count every call. A pass means a deterministic check accepts the result. A right answer in the wrong format is a format miss, not a pass.

Inputs from our hard-task study, 24 single-call attempts each. Both models used Claude Code at its default effort. All rate intervals below are 95% Wilson intervals:

  • Claude Sonnet 5.5: 24 of 24 passes (95% interval 86% to 100%). Cost per pass $0.01435. So all 24 calls cost about $0.344 (calculation: $0.01435 × 24).
  • Claude Haiku 4.5: 11 of 24 passes (46%, interval 28% to 65%). Cost per pass $0.0672. So all 24 calls cost about $0.739 (calculation: $0.0672 × 11).

Costs are list-price calculations, not subscription invoices. Head-to-head costs price reported tokens, with cache reads and writes priced apart. Memory and SWE-bench costs use recorded notional estimates.

Why price per token misleads

Our 2026-09-21 price snapshot lists Haiku 4.5 at $1 per million input tokens and $5 per million output tokens. Sonnet 5.5 lists at $2 and $10. Haiku is half the price (calculation: 1 ÷ 2 and 5 ÷ 10).

  • Output length. On five short tasks, Haiku wrote a median 367 output tokens per call (range 275 to 2,851). Sonnet wrote 107 (range 44 to 745; n = 15 calls each). Both used Claude Code at default effort. That is 3.4 times as many tokens (calculation: 367 ÷ 107).
  • Thinking. Haiku reported a median 297 reasoning tokens per call (range 212 to 2,643; n = 15). These tokens are part of output, not an extra meter to add again.
  • Failures. Failed calls spread spend over fewer passes.

Median cost per call was $0.00566 for Haiku (range $0.00513 to $0.01804). Sonnet cost $0.00360 (range $0.00342 to $0.01021; n = 15 each). The ranges overlap, so this is not a ranking. Ranges show observed limits, not confidence intervals.

Calculation
Claude Fable 5.1
Claude Sonnet 5.5
Claude Opus 5.5 (high)
Claude Opus 5.5
Claude Opus 5.5 (low)
Claude Haiku 4.5
GPT-6.1 Sol (high)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (low)

Every interval overlaps every other: this chart does not order these rows.

List-price calculation, not a run. 9 rows. Highest GPT-6.1 Sol (high) · Codex CLI $0.01 (range $0.0066–$0.028, n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0036 (range $0.0034–$0.01, n 15). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 10–15 per row

Reported tokens × list price; the calls ran on subscriptions

Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.

Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

On the same five tasks, Haiku passed 15 of 15 (80% to 100%) and Sonnet 12 of 15 (55% to 93%). The intervals overlap, so pass rate does not separate them. The short task set hit a ceiling for most configurations. All three Sonnet misses were format misses. Cost per pass was $0.00836 for Haiku and $0.00624 for Sonnet (n = 15 each). Cost has no interval; this gap describes this run.

What the hard tasks showed

On eight hard tasks, the gap grew:

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

  • 4.7 times per pass. Haiku cost $0.0672 per strict pass, Sonnet $0.01435 (calculation: $0.0672 ÷ $0.01435).
  • Pass rate separates them. The intervals, 28% to 65% and 86% to 100%, do not overlap. Sonnet is ahead on pass rate on this set.
  • Tokens. Haiku wrote a median 5,064 output tokens per call (range 1,899 to 9,321), including 4,556 reasoning tokens (range 1,452 to 8,569). Sonnet wrote 1,050 (range 176 to 3,895), including 585 reasoning tokens (range 0 to 3,060; n = 24 each). That is 4.8 times as many for Haiku (calculation: 5,064 ÷ 1,050).

Five more configurations passed every call. These six hit a ceiling, so this set cannot rank their quality. GPT-6.1 Sol (high), Codex CLI: $0.01514 per pass (16/16; 81% to 100%). Opus 5.5: $0.02824; Fable 5.1: $0.09331 (24/24 each; 86% to 100%). Models and routes change together; this cannot isolate a CLI effect.

What moves it

Effort, pass rate and memory can change this cost:

  • Effort. In our effort study, Sonnet 5.5 passed 16/16 (81% to 100%) at each effort. Cost per pass rose from $0.0122 at low to $0.0167 at high, 37% more (calculation from unrounded receipts: ($0.016705225 ÷ $0.01219145 − 1) × 100). The set hit a ceiling; extra spend added no passes. This does not show effort's effect on harder tasks.
  • Pass rate. In the SWE-bench Verified study, our agent resolved 25 of 33 instances (76%, interval 59% to 87%). Notional cost was $2.81 per attempt and $3.71 per resolved instance, 32% more (calculation: (33 ÷ 25 − 1) × 100).
  • Memory. In our agent-memory study, 15 of 15 Sonnet sessions with an 11-line memory file passed in full, at $0.0818 per full pass. With no memory, 9 of 15 passed, at $0.1386. The intervals, 79.6% to 100% and 35.8% to 80.2%, overlap, so the gap is unclear. One small synthetic repository; the curated file used task knowledge. Questions received no answer and counted as failures.

How to compute yours

  1. Log every call. Record tokens by meter (input, cache write, cache read, output). Keep timeouts and empty replies.
  2. Define a pass before you run. Use a deterministic check, such as tests or an exact match. Count format misses apart.
  3. Name your price basis. Use your invoice or list prices. We use list prices.
  4. Divide. Total cost of all calls, divided by passes. Show k of n and the 95% interval beside it.
  5. Compare like with like. Keep the tasks, validators, route and effort the same. See how to read AI benchmarks honestly.
  6. Watch for zero passes. With 0 passes the ratio has no value. Report the spend and n.

Our AI cost calculator prices your own token counts. This page does not model retries. Repeated calls share only five or eight tasks in the head-to-head studies. Their call-level intervals do not show uncertainty across all possible tasks.

Frequently asked questions

Is Claude Haiku cheaper than Sonnet?

Per token, yes. Per pass on these hard tasks, no. We tested Haiku at Claude Code default effort only.

Why can a cheaper model cost more per task?

It can write more tokens, fail more often, or both. The hard-task figures above show both.

Should I count failed attempts?

Yes. Dropping failures hides their cost. Keep unresolved attempts in the total.

Is list price the real cost?

Our runs used subscriptions. These calculations are not invoices. Discounts, cache rules and gateway fees can change your bill.

Receipt sources: short tasks, hard tasks, effort ladder, memory sessions and SWE-bench attempts.

Watch the data

Live story · 101 sClaude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

Claude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

66 comparison rows from 10 studies: 0 rows favour Sonnet 5.5, 0 favour Opus 5.5, 66 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 66 rows · 10 studies. Sonnet 5.5 vs Opus 5.5. A winner only where the 95% intervals or run ranges do not overlap.
  2. 66 comparison rows from 10 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Sonnet 5.5 is ahead: 0 (of 66). Rows where Opus 5.5 is ahead: 0 (of 66). Ties or unclear: 66 (16 ties · 50 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. SWE-bench pairs, interim: pass rate 33% vs 67%, tie: 95% intervals overlap. None of the 3 rows separates them. Table: SWE-bench pairs, interim · Agent · n = 3 per side. Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
  4. Five short tasks: pass rate 80% vs 100%, tie: 95% intervals overlap. None of the 8 rows separates them. Table: Five short tasks · Claude Code · n = 15 per side. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  5. Eight hard tasks: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 6 rows separates them. Table: Eight hard tasks · Claude Code · n = 24 per side. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  6. Coding agents, hidden tests: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code · n = 12 per side. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  7. Effort ladder, default effort: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code · n = 16 per side. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  8. Caching sessions: pass rate not measured. None of the 4 rows separates them. Table: Caching sessions · Claude Code · n = 15, 3, 12 per side. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  9. Prompt cache break-even: after how many reuses does a cached prefix cost less?: pass rate not measured. None of the 9 rows separates them. Table: Prompt cache break-even: after how many reuses does a cached prefix cost less? · no cache · n = per side. Caveat: The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.
  10. How much of an AI bill is thinking? Reasoning tokens by model and effort: pass rate not measured. None of the 8 rows separates them. Table: How much of an AI bill is thinking? Reasoning tokens by model and effort · Claude Code · n = 24, 16, 15 per side. Caveat: The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.
  11. Where the seconds go: first text, output speed and prompt size for 6 LLMs: pass rate 100% vs 56%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: Where the seconds go: first text, output speed and prompt size for 6 LLMs · Claude Code · n = 4, 3, 9 per side. Caveat: First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).
  12. GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: pass rate 38% vs 42%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks · Claude Code · n = per side. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
  13. No winner where the data shows none. Every row and its reason online.

The data behind this explainer

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.