Calculation · recorded strict pass rates

Cost per correct answer: include retries

A low price per attempt can hide failed answers. Use one study’s pass counts and qualified costs to see what independent retries would imply.

This is a calculation on one study, not a vendor invoice, a customer saving or a measured retry campaign.

Choose two recorded routes

Example input · one 8-task study

Calculation · retries under fixed, independent success probability

Cost per correct answer · calculation

Claude Haiku 4.5 · Claude Code

$0.0672

Recorded strict passes: 11/24 (45.83%).

95% Wilson pass-rate interval: 27.89%–64.93%.

Expected attempts per correct answer
2.18
Cost sensitivity range
$0.0474–$0.1104
Median time proxy per correct answer
85.11 seconds

Cost sensitivity holds mean cost fixed; it is not a cost confidence interval. The time proxy is not expected elapsed time.

Cost per correct answer · calculation

Claude Sonnet 5.5 · Claude Code

$0.0144

Recorded strict passes: 24/24 (100%).

95% Wilson pass-rate interval: 86.2%–100%.

Expected attempts per correct answer
1
Cost sensitivity range
$0.0144–$0.0166
Median time proxy per correct answer
7.75 seconds

Cost sensitivity holds mean cost fixed; it is not a cost confidence interval. The time proxy is not expected elapsed time.

Calculation for 100 correct answers
ConfigurationExpected attemptsList-price totalMedian time proxy (seconds)
Claude Haiku 4.5 · Claude Code218.18$6.728,511.27
Claude Sonnet 5.5 · Claude Code100$1.435775

Totals are calculations. There is no guarantee of a finite retry count or a deadline for a particular task.

The recorded inputs

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks · updated 2026-10-06. Each row is one CLI + model route. Sample sizes differ.

Recorded pass counts and medians; qualified and derived list-price costs
ConfigurationStrict passes95% Wilson pass intervalCost per strict pass (calculation)Derived mean cost per attemptRecorded median call time (s)
Claude Sonnet 5.5 · Claude Code24/2486.2%–100%$0.0144$0.01447.75
Claude Opus 5.5 · Claude Code24/2486.2%–100%$0.0282$0.02829.18
Claude Opus 5.5 (high) · Claude Code24/2486.2%–100%$0.0334$0.033411.03
GPT-6.1 Sol (medium) · Codex CLI16/1680.64%–100%$0.0256$0.025613.11
Claude Fable 5.1 · Claude Code24/2486.2%–100%$0.0933$0.093316.13
GPT-6.1 Sol (high) · Codex CLI16/1680.64%–100%$0.0151$0.015118.12
Claude Haiku 4.5 · Claude Code11/2427.89%–64.93%$0.0672$0.030839.01

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Study limits

  • 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  • Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  • Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  • The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.
  • Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  • CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.
  • Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  • List-price costs are calculations; the calls used a flat subscription.

How the calculation works

  1. Read exact strict pass counts k/n and their 95% Wilson intervals from the same study.
  2. Derive mean attempt cost as the study cost per strict pass × k/n.
  3. Expected attempts per correct answer = 1/(k/n), under a fixed independent probability.
  4. Cost per correct answer = mean attempt cost × expected attempts. Multiply by your requested correct answers for a planning total.

Assumptions and limits

  • Each attempt is independent and keeps the same probability of passing. The pooled study rate does not predict retries on one repeatedly failing task.
  • Strict passes include the declared format rules. These results apply to the 8-task study and its CLI + model routes, not all work.
  • Mean cost per attempt is derived from the rounded study cost per strict pass and its exact pass count. Failures and format misses are included.
  • Costs use reported tokens at the recorded list prices. The runs used flat subscriptions. These are not vendor bills or current price quotes.
  • The cost sensitivity range changes only the pass rate over its 95% Wilson interval while holding mean cost fixed. It is not a confidence interval for cost.
  • The time proxy multiplies a recorded median by expected attempts. It is not expected elapsed time, a measured retry latency or a tested speed difference.
  • A zero pass rate or an absent, zero or unqualified cost gives no cost estimate. A missing cost is never treated as free.

Questions

How do retries change cost per correct answer?

With independent attempts at a fixed pass probability p, expected attempts per correct answer are 1/p. Multiply that by the mean cost per attempt. The mean here is derived from the study cost per strict pass and exact count ratio.

Is this a measured retry result?

No. The pass counts and call-time medians are recorded. The retry totals are calculations. A pooled rate across several tasks does not predict whether another attempt will fix a particular failed task.

What does the cost range mean?

It varies the pass probability over its 95% Wilson interval and holds mean cost fixed. It leaves cost variation and repeated-task dependence out. It is a sensitivity range, not a cost confidence interval.

Why can a perfect score still have a range?

A small perfect sample leaves uncertainty. A 24/24 score has a 95% Wilson lower bound of about 86%; a 16/16 score has one of about 81%. Neither proves a 100% future pass rate.

What happens when a value is missing?

A missing pass count removes that configuration. Missing or unqualified cost and time stay unknown. With zero recorded passes, there is no finite point estimate for retries until a correct answer.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.