Recorded pass counts and medians; qualified and derived list-price costs
Configuration
Strict passes
95% Wilson pass interval
Cost per strict pass (calculation)
Derived mean cost per attempt
Recorded median call time (s)
Claude Sonnet 5.5 · Claude Code
24/24
86.2%–100%
$0.0144
$0.0144
7.75
Claude Opus 5.5 · Claude Code
24/24
86.2%–100%
$0.0282
$0.0282
9.18
Claude Opus 5.5 (high) · Claude Code
24/24
86.2%–100%
$0.0334
$0.0334
11.03
GPT-6.1 Sol (medium) · Codex CLI
16/16
80.64%–100%
$0.0256
$0.0256
13.11
Claude Fable 5.1 · Claude Code
24/24
86.2%–100%
$0.0933
$0.0933
16.13
GPT-6.1 Sol (high) · Codex CLI
16/16
80.64%–100%
$0.0151
$0.0151
18.12
Claude Haiku 4.5 · Claude Code
11/24
27.89%–64.93%
$0.0672
$0.0308
39.01
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.
Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.
Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
List-price costs are calculations; the calls used a flat subscription.
How the calculation works
Read exact strict pass counts k/n and their 95% Wilson intervals from the same study.
Derive mean attempt cost as the study cost per strict pass × k/n.
Expected attempts per correct answer = 1/(k/n), under a fixed independent probability.
Cost per correct answer = mean attempt cost × expected attempts. Multiply by your requested correct answers for a planning total.
Assumptions and limits
Each attempt is independent and keeps the same probability of passing. The pooled study rate does not predict retries on one repeatedly failing task.
Strict passes include the declared format rules. These results apply to the 8-task study and its CLI + model routes, not all work.
Mean cost per attempt is derived from the rounded study cost per strict pass and its exact pass count. Failures and format misses are included.
Costs use reported tokens at the recorded list prices. The runs used flat subscriptions. These are not vendor bills or current price quotes.
The cost sensitivity range changes only the pass rate over its 95% Wilson interval while holding mean cost fixed. It is not a confidence interval for cost.
The time proxy multiplies a recorded median by expected attempts. It is not expected elapsed time, a measured retry latency or a tested speed difference.
A zero pass rate or an absent, zero or unqualified cost gives no cost estimate. A missing cost is never treated as free.
Questions
How do retries change cost per correct answer?
With independent attempts at a fixed pass probability p, expected attempts per correct answer are 1/p. Multiply that by the mean cost per attempt. The mean here is derived from the study cost per strict pass and exact count ratio.
Is this a measured retry result?
No. The pass counts and call-time medians are recorded. The retry totals are calculations. A pooled rate across several tasks does not predict whether another attempt will fix a particular failed task.
What does the cost range mean?
It varies the pass probability over its 95% Wilson interval and holds mean cost fixed. It leaves cost variation and repeated-task dependence out. It is a sensitivity range, not a cost confidence interval.
Why can a perfect score still have a range?
A small perfect sample leaves uncertainty. A 24/24 score has a 95% Wilson lower bound of about 86%; a 16/16 score has one of about 81%. Neither proves a 100% future pass rate.
What happens when a value is missing?
A missing pass count removes that configuration. Missing or unqualified cost and time stay unknown. With zero recorded passes, there is no finite point estimate for retries until a correct answer.