{"method":["This study makes no new model call and adds no new raw file. It reuses recorded calls from the hard head-to-head (Haiku 4.5 and Sonnet 5.5 at default effort, 3 calls per task) and the effort ladder (Sonnet 5.5 at low effort, 2 calls per task). All calls ran in Claude Code on the same 8 tasks with the same strict validators.","The source controls ran before counted inference. Each controls receipt records 8 passing references, 26 rejected wrong answers and 8 detected wrapped format misses. The hard-set protocol outcome says 27 wrong answers; the controls receipt supports 26. Both Claude batches completed their declared call caps: 120 hard-set calls and 80 ladder calls, with no trims or errors. They used a 300 s timeout and a 16,000-output-token setting. The selected calls stayed below that setting. No Claude probe is recorded in these source folders; the hard-set Codex probe is outside this calculation.","For each task and model, we take the pass rate p. It is strict passes ÷ calls, with its 95% Wilson interval. A format miss is not a pass.","These are median-input scenarios. A median is not an expected mean. Cost and time can depend on whether a call fails. We do not model that link. We take the cost of one call as the median list-price cost of that task’s calls. List price is reported tokens × price. Cache reads and one-hour cache writes are priced as in the hard head-to-head.","We take the time of one call as the median total time of that task’s calls. It includes CLI start-up.","A policy is a list of tries. A try runs only when every earlier try failed. The policies are: Sonnet every time, Haiku once, Haiku with one retry then Sonnet, Haiku with up to 3 tries then Sonnet, and Sonnet at low effort once.","Per task, let q = 1 − p(Haiku). For Haiku with one retry then Sonnet: cost = C(Haiku) × (1 + q) + C(Sonnet) × q². Time = T(Haiku) × (1 + q) + T(Sonnet) × q². Success = 1 − q² × (1 − p(Sonnet)).","For Haiku with up to 3 tries then Sonnet: cost = C(Haiku) × (1 + q + q²) + C(Sonnet) × q³. Success = 1 − q³ × (1 − p(Sonnet)). Time uses the same weights.","Over the 8 tasks, each task is equally likely. Cost per correct answer = total expected cost ÷ total expected correct answers. We compute time the same way, with the tries in sequence.","We assume three things. A validator detects a failed answer, and running it is free. Tries are independent at each task’s observed rate. A failed last Sonnet try is a miss. Every figure that follows is a calculation.","For the sensitivity, we set Haiku’s pass rate on all tasks at once to the lower end of its 95% Wilson interval, then to the upper end. A fourth setting counts Haiku’s format misses as passes. The result is a sensitivity range, not a confidence interval.","A joint extreme moves both models at once: Haiku’s pass rate to the upper end and Sonnet’s to the lower end of each task’s 95% Wilson interval. It is more extreme than a 95% interval on the total, and it is not a forecast.","For the break-even, we scale Haiku’s call cost (or time) by a factor f and keep all pass rates. We solve for the f at which Haiku with one retry then Sonnet equals Sonnet every time: f = (Σ C(Sonnet) − Σ C(Sonnet) × q²) ÷ Σ C(Haiku) × (1 + q). This is arithmetic, not a forecast."]}