{"i":24,"slug":"haiku-retry-or-escalate","chart":{"id":"retry-escalate-sensitivity","title":"Sensitivity: cost per correct answer if Haiku’s pass rate is higher or lower (calculation)","subtitle":"Dot: Haiku’s pass rate as observed. Whisker: its pass rate on every task set to its Wilson lower and upper bound at once","kind":"dot-range","unit":"usd","yLabel":"USD per correct answer","series":[{"name":"Cost per correct answer (calculation)","points":{"$k":["label","value","n","highlight","lo","hi"],"$r":[["Sonnet 5.5 low effort once (calculation)",0.01219,16,false,"\u0001","\u0001"],["Sonnet 5.5 every time (calculation)",0.01322,24,true,"\u0001","\u0001"],["Haiku, one retry, then Sonnet (calculation)",0.05664,48,false,0.03941,0.06807],["Haiku once (calculation)",0.06772,24,false,0.03908,0.1834],["Haiku up to 3 tries, then Sonnet (calculation)",0.07193,48,false,0.04149,0.09047]]}}],"note":"Calculation. This is a sensitivity range, not a confidence interval: each Haiku policy is recomputed with Haiku’s pass rate on all 8 tasks at once set to the lower and then the upper end of that task’s 95% Wilson interval (3 calls per task, so the ends are far apart), and the whisker spans the lowest and highest result. That is more extreme than a 95% interval on the total. The two Sonnet policies do not depend on Haiku’s pass rate and have no whisker. Highlighted: the baseline, Sonnet 5.5 every time.","whisker":"minmax","sourceIds":["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"]}}