{"i":24,"slug":"haiku-retry-or-escalate","chart":{"id":"retry-escalate-success","title":"Expected success rate by retry policy (calculation)","subtitle":"Share of the eight hard tasks answered correctly under the strict validator","kind":"bar","unit":"rate","polarity":"higher","yLabel":"Correct answers","series":[{"name":"Expected success rate","points":{"$k":["label","value","n","lo","hi","highlight"],"$r":[["Sonnet 5.5 every time (calculation)",1,24,0.862,1,true],["Haiku, one retry, then Sonnet (calculation)",1,48,"\u0001","\u0001",false],["Haiku up to 3 tries, then Sonnet (calculation)",1,48,"\u0001","\u0001",false],["Sonnet 5.5 low effort once (calculation)",1,16,0.8064,1,false],["Haiku once (calculation)",0.4583,24,0.2789,0.6493,false]]}}],"note":"Whiskers are 95% Wilson intervals on the measured strict pass rate and appear only on the three one-try policies (24/24, 11/24 and 16/16). The two escalation policies are calculations without an interval. A value of 100% means no failure was recorded: the model treats Sonnet 5.5 as always right (24/24, 95% interval 86% to 100%), so the escalation policies cannot miss. n is the number of recorded calls behind each bar.","whisker":"ci95","sourceIds":["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"]}}