{"i":4,"slug":"hard-model-head-to-head","chart":{"id":"hard-h2h-first-useful-latency","title":"Time to first useful output on hard tasks","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Time to first useful output on hard tasks","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",5.95,0.86,30.57,24,true],["Claude Opus 5.5 · Claude Code",6.78,2.39,21.77,24,true],["Claude Opus 5.5 (high) · Claude Code",7.13,2.15,56.23,24,true],["GPT-6.1 Sol (medium) · Codex CLI",10.23,6.09,40.41,16,true],["Claude Fable 5.1 · Claude Code",11.63,2,85.33,24,true],["GPT-6.1 Sol (high) · Codex CLI",12.69,8.93,75.91,16,true],["Claude Haiku 4.5 · Claude Code",35.54,12.88,70.31,24,false]]}}],"note":"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.","sourceIds":["agent-provider-h2h-hard"]}}