{"i":17,"slug":"haiku-thinking-on-off","chart":{"id":"haiku-thinking-router-latency","title":"Haiku thinking study: time per routing decision","subtitle":"Median wall time and model (API) time; whisker to the 95th percentile","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Wall time (CLI)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",4.66,4.66,8.18,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",12.54,12.54,34.48,82,false],["Claude Sonnet 5.5 (low) · Claude Code",2.6,2.6,4.3,82,false]]}},{"name":"Model time (API)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",3.79,3.79,7.43,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",10.51,10.51,32.13,82,false],["Claude Sonnet 5.5 (low) · Claude Code",1.6,1.6,2.58,82,false]]}}],"note":"Whiskers run from p50 to p95, not a confidence interval. Percentiles use the nearest-rank rule of the routing-overhead study, so the thinking-on and Sonnet medians match that study (the Jev-vs-LLM page interpolates between ranks and shows slightly different values). One call at a time through the Claude Code CLI; the thinking-on and Sonnet arms ran on another day.","whisker":"p50-p95","factContext":"typed routing decisions, thinking on vs off","sourceIds":["agent-haiku-thinking","agent-routing"]}}