{"i":11,"slug":"routing-overhead","chart":{"id":"router-overhead-cli-vs-model-time","title":"Where an LLM router’s time goes: model vs CLI","subtitle":"Median per call; whiskers = median to 95th percentile","kind":"dot-range","unit":"ms","yLabel":"Time per call","whisker":"p50-p95","series":[{"name":"Model API time","points":[{"label":"Claude Sonnet 5.5 (effort low, via Claude Code)","value":1596,"lo":1596,"hi":2583,"n":82},{"label":"Claude Haiku 4.5 (thinking on, via Claude Code)","value":10508,"lo":10508,"hi":32132,"n":82}]},{"name":"CLI and harness time","points":[{"label":"Claude Sonnet 5.5 (effort low, via Claude Code)","value":973,"lo":973,"hi":1277,"n":82},{"label":"Claude Haiku 4.5 (thinking on, via Claude Code)","value":1698,"lo":1698,"hi":2677,"n":82}]}],"note":"Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.","sourceIds":["agent-routing-overhead","agent-routing"]}}