{"i":23,"slug":"llm-speed-anatomy","chart":{"id":"speed-anatomy-prompt-size","title":"Time to first text as the prompt grows","subtitle":"Median of 3 calls per size; whiskers = fastest and slowest call","kind":"line","unit":"seconds","xLabel":"Prompt-size target (approximate Haiku tokens; calibration calculation)","yLabel":"Seconds to first text","series":{"$k":["name","points"],"$r":[["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.93,1.85,2.04,3],["16k",2.27,2.22,2.47,3],["64k",2.78,2.45,2.89,3]]}],["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.45,1.23,1.72,3],["16k",1.78,1.64,2.11,3],["64k",3.07,1.38,3.61,3]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.51,1.46,2.01,3],["16k",1.74,1.7,2.97,3],["64k",1.79,1.72,3.72,3]]}],["GPT-6.1 Sol (low) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["1k",3.36,3.36,4.75,3],["16k",4.02,3.3,4.28,3],["64k",3.93,3.42,4.38,3]]}]]},"note":"Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}}