{"$k":["slug","chart"],"$r":[["cli-model-latency-tokens",{"id":"cli-vs-api-exact-reply-latency","title":"CLI vs API: time for a one-line answer","subtitle":"Matched cohort, fixed exact reply, 5 runs per configuration","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.97,0.65,1.5,5],["OpenAI API · GPT-6.1 Sol · low",1.02,0.96,1.87,5],["OpenAI API · GPT-6.1 Sol · high",1.52,1.35,2.23,5],["Codex CLI · GPT-6 Luna · none",3.19,2.88,3.83,5],["Codex CLI · GPT-6.1 Sol · low",4.18,3.86,4.53,5],["Codex CLI · GPT-6.1 Sol · high",4.19,3.81,4.69,5]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.82,0.51,1.37,5],["OpenAI API · GPT-6.1 Sol · low",0.87,0.84,1.74,5],["OpenAI API · GPT-6.1 Sol · high",1.34,1.26,2.12,5],["Codex CLI · GPT-6 Luna · none",2.79,2.46,3.42,5],["Codex CLI · GPT-6.1 Sol · low",3.75,3.44,4.1,5],["Codex CLI · GPT-6.1 Sol · high",3.79,3.37,4.3,5]]}}],"note":"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.","sourceIds":["agent-provider-explorer"]}],["cli-model-latency-tokens",{"id":"cli-vs-api-small-coding-latency","title":"CLI vs API: time for a small coding task","subtitle":"Matched cohort, small coding task, 3 runs per configuration","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",4.01,3.83,4.35,3],["OpenAI API · GPT-6.1 Sol · low",6,5.44,6.2,3],["Codex CLI · GPT-6 Luna · none",9.23,8.99,11.68,3],["OpenAI API · GPT-6.1 Sol · high",9.56,9.44,10.94,3],["Codex CLI · GPT-6.1 Sol · low",14.15,13.02,14.41,3],["Codex CLI · GPT-6.1 Sol · high",17.85,17.68,22.42,3]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.67,0.62,0.81,3],["OpenAI API · GPT-6.1 Sol · low",1.05,0.97,1.4,3],["Codex CLI · GPT-6 Luna · none",8.68,8.27,11.01,3],["OpenAI API · GPT-6.1 Sol · high",5.31,4.99,6.42,3],["Codex CLI · GPT-6.1 Sol · low",13.6,12.52,13.83,3],["Codex CLI · GPT-6.1 Sol · high",17.27,17.13,21.86,3]]}}],"note":"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.","sourceIds":["agent-provider-explorer"]}],["cli-model-latency-tokens",{"id":"cli-vs-api-prompt-overhead","title":"Hidden prompt: input tokens for the same one-line request","subtitle":"Reported input tokens, matched cohort","kind":"bar","unit":"tokens","yLabel":"Input tokens per call","series":[{"name":"Input tokens","points":{"$k":["label","value","n"],"$r":[["OpenAI API · GPT-6 Luna · none",17,5],["OpenAI API · GPT-6.1 Sol · low",17,5],["OpenAI API · GPT-6.1 Sol · high",17,5],["Codex CLI · GPT-6 Luna · none",18859,5],["Codex CLI · GPT-6.1 Sol · low",19551,5],["Codex CLI · GPT-6.1 Sol · high",19555,5]]}}],"note":"The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.","sourceIds":["agent-provider-explorer"]}],["llm-speed-anatomy",{"id":"speed-anatomy-first-text","title":"Time to first text: a 250-line answer, six models","subtitle":"Median of 4 calls per model; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds to first text","series":[{"name":"Time to first text","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",4,2.84,6.38,4],["Claude Sonnet 5.5 · Claude Code",1.96,0.88,4.09,4],["Claude Opus 5.5 · Claude Code",1.97,1.7,2.35,4],["Claude Fable 5.1 · Claude Code",4.43,2.27,4.64,4],["GPT-6.1 Sol (low) · Codex CLI",3.52,2.75,4.42,4],["GPT-6 Luna (low) · Codex CLI",3.3,3.19,3.47,4]]}}],"note":"Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-output-speed","title":"Output speed after the first text: visible tokens per second (calculation)","subtitle":"Median of 4 calls per model; whiskers = slowest and fastest call","kind":"dot-range","unit":"tokens","yLabel":"Visible tokens per second","series":[{"name":"Visible tokens per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",153.2,152.6,216.1,4],["Claude Sonnet 5.5 · Claude Code",231.7,230.3,233,4],["Claude Opus 5.5 · Claude Code",155.5,154.6,156.4,4],["Claude Fable 5.1 · Claude Code",122.6,120.9,131.4,4],["GPT-6.1 Sol (low) · Codex CLI",79.6,71.6,80.5,4],["GPT-6 Luna (low) · Codex CLI",129.1,55.5,259.1,4]]}}],"note":"Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-chars-per-second","title":"Output speed in characters per second after the first text (calculation)","subtitle":"Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call","kind":"dot-range","unit":"count","yLabel":"Characters per second","series":[{"name":"Characters per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",547,546,548,3],["Claude Sonnet 5.5 · Claude Code",517,513,519,4],["Claude Opus 5.5 · Claude Code",347,345,349,4],["Claude Fable 5.1 · Claude Code",273,270,293,4],["GPT-6.1 Sol (low) · Codex CLI",323,291,327,4],["GPT-6 Luna (low) · Codex CLI",524,225,1052,4]]}}],"note":"Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}]]}