{"i":14,"study":{"slug":"cli-model-latency-tokens","title":"Claude Code CLI vs Codex CLI vs the API: latency and tokens","seoTitle":"Claude Code vs Codex CLI vs API: latency and token overhead","description":"194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.","question":"How much time and how many tokens does a coding CLI add on top of the model, and how do Claude Code and Codex compare on the same repair task?","answer":"For a one-line answer, the Codex CLI took a median 3.5 times as long as the OpenAI API with the same model and effort, and it sent about 19,551 input tokens instead of 17. On a dependency-aware scheduler repair with 296 checks, all 9 runs passed: Claude Code with Sonnet 5.5 took a median 15.0 s, the OpenAI API with GPT-6.1 Sol 17.3 s and the Codex CLI with GPT-6.1 Sol 61.2 s. With 3 to 5 runs per cell these are directional measurements, not rankings.","date":"2026-10-03","updated":"2026-10-05","tags":["claude-code","codex","latency","tokens","cli-vs-api"],"caveats":["Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.","The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.","The Claude CLI receipts record only the uncached remainder of the input (2 tokens), so the Claude input column is not comparable.","All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.","API costs in the raw file are list-price estimates; CLI runs are subscription calls with no per-call price."],"sourceIds":["agent-provider-explorer"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["explorer-pass-rate","Evaluated runs that passed their validator",1,"rate","100% (194/194)",194,[0.9806,1],"18 excluded and 18 diagnostic receipts are not counted."],["exact-reply-cli-over-api","Codex CLI vs OpenAI API, median total time for a one-line answer",3.49,"ratio","3.5x slower",30,"\u0001","Medians 3.9 s (CLI) vs 1.1 s (API), same models and efforts."],["scheduler-claude-cli-median","Claude Code CLI (Sonnet 5.5) median time to repair the scheduler",15,"seconds","15.0 s",3,"\u0001","\u0001"],["scheduler-codex-cli-median","Codex CLI (GPT-6.1 Sol) median time to repair the scheduler",61.2,"seconds","61.2 s",3,"\u0001","\u0001"],["cli-hidden-prompt","Median input tokens the Codex CLI sends for a one-line request",19551,"tokens","19,551",15,"\u0001","The API sends 17 tokens for the same request."]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["cli-vs-api-exact-reply-latency","CLI vs API: time for a one-line answer","Matched cohort, fixed exact reply, 5 runs per configuration","dot-range","seconds","Seconds",[{"name":"Total time","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.97,0.65,1.5,5],["OpenAI API · GPT-6.1 Sol · low",1.02,0.96,1.87,5],["OpenAI API · GPT-6.1 Sol · high",1.52,1.35,2.23,5],["Codex CLI · GPT-6 Luna · none",3.19,2.88,3.83,5],["Codex CLI · GPT-6.1 Sol · low",4.18,3.86,4.53,5],["Codex CLI · GPT-6.1 Sol · high",4.19,3.81,4.69,5]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.82,0.51,1.37,5],["OpenAI API · GPT-6.1 Sol · low",0.87,0.84,1.74,5],["OpenAI API · GPT-6.1 Sol · high",1.34,1.26,2.12,5],["Codex CLI · GPT-6 Luna · none",2.79,2.46,3.42,5],["Codex CLI · GPT-6.1 Sol · low",3.75,3.44,4.1,5],["Codex CLI · GPT-6.1 Sol · high",3.79,3.37,4.3,5]]}}],"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.",["agent-provider-explorer"]],["cli-vs-api-small-coding-latency","CLI vs API: time for a small coding task","Matched cohort, small coding task, 3 runs per configuration","dot-range","seconds","Seconds",[{"name":"Total time","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",4.01,3.83,4.35,3],["OpenAI API · GPT-6.1 Sol · low",6,5.44,6.2,3],["Codex CLI · GPT-6 Luna · none",9.23,8.99,11.68,3],["OpenAI API · GPT-6.1 Sol · high",9.56,9.44,10.94,3],["Codex CLI · GPT-6.1 Sol · low",14.15,13.02,14.41,3],["Codex CLI · GPT-6.1 Sol · high",17.85,17.68,22.42,3]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.67,0.62,0.81,3],["OpenAI API · GPT-6.1 Sol · low",1.05,0.97,1.4,3],["Codex CLI · GPT-6 Luna · none",8.68,8.27,11.01,3],["OpenAI API · GPT-6.1 Sol · high",5.31,4.99,6.42,3],["Codex CLI · GPT-6.1 Sol · low",13.6,12.52,13.83,3],["Codex CLI · GPT-6.1 Sol · high",17.27,17.13,21.86,3]]}}],"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.",["agent-provider-explorer"]],["cli-vs-api-prompt-overhead","Hidden prompt: input tokens for the same one-line request","Reported input tokens, matched cohort","bar","tokens","Input tokens per call",[{"name":"Input tokens","points":{"$k":["label","value","n"],"$r":[["OpenAI API · GPT-6 Luna · none",17,5],["OpenAI API · GPT-6.1 Sol · low",17,5],["OpenAI API · GPT-6.1 Sol · high",17,5],["Codex CLI · GPT-6 Luna · none",18859,5],["Codex CLI · GPT-6.1 Sol · low",19551,5],["Codex CLI · GPT-6.1 Sol · high",19555,5]]}}],"The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.",["agent-provider-explorer"]],["scheduler-repair-claude-vs-codex","Repairing a scheduler: Claude Code vs Codex vs API","Same prompt, medium effort, 296 behavioral checks, 3 runs each","dot-range","seconds","Seconds",[{"name":"Total time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Code CLI · Sonnet 5.5 · medium",15,13.89,15.89,3,true],["Codex CLI · GPT-6.1 Sol · medium",61.16,59.9,69.51,3,"\u0001"],["OpenAI API · GPT-6.1 Sol · medium",17.32,16.28,18.61,3,"\u0001"]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Code CLI · Sonnet 5.5 · medium",7.55,6.77,7.63,3],["Codex CLI · GPT-6.1 Sol · medium",15.56,13.65,23.04,3],["OpenAI API · GPT-6.1 Sol · medium",7.46,6.68,9.05,3]]}}],"All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.",["agent-provider-explorer"]],["scheduler-repair-output-tokens","Output tokens to repair the scheduler","Median per run; reasoning tokens shown separately where reported","grouped-bar","tokens","Tokens",[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Code CLI · Sonnet 5.5 · medium",2227,3],["Codex CLI · GPT-6.1 Sol · medium",1181,3],["OpenAI API · GPT-6.1 Sol · medium",1313,3]]}},{"name":"Reasoning tokens (reported)","points":{"$k":["label","value","n"],"$r":[["Claude Code CLI · Sonnet 5.5 · medium",0,0],["Codex CLI · GPT-6.1 Sol · medium",156,3],["OpenAI API · GPT-6.1 Sol · medium",267,3]]}}],"The Claude CLI does not report reasoning tokens separately; 0 there means \"not reported\", not \"none\".",["agent-provider-explorer"]]]}}}