[{"i":13,"story":{"id":"cli-vs-api-latency-race","title":"Claude Code CLI vs Codex CLI vs the API: a latency race","description":"For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.","studySlug":"cli-model-latency-tokens","chartIds":["cli-vs-api-exact-reply-latency","cli-vs-api-prompt-overhead","scheduler-repair-claude-vs-codex"],"durationSeconds":32.8,"transcript":["Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.","A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.","It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.","A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.","Open benchmarks: intervals, sources and every failure kept."]}}]