Why is Claude Code slow? Where the seconds go in a coding CLI call, and what to change
A one-word Claude Code call took a median 2.5 s, with 1.7 s outside the model (n = 5). Where the rest goes: thinking, effort, route. Measured splits and fixes.
TL;DR
- Part of a call runs outside the model. Claude Code answered a one-word prompt in a median 2.53 s (n = 5, range 2.27 to 3.38 s). The study counts a median 1.69 s outside the model (range 1.53 to 1.81 s). That time includes start-up and a step after the answer, so it is not all start-up.
- Thinking can fill a call. On hard tasks, Haiku 4.5 wrote a median 4,556 reasoning tokens and took a median 39.01 s per call. Sonnet 5.5 took a median 7.75 s. The ranges overlap, so this is not a tested gap.
- The median time rose with effort on our set, but the ranges overlap, so it is not a reliable gap. Sonnet 5.5 took a median 5.82 s at low effort and 8.81 s at high. Both passed 16 of 16 (95% interval 81% to 100%).
- The route matters for tiny calls. For a one-line answer, the Codex CLI took 3.5 times as long as the OpenAI API with the same model (a calculation on medians). We did not measure a direct Claude API call.
- Our reading, untested before and after: try a lower effort where the task allows it, and test it on your own task. Check a model's default thinking before a fast step. Consider an API for tiny calls.
The data: routing overhead and CLI vs API latency.
Claude Code CLI vs Codex CLI vs the API: a latency race
For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.
Transcript
- Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
- A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
- It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
- A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
- Open benchmarks: intervals, sources and every failure kept.
Each section below covers one place where time can go, with its measured number. Section 7 lists the changes we would try.
1. Time outside the model, including start-up
We asked Claude Code (Haiku 4.5) for one word, with tools, MCP servers and sessions off. The median of 5 runs was 563 ms to the first output event and 1,461 ms to the first model output. The call ended at 2,529 ms. The whiskers are the fastest and slowest run, not an interval.
The study counts a median 1,690 ms outside the model (range 1,533 to 1,811 ms). It measures this as wall time minus the API time the CLI reported. It covers process start, init and a post-turn summary step before exit, so not all of it is start-up. The study does not split it into those parts. Time outside the model is about two thirds of the total (a calculation on two medians). It is not all start-up.
The CLI sent 6,761 input tokens before the one word (cache reads and writes included). The Codex CLI sent 17,051 and took 5,999 ms (n = 5 each). The two CLIs ran different models, so CLI and model are not separated. The comparison page puts Claude Code ahead on first model output and total time, where the run ranges do not overlap. The rows for first output event and input tokens have no winner.
2. The CLI adds time on every call
We recorded 82 routing calls through Claude Code with Sonnet 5.5 at low effort. Median model time was 1,596 ms (p95 2,583 ms). Median CLI and harness time was 973 ms (p95 1,277 ms). The whole call took a median 2,597 ms (p95 4,298 ms), so the CLI share is about 37% (a calculation on medians). As in section 1, CLI and harness time is wall time minus the API time the CLI reported. Medians of parts do not add up to the median of the whole: 1,596 + 973 is 2,569 ms (a calculation).
For contrast, Jev 1.13, a hosted decision model, was called directly over HTTPS from the same Mac. It took a median 136.5 ms per call (p95 195.7 ms, n = 246 calls; study table). That is client wall time with the network inside it. It is a different route from the CLI, so it shows what a caller waits per decision, not model compute time.
3. Thinking tokens: the model works before it answers
On 8 hard tasks (study), Haiku 4.5 wrote a median 5,064 output tokens per call (n = 24). Its median reasoning tokens, which count inside the output, were 4,556. Sonnet 5.5 wrote a median 1,050, with 585 reasoning tokens.
Haiku's median time was 39.01 s (range 15.27 to 75.13 s) and Sonnet's was 7.75 s (range 2.26 to 34.79 s). Haiku's median is about 5 times as long (a calculation). The ranges overlap, so the comparison page marks this row unclear.
The comparison page has a winner for the routing row, because the p50 to p95 bands do not overlap. Through the same CLI, Haiku took a median 12,543 ms (p95 34,481 ms). Sonnet took 2,597 ms (p95 4,298 ms). Each has n = 82.
In the routing runs, two settings changed at once: Haiku ran with the CLI default thinking (the study labels it thinking on) and Sonnet ran at low effort. In the hard set, Haiku and Sonnet both ran without an effort flag. The studies cited in this post did not test Haiku with thinking off.
The slower model did not do better. On the hard set, Haiku passed 11 of 24 strictly (46%, 95% interval 28% to 65%) and Sonnet 24 of 24 (86% to 100%). Sonnet is ahead on that row. On the 82 routing decisions (study), Haiku was exactly right on 73 (89%, 80% to 94%) and Sonnet on 77 (94%, 87% to 97%). That row is a tie.
4. Effort: a higher median, not a tested gap
Sonnet 5.5 ran the same 8 hard tasks at three efforts (study, 16 calls each). The median time per call was 5.82 s at low, 7.63 s at medium and 8.81 s at high. Median output tokens were 667, 770 and 1,192. Every cell passed 16 of 16 (95% interval 81% to 100%). The ranges overlap (low 2.78 to 19.96 s, high 2.93 to 35.81 s), so these medians are not a ranking. The dataset marks the time rows for low vs medium and low vs high effort unclear for that reason.
With no effort flag, Sonnet's median was 7.97 s over 16 calls (the 24-call median in section 3 is 7.75 s). This set has a ceiling, so it cannot show whether low effort is enough for harder work. See Reasoning effort, explained.
5. Model choice on short calls
On five short tasks (study), the Claude Code medians ran from 1.94 s for Fable 5.1 to 4.43 s for Haiku 4.5. Their ranges are 1.41 to 9.83 s and 3.16 to 23.57 s. Sonnet 5.5 took a median 2.31 s (n = 15 per cell). The ranges overlap, so we do not rank the models.
6. Route: CLI or API
For a one-line answer, with the same models and efforts, the Codex CLI took a median 3.9 s and the OpenAI API 1.1 s. That is 3.5 times as long (a calculation on the medians of 15 runs per route). For GPT-6.1 Sol at high effort, the medians were 4.19 s and 1.52 s, and the ranges do not overlap (n = 5 each). The comparison page puts the API ahead on that row. The Codex CLI sent a median 19,551 input tokens for the line (n = 15 runs), and the API sent 17.
The story above ends on a scheduler repair. Claude Code with Sonnet 5.5 took a median 15.0 s. The OpenAI API with GPT-6.1 Sol took 17.3 s (n = 3 runs each, all passed). The models differ and the runs are few, so the comparison page marks that row unclear. It does not show that Claude Code is faster than an API call.
We did not measure a direct Claude API call. This pair shows that the Codex CLI route took seconds longer than the API route for the same model. It does not show how many seconds Claude Code adds.
7. What to change (our reading of the data)
We did not test these changes before and after. Each rests on a number above.
- Try a lower effort where the task allows it (section 4). Sonnet's median was 5.82 s at low and 8.81 s at high, and both passed 16 of 16. The ranges overlap, so test it on your own task. Try
--effort low. - Check default thinking before you pick a model for a fast step (section 3). Haiku's default run took a median 12,543 ms per routing call. The studies cited here did not test thinking off, so test it yourself.
- Use rules or a hosted decision model for routing (sections 2 and 3). The in-process rule policy decided in a median 1.42 µs (p95 2.33 µs, n = 20,000). Jev took a median 136.5 ms per call over HTTPS. Routing every call through Sonnet (a median 49.5 calls per task) would add up to 128.6 s of waiting per task. Through Jev it would add up to 6.76 s. Those are calculations and upper bounds: they assume each decision waits for the one before. On the 82 typed decisions, Jev was exactly right on 90% (221 of 246 live calls). The 95% interval is taken at 82 decisions and runs from 82% to 95%. Sonnet 5.5 was exactly right on 77 of 82 (94%, 87% to 97%). The intervals overlap, so accuracy does not separate them. Jev's case sets were tuned against Jev answers. The routing-overhead study gives no accuracy score for the rules.
- Consider the API for tiny calls (section 6). The Codex CLI took a median 3.9 s and the OpenAI API 1.1 s for one line (15 runs per route). Time a direct Claude API call on your own host.
- Try a session that stays open instead of a new process for every call (section 1). Part of the 1,690 ms outside the model is process start and init. Part is a post-turn step that may stay in a warm session. We did not time a warm session.
How we measured
- One-word run: a one-word prompt, 5 runs per CLI, one call at a time, on one Mac. Time outside the model is wall time minus the API time the CLI reported. It includes process start, init and a post-turn summary step before exit.
- Routing: 82 recorded calls per model. CLI and harness time is wall time minus the API time the CLI reported. Jev: 246 live calls (3 repeats of the same 82 decisions), one at a time over HTTPS from the same Mac, on 2026-10-06.
- Hard, effort and short sets: validated tasks, fresh folder, tools off, one turn. Calls per Claude Code configuration: hard 24, effort 16, short 15.
- CLI vs API: a matched cohort on one host on 2026-10-03: 3 model and effort pairs per route, 5 runs each.
- Ratios are our arithmetic on published medians. A range is not a confidence interval, and p95 is a percentile, not an interval. Pass rates carry 95% Wilson intervals.
Caveats
- One host, small samples. The one-word and CLI-vs-API cells have 5 runs each. Provider speed changes through the day.
- Isolated one-word run. A normal session with tools, MCP servers or project files may send more. We did not measure that. The 1,690 ms outside the model is not split into start-up and the step after the answer.
- Two settings changed at once in the routing runs (Haiku thinking on, Sonnet low effort).
- Jev is a different route. Its timing is one 35-second window from one Mac on a home network, and its 246 calls are 3 repeats of 82 requests, so they are not independent draws. The API reports no server time.
- Ceiling. 6 of 7 hard-set configurations passed every call, and every effort cell passed 16 of 16.
- Batch timing. The default-effort Sonnet cell ran in a different batch and hour from the other three.
- Not tested in the studies cited here: a direct Claude API call, a warm session, and Haiku with thinking off.
What to read next
- Claude Code vs Codex CLI vs the API: latency, time to first token and hidden prompts
- Tokens per call and the CLI context tax
- Does reasoning effort buy quality? Claude and Codex on hard tasks, low to high
- What does a router cost you? Rules vs Jev vs an LLM router
Find where your own seconds go
Agent records the model, the tokens and the result of every step. Try Agent and see which of your calls earn their seconds.