• Latency
  • Claude Code
  • CLI vs API
  • Reasoning Effort

Why is Claude Code slow? Where the seconds go in a coding CLI call, and what to change

A one-word Claude Code call took a median 2.5 s, with 1.7 s outside the model (n = 5). Where the rest goes: thinking, effort, route. Measured splits and fixes.

TL;DR

  • Part of a call runs outside the model. Claude Code answered a one-word prompt in a median 2.53 s (n = 5, range 2.27 to 3.38 s). The study counts a median 1.69 s outside the model (range 1.53 to 1.81 s). That time includes start-up and a step after the answer, so it is not all start-up.
  • Thinking can fill a call. On hard tasks, Haiku 4.5 wrote a median 4,556 reasoning tokens and took a median 39.01 s per call. Sonnet 5.5 took a median 7.75 s. The ranges overlap, so this is not a tested gap.
  • The median time rose with effort on our set, but the ranges overlap, so it is not a reliable gap. Sonnet 5.5 took a median 5.82 s at low effort and 8.81 s at high. Both passed 16 of 16 (95% interval 81% to 100%).
  • The route matters for tiny calls. For a one-line answer, the Codex CLI took 3.5 times as long as the OpenAI API with the same model (a calculation on medians). We did not measure a direct Claude API call.
  • Our reading, untested before and after: try a lower effort where the task allows it, and test it on your own task. Check a model's default thinking before a fast step. Consider an API for tiny calls.

The data: routing overhead and CLI vs API latency.

Live story · 33 sClaude Code CLI vs Codex CLI vs the API: a latency race

Claude Code CLI vs Codex CLI vs the API: a latency race

For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.

Transcript
  1. Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
  2. A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
  3. It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  4. A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
  5. Open benchmarks: intervals, sources and every failure kept.

Each section below covers one place where time can go, with its measured number. Section 7 lists the changes we would try.

1. Time outside the model, including start-up

  • First output event
  • First model output
  • Total wall time
Entrance: medians race at 4.3× real timeMotion reduced: press Replay to animateThe slowest median is 6 s. The clock runs at the recorded speed.
Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Median of 5 runs; whiskers = fastest and slowest run

Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

We asked Claude Code (Haiku 4.5) for one word, with tools, MCP servers and sessions off. The median of 5 runs was 563 ms to the first output event and 1,461 ms to the first model output. The call ended at 2,529 ms. The whiskers are the fastest and slowest run, not an interval.

The study counts a median 1,690 ms outside the model (range 1,533 to 1,811 ms). It measures this as wall time minus the API time the CLI reported. It covers process start, init and a post-turn summary step before exit, so not all of it is start-up. The study does not split it into those parts. Time outside the model is about two thirds of the total (a calculation on two medians). It is not all start-up.

Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows. Highest Codex CLI (default model) 17,051 (n 5). Lowest Claude Code · Claude Haiku 4.5 6,761 (n 5).

Notesn = 5 per row

Per call, mostly the CLI’s own system prompt and tool definitions

Claude Code sums its disjoint input, cache-read and cache-write fields. Codex CLI reports 17,051 input tokens, 13,184 of them read from the cache. The prompt itself is a few tokens.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

The CLI sent 6,761 input tokens before the one word (cache reads and writes included). The Codex CLI sent 17,051 and took 5,999 ms (n = 5 each). The two CLIs ran different models, so CLI and model are not separated. The comparison page puts Claude Code ahead on first model output and total time, where the run ranges do not overlap. The rows for first output event and input tokens have no winner.

2. The CLI adds time on every call

  • Model API time
  • CLI and harness time
Entrance: medians race at 7.5× real timeMotion reduced: press Replay to animateThe slowest median is 10.51 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

2 rows, 2 series: Model API time, CLI and harness time. Model API time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 10.51 s (median to p95 10.51 s–32.13 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 1.6 s (median to p95 1.6 s–2.58 s, n 82). Not all run ranges overlap. CLI and harness time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 1.7 s (median to p95 1.7 s–2.68 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 973 ms (median to p95 973 ms–1.28 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n = 82 per row

Median per call; whiskers = median to 95th percentile

Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing

We recorded 82 routing calls through Claude Code with Sonnet 5.5 at low effort. Median model time was 1,596 ms (p95 2,583 ms). Median CLI and harness time was 973 ms (p95 1,277 ms). The whole call took a median 2,597 ms (p95 4,298 ms), so the CLI share is about 37% (a calculation on medians). As in section 1, CLI and harness time is wall time minus the API time the CLI reported. Medians of parts do not add up to the median of the whole: 1,596 + 973 is 2,569 ms (a calculation).

For contrast, Jev 1.13, a hosted decision model, was called directly over HTTPS from the same Mac. It took a median 136.5 ms per call (p95 195.7 ms, n = 246 calls; study table). That is client wall time with the network inside it. It is a different route from the CLI, so it shows what a caller waits per decision, not model compute time.

3. Thinking tokens: the model works before it answers

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).

Notesn 16–24 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

On 8 hard tasks (study), Haiku 4.5 wrote a median 5,064 output tokens per call (n = 24). Its median reasoning tokens, which count inside the output, were 4,556. Sonnet 5.5 wrote a median 1,050, with 585 reasoning tokens.

Haiku's median time was 39.01 s (range 15.27 to 75.13 s) and Sonnet's was 7.75 s (range 2.26 to 34.79 s). Haiku's median is about 5 times as long (a calculation). The ranges overlap, so the comparison page marks this row unclear.

The comparison page has a winner for the routing row, because the p50 to p95 bands do not overlap. Through the same CLI, Haiku took a median 12,543 ms (p95 34,481 ms). Sonnet took 2,597 ms (p95 4,298 ms). Each has n = 82.

In the routing runs, two settings changed at once: Haiku ran with the CLI default thinking (the study labels it thinking on) and Sonnet ran at low effort. In the hard set, Haiku and Sonnet both ran without an effort flag. The studies cited in this post did not test Haiku with thinking off.

The slower model did not do better. On the hard set, Haiku passed 11 of 24 strictly (46%, 95% interval 28% to 65%) and Sonnet 24 of 24 (86% to 100%). Sonnet is ahead on that row. On the 82 routing decisions (study), Haiku was exactly right on 73 (89%, 80% to 94%) and Sonnet on 77 (94%, 87% to 97%). That row is a tie.

4. Effort: a higher median, not a tested gap

Seconds (median)

* default: the effort flag was not passed; its level is not known, so no line joins it.

4 efforts, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest high 8.8 s (n 16). Fastest low 5.8 s (n 16). Claude Opus 5.5 · Claude Code: slowest high 10.1 s (n 16). Fastest low 7.5 s (n 16).

Notesn = 16 per row

One line per model and route; default = the effort flag was not passed

Medians only; the per-call ranges are in the total-time chart and they overlap. "default" is placed last because its level is not known: the CLI chose it.

Source: Effort ladder: the hard task set at each effort level

Sonnet 5.5 ran the same 8 hard tasks at three efforts (study, 16 calls each). The median time per call was 5.82 s at low, 7.63 s at medium and 8.81 s at high. Median output tokens were 667, 770 and 1,192. Every cell passed 16 of 16 (95% interval 81% to 100%). The ranges overlap (low 2.78 to 19.96 s, high 2.93 to 35.81 s), so these medians are not a ranking. The dataset marks the time rows for low vs medium and low vs high effort unclear for that reason.

With no effort flag, Sonnet's median was 7.97 s over 16 calls (the 24-call median in section 3 is 7.75 s). This set has a ceiling, so it cannot show whether low effort is enough for harder work. See Reasoning effort, explained.

5. Model choice on short calls

Entrance: medians race at 4.5× real timeMotion reduced: press Replay to animateThe slowest median is 6.3 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

9 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 6.3 s (range 4.7 s–10.5 s, n 10). Fastest Claude Fable 5.1 · Claude Code 1.9 s (range 1.4 s–9.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

On five short tasks (study), the Claude Code medians ran from 1.94 s for Fable 5.1 to 4.43 s for Haiku 4.5. Their ranges are 1.41 to 9.83 s and 3.16 to 23.57 s. Sonnet 5.5 took a median 2.31 s (n = 15 per cell). The ranges overlap, so we do not rank the models.

6. Route: CLI or API

  • Total time
  • First useful output
Entrance: medians race at 3× real timeMotion reduced: press Replay to animateThe slowest median is 4.2 s. The clock runs at the recorded speed.
OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6 Luna · none
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

6 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 4.2 s (range 3.8 s–4.7 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 1 s (range 0.7 s–1.5 s, n 5). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 3.8 s (range 3.4 s–4.3 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 0.8 s (range 0.5 s–1.4 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Matched cohort, fixed exact reply, 5 runs per configuration

Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.

Source: Provider explorer receipts: CLI vs API

For a one-line answer, with the same models and efforts, the Codex CLI took a median 3.9 s and the OpenAI API 1.1 s. That is 3.5 times as long (a calculation on the medians of 15 runs per route). For GPT-6.1 Sol at high effort, the medians were 4.19 s and 1.52 s, and the ranges do not overlap (n = 5 each). The comparison page puts the API ahead on that row. The Codex CLI sent a median 19,551 input tokens for the line (n = 15 runs), and the API sent 17.

The story above ends on a scheduler repair. Claude Code with Sonnet 5.5 took a median 15.0 s. The OpenAI API with GPT-6.1 Sol took 17.3 s (n = 3 runs each, all passed). The models differ and the runs are few, so the comparison page marks that row unclear. It does not show that Claude Code is faster than an API call.

We did not measure a direct Claude API call. This pair shows that the Codex CLI route took seconds longer than the API route for the same model. It does not show how many seconds Claude Code adds.

7. What to change (our reading of the data)

We did not test these changes before and after. Each rests on a number above.

  1. Try a lower effort where the task allows it (section 4). Sonnet's median was 5.82 s at low and 8.81 s at high, and both passed 16 of 16. The ranges overlap, so test it on your own task. Try --effort low.
  2. Check default thinking before you pick a model for a fast step (section 3). Haiku's default run took a median 12,543 ms per routing call. The studies cited here did not test thinking off, so test it yourself.
  3. Use rules or a hosted decision model for routing (sections 2 and 3). The in-process rule policy decided in a median 1.42 µs (p95 2.33 µs, n = 20,000). Jev took a median 136.5 ms per call over HTTPS. Routing every call through Sonnet (a median 49.5 calls per task) would add up to 128.6 s of waiting per task. Through Jev it would add up to 6.76 s. Those are calculations and upper bounds: they assume each decision waits for the one before. On the 82 typed decisions, Jev was exactly right on 90% (221 of 246 live calls). The 95% interval is taken at 82 decisions and runs from 82% to 95%. Sonnet 5.5 was exactly right on 77 of 82 (94%, 87% to 97%). The intervals overlap, so accuracy does not separate them. Jev's case sets were tuned against Jev answers. The routing-overhead study gives no accuracy score for the rules.
  4. Consider the API for tiny calls (section 6). The Codex CLI took a median 3.9 s and the OpenAI API 1.1 s for one line (15 runs per route). Time a direct Claude API call on your own host.
  5. Try a session that stays open instead of a new process for every call (section 1). Part of the 1,690 ms outside the model is process start and init. Part is a post-turn step that may stay in a warm session. We did not time a warm session.

How we measured

  • One-word run: a one-word prompt, 5 runs per CLI, one call at a time, on one Mac. Time outside the model is wall time minus the API time the CLI reported. It includes process start, init and a post-turn summary step before exit.
  • Routing: 82 recorded calls per model. CLI and harness time is wall time minus the API time the CLI reported. Jev: 246 live calls (3 repeats of the same 82 decisions), one at a time over HTTPS from the same Mac, on 2026-10-06.
  • Hard, effort and short sets: validated tasks, fresh folder, tools off, one turn. Calls per Claude Code configuration: hard 24, effort 16, short 15.
  • CLI vs API: a matched cohort on one host on 2026-10-03: 3 model and effort pairs per route, 5 runs each.
  • Ratios are our arithmetic on published medians. A range is not a confidence interval, and p95 is a percentile, not an interval. Pass rates carry 95% Wilson intervals.

Caveats

  • One host, small samples. The one-word and CLI-vs-API cells have 5 runs each. Provider speed changes through the day.
  • Isolated one-word run. A normal session with tools, MCP servers or project files may send more. We did not measure that. The 1,690 ms outside the model is not split into start-up and the step after the answer.
  • Two settings changed at once in the routing runs (Haiku thinking on, Sonnet low effort).
  • Jev is a different route. Its timing is one 35-second window from one Mac on a home network, and its 246 calls are 3 repeats of 82 requests, so they are not independent draws. The API reports no server time.
  • Ceiling. 6 of 7 hard-set configurations passed every call, and every effort cell passed 16 of 16.
  • Batch timing. The default-effort Sonnet cell ran in a different batch and hour from the other three.
  • Not tested in the studies cited here: a direct Claude API call, a warm session, and Haiku with thinking off.

Find where your own seconds go

Agent records the model, the tokens and the result of every step. Try Agent and see which of your calls earn their seconds.

The data behind this post

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.