Explainer · CLI context tax
Tokens per call and the CLI context tax
Definition
Tokens per call is the number of input and output tokens one model request uses, and the CLI context tax is the part of the input that a coding CLI such as Claude Code or the Codex CLI adds on its own: its system prompt, tool definitions and environment context, sent with every request before your prompt. You pay for these tokens (or spend subscription quota on them), and the model must read them, even when your question is one line.
Agent team · · 4 min read · Every number is from the public studies
Interactive
The same one-line request, with and without a CLI around it
Each block is 1,000 input tokens. Most of what the CLI sends is its own system prompt and tool context.
OpenAI API17 tokens
Codex CLI19,551 tokens
1.2 thousand x the input for the same request (ratio of the two values shown)
| Route and model | Input tokens | n (calls) |
|---|---|---|
| OpenAI API · GPT-6 Luna · none | 17 | 5 |
| OpenAI API · GPT-6.1 Sol · low | 17 | 5 |
| OpenAI API · GPT-6.1 Sol · high | 17 | 5 |
| Codex CLI · GPT-6 Luna · none | 18,859 | 5 |
| Codex CLI · GPT-6.1 Sol · low | 19,551 | 5 |
| Codex CLI · GPT-6.1 Sol · high | 19,555 | 5 |
GPT-6.1 Sol · low: OpenAI API sent 17 input tokens, Codex CLI sent 19,551 (n = 5 calls each). That is 1.2 thousand x the input, as the ratio of the two values shown.
The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.
Source: CLI vs API study
Where the tokens come from
A request through a bare API holds what you send: a system message if you write one, and the user message. A request through a coding CLI holds much more:
- The CLI's system prompt: how to behave, how to format, how to use tools.
- Tool definitions: a schema for every tool (read file, edit file, run shell, search) the model may call.
- Environment context: the working directory, the platform, sometimes project instructions.
- The conversation so far in a multi-turn session.
All of this counts as input tokens. Prompt caching can make the repeated part cheaper, but it is still sent and still counted.
How big is the tax?
The cleanest measure is the same one-line request sent through the OpenAI API and through the Codex CLI, with the same model and effort:
- The OpenAI API reported 17 input tokens for the one-line request.
- The Codex CLI reported about 19,551 for the same request with GPT-6.1 Sol at low effort (n = 5 per cell).
On short tasks in our head-to-head, the median call through the Codex CLI used 12,124 input tokens, against 2,130 through Claude Code:
The prompts themselves were a few hundred tokens. Most input was the CLI's own context. A large part was served from the prompt cache, which lowers its price but not its count.
For a one-word answer, Claude Code with Haiku 4.5 sent 6,761 input tokens per call and the Codex CLI (default model) sent 17,051 (5 runs each, different models).
The tax is also time
Tokens are not the only overhead. A CLI starts a process, loads its configuration and tools, and may run more than one internal step:
- For a one-line answer, the Codex CLI took a median 3.5 times as long as the OpenAI API with the same model and effort.
- On a one-word answer, Claude Code spent a median 1,690 ms outside the model (1,533 to 1,811 ms, n = 5).
On real work the picture changes. Repairing a scheduler with 296 checks, all 9 runs passed: Claude Code with Sonnet 5.5 took a median 15.0 s, the OpenAI API with GPT-6.1 Sol 17.3 s and the Codex CLI with GPT-6.1 Sol 61.2 s. With 3 to 5 runs per cell these are directional, not rankings. On a longer task the fixed CLI overhead is a smaller share of the total, and the routes differ more by how each one works through the task.
When the tax matters
- Many short calls. Classification, routing and one-line answers pay the full context for a tiny output. Here a direct API call is much leaner. See what is an LLM router: about a second of each Sonnet router decision was CLI time.
- Per-token billing. On an API key, about 19,500 extra input tokens per call (median 19,551 vs 17 tokens, n = 15) is real money at scale, cached or not. On a flat subscription, it uses quota.
- Model comparisons. A Claude-vs-Codex cost gap measured through CLIs is partly the CLIs. Compare models on the same route, or say that the route differs.
It matters less for long agent tasks, where the CLI's tools earn their context and the cache absorbs most of the repeat.
How to reduce it
- Use the bare API for short, tool-free calls.
- Keep long-running CLI sessions open, so the prefix is cached once and reused. See prompt caching, explained.
- Turn off tools and project instructions a step does not need, where the CLI allows it.
- Read the usage fields per call (input, cache read, cache write, output), not only the total.
Frequently asked questions
How many tokens does Claude Code send per request?
On our short tasks, a median of 2,130 input tokens per call, most of it Claude Code's own context, much of it from the prompt cache. For a one-word answer with Haiku 4.5 it sent 6,761 input tokens.
Why does the Codex CLI use so many input tokens?
It sends its own system prompt, tool schemas and environment context with every request. For a one-line request it reported about 19,551 input tokens, where the OpenAI API reported 17 for the same request.
Does prompt caching remove the CLI context tax?
It lowers the price of the repeated part, because cache reads are cheaper than normal input. It does not remove the tokens: they are still sent, counted and read by the model.
Is a CLI slower than the API?
For tiny requests, yes: the Codex CLI took a median 3.5 times as long as the API for a one-line answer. On a real repair task, Claude Code had the lowest median of the three routes we timed, but with 3 runs per cell that is directional, not a ranking.
Watch the data
Claude Code CLI vs Codex CLI vs the API: a latency race
For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.
Transcript
- Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
- A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
- It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
- A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
- Open benchmarks: intervals, sources and every failure kept.