Explainer · CLI context tax

Tokens per call and the CLI context tax

Definition

Tokens per call is the number of input and output tokens one model request uses, and the CLI context tax is the part of the input that a coding CLI such as Claude Code or the Codex CLI adds on its own: its system prompt, tool definitions and environment context, sent with every request before your prompt. You pay for these tokens (or spend subscription quota on them), and the model must read them, even when your question is one line.

Agent team · · 4 min read · Every number is from the public studies

Interactive

The same one-line request, with and without a CLI around it

Each block is 1,000 input tokens. Most of what the CLI sends is its own system prompt and tool context.

Measured
Model

OpenAI API17 tokens

Codex CLI19,551 tokens

1.2 thousand x the input for the same request (ratio of the two values shown)

GPT-6.1 Sol · low: OpenAI API sent 17 input tokens, Codex CLI sent 19,551 (n = 5 calls each). That is 1.2 thousand x the input, as the ratio of the two values shown.

The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.

Source: CLI vs API study

Where the tokens come from

A request through a bare API holds what you send: a system message if you write one, and the user message. A request through a coding CLI holds much more:

  • The CLI's system prompt: how to behave, how to format, how to use tools.
  • Tool definitions: a schema for every tool (read file, edit file, run shell, search) the model may call.
  • Environment context: the working directory, the platform, sometimes project instructions.
  • The conversation so far in a multi-turn session.

All of this counts as input tokens. Prompt caching can make the repeated part cheaper, but it is still sent and still counted.

How big is the tax?

The cleanest measure is the same one-line request sent through the OpenAI API and through the Codex CLI, with the same model and effort:

OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6 Luna · none
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

6 rows. Highest Codex CLI · GPT-6.1 Sol · high 19,555 (n 5). Lowest OpenAI API · GPT-6.1 Sol · high 17 (n 5).

Notesn = 5 per row

Reported input tokens, matched cohort

The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.

Source: Provider explorer receipts: CLI vs API

  • The OpenAI API reported 17 input tokens for the one-line request.
  • The Codex CLI reported about 19,551 for the same request with GPT-6.1 Sol at low effort (n = 5 per cell).

On short tasks in our head-to-head, the median call through the Codex CLI used 12,124 input tokens, against 2,130 through Claude Code:

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

9 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (low) · Codex CLI 8,064 (n 10). Lowest Claude Haiku 4.5 · Claude Code 0 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Fable 5.1 · Claude Code 473 (n 15).

Notesn 10–15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

The prompts themselves were a few hundred tokens. Most input was the CLI's own context. A large part was served from the prompt cache, which lowers its price but not its count.

For a one-word answer, Claude Code with Haiku 4.5 sent 6,761 input tokens per call and the Codex CLI (default model) sent 17,051 (5 runs each, different models).

The tax is also time

Tokens are not the only overhead. A CLI starts a process, loads its configuration and tools, and may run more than one internal step:

  • For a one-line answer, the Codex CLI took a median 3.5 times as long as the OpenAI API with the same model and effort.
  • On a one-word answer, Claude Code spent a median 1,690 ms outside the model (1,533 to 1,811 ms, n = 5).

On real work the picture changes. Repairing a scheduler with 296 checks, all 9 runs passed: Claude Code with Sonnet 5.5 took a median 15.0 s, the OpenAI API with GPT-6.1 Sol 17.3 s and the Codex CLI with GPT-6.1 Sol 61.2 s. With 3 to 5 runs per cell these are directional, not rankings. On a longer task the fixed CLI overhead is a smaller share of the total, and the routes differ more by how each one works through the task.

When the tax matters

  • Many short calls. Classification, routing and one-line answers pay the full context for a tiny output. Here a direct API call is much leaner. See what is an LLM router: about a second of each Sonnet router decision was CLI time.
  • Per-token billing. On an API key, about 19,500 extra input tokens per call (median 19,551 vs 17 tokens, n = 15) is real money at scale, cached or not. On a flat subscription, it uses quota.
  • Model comparisons. A Claude-vs-Codex cost gap measured through CLIs is partly the CLIs. Compare models on the same route, or say that the route differs.

It matters less for long agent tasks, where the CLI's tools earn their context and the cache absorbs most of the repeat.

How to reduce it

  • Use the bare API for short, tool-free calls.
  • Keep long-running CLI sessions open, so the prefix is cached once and reused. See prompt caching, explained.
  • Turn off tools and project instructions a step does not need, where the CLI allows it.
  • Read the usage fields per call (input, cache read, cache write, output), not only the total.

Frequently asked questions

How many tokens does Claude Code send per request?

On our short tasks, a median of 2,130 input tokens per call, most of it Claude Code's own context, much of it from the prompt cache. For a one-word answer with Haiku 4.5 it sent 6,761 input tokens.

Why does the Codex CLI use so many input tokens?

It sends its own system prompt, tool schemas and environment context with every request. For a one-line request it reported about 19,551 input tokens, where the OpenAI API reported 17 for the same request.

Does prompt caching remove the CLI context tax?

It lowers the price of the repeated part, because cache reads are cheaper than normal input. It does not remove the tokens: they are still sent, counted and read by the model.

Is a CLI slower than the API?

For tiny requests, yes: the Codex CLI took a median 3.5 times as long as the API for a one-line answer. On a real repair task, Claude Code had the lowest median of the three routes we timed, but with 3 runs per cell that is directional, not a ranking.

Watch the data

Live story · 33 sClaude Code CLI vs Codex CLI vs the API: a latency race

Claude Code CLI vs Codex CLI vs the API: a latency race

For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.

Transcript
  1. Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
  2. A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
  3. It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  4. A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
  5. Open benchmarks: intervals, sources and every failure kept.

The data behind this explainer

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.