• Claude Code
  • Codex
  • CLI vs API
  • Tokens

Claude Code vs Codex CLI: the hidden context tax and what it does to speed

Codex CLI sent a median 12,124 input tokens per call; Claude Code sent 2,130. What a coding CLI's hidden context costs, and how much time the route adds.

TL;DR

  • A coding CLI wraps your prompt in its own system prompt and tool context. We call that the context tax.
  • For a one-line request, the Codex CLI sent 19,551 input tokens. The OpenAI API sent 17 for the same request.
  • Across our five-task head-to-head, Codex CLI sent a median 12,124 input tokens per call. Claude Code sent 2,130.
  • At GPT-6.1 Sol list prices, 19,551 tokens of context cost between $0.002 and $0.039 per request, depending on how much the cache serves. That is $1.96 to $39.10 per 1,000 requests (our calculation).
  • On a scheduler repair with 296 checks, Claude Code with Sonnet 5.5 took a median 15.0 s, the OpenAI API with GPT-6.1 Sol 17.3 s and the Codex CLI with the same GPT-6.1 Sol 61.2 s. The run ranges do not overlap, but each cell has only 3 runs.
  • Model and route change together in the Claude-vs-Codex rows. This data compares route plus model pairs, not the CLIs alone.

Side by side: /compare/claude-code-cli-vs-codex-cli.

What a CLI adds to every call

When you type a request into a coding CLI, the model does not see only your request. It sees:

  • the CLI's system prompt;
  • the definitions of every tool the CLI can call;
  • environment and session context.

All of that is input. It is sent on every call, and the model processes it before it writes your answer. Part of it comes from the prompt cache, which makes it cheaper but not free.

We measured this tax two ways: the same model through its API and through its CLI, and two CLIs on the same tasks.

Same model, API vs CLI

OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6 Luna · none
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

6 rows. Highest Codex CLI · GPT-6.1 Sol · high 19,555 (n 5). Lowest OpenAI API · GPT-6.1 Sol · high 17 (n 5).

Notesn = 5 per row

Reported input tokens, matched cohort

The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.

Source: Provider explorer receipts: CLI vs API

For one fixed one-line request, five runs per configuration:

  • OpenAI API: 17 input tokens, for every model and effort.
  • Codex CLI: 18,859 input tokens with GPT-6 Luna, 19,551 with GPT-6.1 Sol at low effort and 19,555 at high.

So the request was less than 0.1% of what the CLI sent (17 of 19,551, our calculation). The rest is the CLI's own context.

Time followed the tokens. For the one-line answer, the Codex CLI took a median 3.5 times as long as the API with the same model and effort. With GPT-6.1 Sol at low effort, the API took 1.02 s and the CLI 4.18 s. The ranges do not touch: the API's slowest run took 1.87 s and the CLI's fastest 3.86 s.

Live story · 33 sClaude Code CLI vs Codex CLI vs the API: a latency race

Claude Code CLI vs Codex CLI vs the API: a latency race

For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.

Transcript
  1. Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
  2. A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
  3. It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  4. A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
  5. Open benchmarks: intervals, sources and every failure kept.

On a small coding task the gap grew in seconds: GPT-6.1 Sol low took 6.0 s through the API and 14.15 s through the CLI.

Two CLIs on the same tasks

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

9 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (low) · Codex CLI 8,064 (n 10). Lowest Claude Haiku 4.5 · Claude Code 0 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Fable 5.1 · Claude Code 473 (n 15).

Notesn 10–15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

In our five-task head-to-head, every call ran through a CLI: Claude Code for the Claude models, Codex CLI for GPT-6.1 Sol. The mean input per call:

  • Claude Code with Sonnet 5.5: 1,401 cache-read tokens + 685 other input.
  • Claude Code with Opus 5.5: 1,401 + 680 at default effort.
  • Codex CLI with GPT-6.1 Sol: 5,180 to 8,064 cache-read tokens + 4,059 to 6,943 other input, depending on effort.

The three Codex efforts all add up to about 12,120 tokens (our calculation); only the cache split changes. The median over all calls was 12,124 for Codex CLI and 2,130 for Claude Code, about 5.7 times as much (our calculation).

The prompts were a few hundred tokens. Both CLIs add context; the Codex CLI adds several times more.

What the tax costs (a calculation)

The calls ran on subscriptions, so no one paid per token. To show the price of the context, we apply list prices. GPT-6.1 Sol lists at $2 per million input tokens and $0.10 per million cache reads.

For the 19,551-token one-liner:

  • All served from cache: 19,551 × $0.10 / 1M = $0.00196 per request.
  • None served from cache: 19,551 × $2 / 1M = $0.0391 per request.
  • Per 1,000 requests: $1.96 to $39.10 (our calculation), before the model writes a single output token.

The real figure sits between these, because part of the context is cached. In the head-to-head, 43% to 67% of the Codex CLI input was cache reads, depending on effort (our calculation from the means above).

The list price per call shows the same effect. GPT-6.1 Sol and Claude Sonnet 5.5 have the same headline prices: $2 input and $10 output per million. Yet the median call cost $0.0077 to $0.0105 through Codex CLI and $0.0036 for Sonnet in Claude Code. The Codex calls wrote fewer output tokens (a median 42 against 107). So the extra cost came from input, which is mostly CLI context. The ranges overlap, so the comparison page marks these rows "unclear".

Speed on real work

A one-line test exaggerates overhead. So we also gave three routes one real repair: a dependency-aware job scheduler, medium effort, 296 behavioral checks, 3 runs each.

  • Total time
  • First useful output
Entrance: medians race at 44× real time
Claude Code CLI · Sonnet 5.5 · medium
Codex CLI · GPT-6.1 Sol · medium
OpenAI API · GPT-6.1 Sol · medium

3 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · medium 61.2 s (range 59.9 s–69.5 s, n 3). Fastest Claude Code CLI · Sonnet 5.5 · medium 15 s (range 13.9 s–15.9 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · medium 15.6 s (range 13.7 s–23 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · medium 7.5 s (range 6.7 s–9.1 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Same prompt, medium effort, 296 behavioral checks, 3 runs each

All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.

Source: Provider explorer receipts: CLI vs API

All 9 runs passed all 296 checks. Median total time:

  • Claude Code · Sonnet 5.5: 15.0 s (13.89 to 15.89 s)
  • OpenAI API · GPT-6.1 Sol: 17.32 s (16.28 to 18.61 s)
  • Codex CLI · GPT-6.1 Sol: 61.16 s (59.9 to 69.51 s)

First useful output: 7.55 s, 7.46 s and 15.56 s.

The same GPT-6.1 Sol model was close to Claude Code through its API and about 3.5 times slower through its CLI (61.16 s against 17.32 s, our calculation). On this task, the route mattered more than the model. These are the only two rows on the comparison page where Claude Code leads: the run ranges do not overlap. With 3 runs per cell, treat them as strong direction, not a measured rate.

In the five-task head-to-head the pattern held at a smaller scale. Claude Code medians ran from 1.9 s to 4.4 s; Codex CLI medians from 5.6 s to 6.3 s. There the single-call ranges overlap, so the rows are "unclear".

And on hard tasks?

Updated 2026-10-06. The first 30 Codex CLI attempts in the hard head-to-head were blocked before any model call, and they stay in the record as blocked. A later batch reached the model: GPT-6.1 Sol in the Codex CLI passed 16 of 16 at medium and at high effort, the same ceiling as Sonnet, Opus and Fable at 24 of 24. Median time per call was 13.1 s and 18.1 s against 7.7 s for Sonnet in Claude Code; the ranges overlap. The context tax shows here too: on a one-word answer, the Codex CLI sent 17,051 input tokens and Claude Code 6,761 (different models). Full story: Claude vs Codex on hard tasks.

What to do about the context tax

  • Measure the route, not just the model. The same model took about 3.5 times as long through one wrapper, on two different tasks.
  • Check cache behaviour. The tax is 20 times cheaper when the cache serves it ($0.10 against $2 per million for GPT-6.1 Sol).
  • Watch first output. Through a CLI, first useful text can arrive late in the run.
  • Use the bare API for tiny, frequent calls. A one-line classification does not need 19,000 tokens of tool context.
  • Count CLI context in your budget. Our cost estimate walkthrough shows how.

Entity pages: Claude Code, Codex CLI and GPT-6.1 Sol in Codex CLI.

How we measured

  • Receipts: route (CLI or API), model, effort, timings, reported tokens and a deterministic validator result for every run.
  • Isolation: fresh folder, one turn, back-to-back cohorts on one host.
  • Costs: reported tokens × list price. Calculations; subscription calls carry no per-call price.
  • Ratios and per-1,000 figures in this post are our arithmetic on published values.

Caveats

  • Route and model change together in every Claude Code vs Codex CLI row.
  • Small samples: 3 to 5 runs per configuration for the CLI-vs-API study; 10 to 15 per configuration in the head-to-head. Medians with ranges.
  • Claude Code input in the CLI-vs-API study is not comparable: those receipts record only the uncached remainder. We use the head-to-head input counts instead.
  • One host, one network, one day per study. Vendor latency changes over the day.
  • Hard-task Codex cells are small: n = 16 per configuration.

Know what every call carries

Agent records the route, the input and output tokens, the cache hits and the time of every model call. Try Agent and see the context tax on your own work.

The data behind this post

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.