• Latency
  • Claude Code
  • Codex
  • CLI vs API

Claude Code vs Codex CLI vs the API: latency, time to first token and hidden prompts

194 timed runs. Codex CLI took 3.5x as long as the OpenAI API for a one-line answer and sent 19,551 input tokens instead of 17. What a coding CLI adds.

TL;DR

  • For a one-line answer, the Codex CLI took a median 3.5x as long as the OpenAI API with the same model and effort: 3.9 s against 1.1 s.
  • The CLI sent about 19,551 input tokens for that one line. The API sent 17.
  • On a real repair task with 296 behavioral checks, all 9 runs passed. Claude Code with Sonnet 5.5 took a median 15.0 s, the OpenAI API with GPT-6.1 Sol 17.3 s and the Codex CLI with GPT-6.1 Sol 61.2 s.
  • Time to first useful output in our model head-to-head ran from 1.2 s (Fable 5.1 in Claude Code) to 5.3 s (GPT-6.1 Sol high in Codex CLI).
  • All 194 evaluated runs passed their validator. With 3 to 5 runs per cell, these are directional, not rankings.

Every receipt: /benchmarks/cli-model-latency-tokens.

Why we measured the wrapper

Coding agents rarely call a model directly. They call a CLI, such as Claude Code or Codex, which calls the model. The CLI adds a system prompt, tool definitions, start-up time and its own streaming behavior. When people compare "Claude vs GPT" through two CLIs, they often compare two wrappers as much as two models.

So we split the question in two:

  1. Same model, different route. OpenAI API vs Codex CLI, same model, same effort, same prompt.
  2. Same task, different route and model. Claude Code vs Codex CLI vs the API on one repair task.

Same model, same prompt: API vs CLI

  • Total time
  • First useful output
Entrance: medians race at 3× real timeMotion reduced: press Replay to animateThe slowest median is 4.2 s. The clock runs at the recorded speed.
OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6 Luna · none
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

6 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 4.2 s (range 3.8 s–4.7 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 1 s (range 0.7 s–1.5 s, n 5). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 3.8 s (range 3.4 s–4.3 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 0.8 s (range 0.5 s–1.4 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Matched cohort, fixed exact reply, 5 runs per configuration

Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.

Source: Provider explorer receipts: CLI vs API

Live story · 33 sClaude Code CLI vs Codex CLI vs the API: a latency race

Claude Code CLI vs Codex CLI vs the API: a latency race

For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.

Transcript
  1. Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
  2. A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
  3. It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  4. A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
  5. Open benchmarks: intervals, sources and every failure kept.

The task: reply with one fixed line. Five runs per configuration, back to back on the same host.

  • OpenAI API: GPT-6 Luna (no reasoning) 0.97 s, GPT-6.1 Sol low 1.02 s, GPT-6.1 Sol high 1.52 s.
  • Codex CLI: GPT-6 Luna 3.19 s, GPT-6.1 Sol low 4.18 s, GPT-6.1 Sol high 4.19 s.

Time to first useful output follows the same pattern. The API showed first text after 0.82 to 1.34 s. The CLI showed it after 2.79 to 3.79 s.

For a small coding task, the gap grows. The API with GPT-6.1 Sol low took a median 6.0 s, and the Codex CLI took 14.15 s. At high effort it was 9.56 s against 17.85 s. One detail stands out: through the CLI, first useful output came close to the end of the run. For GPT-6.1 Sol high, the first text appeared at 17.27 s of 17.85 s. The API streamed its first text much earlier. If your product shows progress to a person, that difference matters as much as the total.

The hidden prompt

OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6 Luna · none
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

6 rows. Highest Codex CLI · GPT-6.1 Sol · high 19,555 (n 5). Lowest OpenAI API · GPT-6.1 Sol · high 17 (n 5).

Notesn = 5 per row

Reported input tokens, matched cohort

The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.

Source: Provider explorer receipts: CLI vs API

For the same one-line request:

  • OpenAI API: 17 input tokens.
  • Codex CLI: 18,859 to 19,555 input tokens, depending on model and effort.

The CLI wraps every request in its own system prompt and tool context. Part of that is served from the prompt cache, so it is cheaper than its size suggests. It still has to be processed on every call.

Is the token count the whole story? Not quite. We also ran a "wire-matched" cohort, where the API got the same large input the CLI sends, about 8,297 tokens. The API took a median 1.7 s and the CLI 2.76 s. So the extra input explains part of the gap, and CLI start-up and orchestration explain the rest.

The same effect shows up in our model head-to-head. Through the CLI, Codex sent a median 12,124 input tokens per call and Claude Code 2,130. See Claude Haiku vs Sonnet vs Opus vs Fable vs Codex.

A real task: repairing a scheduler

The one-line test shows overhead. A real task shows whether it matters. We gave three routes the same prompt: repair a dependency-aware job scheduler. Medium effort. 296 behavioral checks. Three runs each.

  • Total time
  • First useful output
Entrance: medians race at 44× real time
Claude Code CLI · Sonnet 5.5 · medium
Codex CLI · GPT-6.1 Sol · medium
OpenAI API · GPT-6.1 Sol · medium

3 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · medium 61.2 s (range 59.9 s–69.5 s, n 3). Fastest Claude Code CLI · Sonnet 5.5 · medium 15 s (range 13.9 s–15.9 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · medium 15.6 s (range 13.7 s–23 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · medium 7.5 s (range 6.7 s–9.1 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Same prompt, medium effort, 296 behavioral checks, 3 runs each

All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.

Source: Provider explorer receipts: CLI vs API

All 9 runs passed all 296 checks. Median total time:

  • Claude Code CLI · Sonnet 5.5: 15.0 s (range 13.89 to 15.89 s)
  • OpenAI API · GPT-6.1 Sol: 17.32 s (range 16.28 to 18.61 s)
  • Codex CLI · GPT-6.1 Sol: 61.16 s (range 59.9 to 69.51 s)

First useful output: Claude Code 7.55 s, the API 7.46 s and the Codex CLI 15.56 s.

The OpenAI model through its API was close to Claude Code. The same model through the Codex CLI took several times as long. On this task, the route made a bigger difference than the model.

Output size also differed. Claude Code wrote a median 2,227 output tokens. The Codex CLI wrote 1,181 plus 156 reported reasoning tokens, and the API 1,313 plus 267. The Claude CLI does not report reasoning tokens separately, so a 0 there means "not reported", not "none".

Important: this compares route plus model pairs. Claude Code ran Sonnet 5.5 and the other two ran GPT-6.1 Sol. It is not a pure test of the CLIs.

Time to first useful output across models

How fast does each configuration start talking? This comes from our five-task model head-to-head, through each vendor's CLI.

Entrance: medians race at 3.8× real timeMotion reduced: press Replay to animateThe slowest median is 5.3 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

9 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 5.3 s (range 3.6 s–16.4 s, n 15). Fastest Claude Fable 5.1 · Claude Code 1.2 s (range 1 s–7.9 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Median time to first useful output:

  • Claude Fable 5.1: 1.2 s
  • Claude Sonnet 5.5: 1.56 s
  • Claude Opus 5.5: 1.92 s (low effort 2.39 s, high effort 2.04 s)
  • Claude Haiku 4.5: 3.63 s, slowed by reasoning under the CLI default
  • GPT-6.1 Sol in Codex CLI: 5.05 to 5.32 s across efforts

The whiskers are wide. Haiku's slowest first output took 22.27 s, and the slowest Codex one 17.82 s. Single calls vary a lot, so look at the medians and the ranges together.

What this means for builders

  • If latency matters, measure the route, not just the model. The same GPT-6.1 Sol model took 17.32 s through the API and 61.16 s through the CLI on the same repair task.
  • Watch time to first output, not only total time. A CLI can hold back text until late in the run.
  • Count the hidden prompt. Thousands of tokens of system and tool context ride along with every call. Caching softens the price, not the processing.
  • Effort settings did little on short tasks. Across the CLIs, low and high effort were close on short prompts. The API showed a larger effort effect on the small coding task.

How we measured

  • Receipts. Each run records the route (CLI or API), model, effort, timings, reported tokens and a deterministic validator result.
  • First useful output is the first streamed text that belongs to the answer. Total time runs from launch to exit, including CLI start-up.
  • Counted runs. Only receipts classified "evaluated" count: 194 runs, all of which passed. 18 excluded and 18 diagnostic receipts stay in the raw file but are not charted.
  • Matched cohorts ran every configuration back to back on the same host with the same prompt.

Caveats

  • Small samples. 3 to 5 runs per configuration. Medians with ranges, not confidence intervals.
  • Route and model change together in the scheduler comparison.
  • Claude input not comparable. The Claude CLI receipts record only the uncached remainder of the input (2 tokens), so we do not compare Claude input counts in this study.
  • One host, one network, one day (2026-10-03). Vendor latency changes over the day.
  • Costs. API costs in the raw file are list-price estimates. CLI runs used subscriptions with no per-call price.

Pick the route with data

Agent records the route, time to first output, total time and tokens for every model call it makes. Try Agent and see which route is fastest for your work.

The data behind this post

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.