• Claude Code
  • Codex
  • Latency
  • Tokens
  • CLI vs API

Claude Code CLI vs Codex CLI vs the API: latency and tokens

How much time and how many tokens does a coding CLI add on top of the model, and how do Claude Code and Codex compare on the same repair task?

Published · Updated · 5 charts · Download the data or a carousel

100%

95% CI 98%–100% · n = 194

194/194 · Evaluated runs that passed their validator

18 excluded and 18 diagnostic receipts are not counted.

The answer

For a one-line answer, the Codex CLI took a median 3.5 times as long as the OpenAI API with the same model and effort, and it sent about 19,551 input tokens instead of 17. On a dependency-aware scheduler repair with 296 checks, all 9 runs passed: Claude Code with Sonnet 5.5 took a median 15.0 s, the OpenAI API with GPT-6.1 Sol 17.3 s and the Codex CLI with GPT-6.1 Sol 61.2 s. With 3 to 5 runs per cell these are directional measurements, not rankings.

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 33 sClaude Code CLI vs Codex CLI vs the API: a latency race

Claude Code CLI vs Codex CLI vs the API: a latency race

For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.

Transcript
  1. Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
  2. A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
  3. It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  4. A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
  5. Open benchmarks: intervals, sources and every failure kept.

Key numbers

3.5x

Codex CLI vs OpenAI API, median total time for a one-line answer

slower · n = 30

15.0s

Claude Code CLI (Sonnet 5.5) median time to repair the scheduler

n = 3

61.2s

Codex CLI (GPT-6.1 Sol) median time to repair the scheduler

n = 3

19,551

Median input tokens the Codex CLI sends for a one-line request

n = 15

Estimate what serial waiting costs your team

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

  • Total time
  • First useful output
Entrance: medians race at 3× real timeMotion reduced: press Replay to animateThe slowest median is 4.2 s. The clock runs at the recorded speed.
OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6 Luna · none
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

6 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 4.2 s (range 3.8 s–4.7 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 1 s (range 0.7 s–1.5 s, n 5). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 3.8 s (range 3.4 s–4.3 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 0.8 s (range 0.5 s–1.4 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Matched cohort, fixed exact reply, 5 runs per configuration

Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.

Source: Provider explorer receipts: CLI vs API

Share card (PNG)
  • Total time
  • First useful output
OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
Codex CLI · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

Seconds · log scale: each gridline is 10 times the one before

6 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 17.9 s (range 17.7 s–22.4 s, n 3). Fastest OpenAI API · GPT-6 Luna · none 4 s (range 3.8 s–4.4 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 17.3 s (range 17.1 s–21.9 s, n 3). Fastest OpenAI API · GPT-6 Luna · none 0.7 s (range 0.6 s–0.8 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Matched cohort, small coding task, 3 runs per configuration

Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.

Source: Provider explorer receipts: CLI vs API

Share card (PNG)
OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6 Luna · none
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

6 rows. Highest Codex CLI · GPT-6.1 Sol · high 19,555 (n 5). Lowest OpenAI API · GPT-6.1 Sol · high 17 (n 5).

Notesn = 5 per row

Reported input tokens, matched cohort

The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.

Source: Provider explorer receipts: CLI vs API

Share card (PNG)
  • Total time
  • First useful output
Entrance: medians race at 44× real time
Claude Code CLI · Sonnet 5.5 · medium
Codex CLI · GPT-6.1 Sol · medium
OpenAI API · GPT-6.1 Sol · medium

3 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · medium 61.2 s (range 59.9 s–69.5 s, n 3). Fastest Claude Code CLI · Sonnet 5.5 · medium 15 s (range 13.9 s–15.9 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · medium 15.6 s (range 13.7 s–23 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · medium 7.5 s (range 6.7 s–9.1 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Same prompt, medium effort, 296 behavioral checks, 3 runs each

All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.

Source: Provider explorer receipts: CLI vs API

Share card (PNG)
  • Output tokens
  • of which reasoning tokens (reported) (inner bar)
Claude Code CLI · Sonnet 5.5 · medium
Codex CLI · GPT-6.1 Sol · medium
OpenAI API · GPT-6.1 Sol · medium

3 rows, 2 series: Output tokens, Reasoning tokens (reported). Output tokens: highest Claude Code CLI · Sonnet 5.5 · medium 2,227 (n 3). Lowest Codex CLI · GPT-6.1 Sol · medium 1,181 (n 3). Reasoning tokens (reported): highest OpenAI API · GPT-6.1 Sol · medium 267 (n 3). Lowest Claude Code CLI · Sonnet 5.5 · medium 0 (n 0).

Notesn 0–3 per row

Median per run; reasoning tokens shown separately where reported

The Claude CLI does not report reasoning tokens separately; 0 there means "not reported", not "none".

Source: Provider explorer receipts: CLI vs API

Share card (PNG)

Tables

Every evaluated configuration

TaskRoute · model · effortCohortRunsPassedMedian first useful (s)Median total (s)Median input tokensMedian output tokens
Fixed 243-token answerOpenAI API · GPT-6.1 Sol · low · flexflex-warm-balanced661.2 s3.3 s279243
Fixed 243-token answerOpenAI API · GPT-6.1 Sol · low · flexflex-initial331.6 s3.7 s279243
Fixed 243-token answerOpenAI API · GPT-6.1 Sol · lowflex-warm-balanced661.1 s3.9 s279243
Fixed 243-token answerOpenAI API · GPT-6.1 Sol · lowflex-initial331.3 s4.1 s279243
Fixed 243-token answerOpenAI API · GPT-6.1 Sol · lowbare-api331.4 s4.1 s279243
Fixed 243-token answerCodex CLI · GPT-6.1 Sol · lowfollowup-stream12122.5 s7.9 s8,941243
Fixed 243-token answerCodex CLI · GPT-6.1 Sol · lownative44—8.9 s4,945243
Fixed 243-token answerCodex CLI · GPT-6.1 Sol · lowfollowup-minimal662.8 s9.3 s4,945243
Fixed exact replyOpenAI API · GPT-6 Luna · nonematched550.8 s1 s177
Fixed exact replyOpenAI API · GPT-6.1 Sol · lowmatched550.9 s1 s177
Fixed exact replyOpenAI API · GPT-6.1 Sol · lowparity550.9 s1.1 s1667
Fixed exact replyOpenAI API · GPT-6.1 Sol · highmatched551.3 s1.5 s1717
Fixed exact replyOpenAI API · GPT-6.1 Sol · lowwire-matched331.6 s1.7 s8,2977
Fixed exact replyOpenAI API · GPT-6.1 Sol · highapi-initial111.8 s1.9 s1717
Fixed exact replyCodex CLI · GPT-6.1 Sol · lowwire-matched332.5 s2.8 s8,2977
Fixed exact replyCodex CLI · GPT-6.1 Sol · lowparity10102.8 s3.1 s8,2967
Fixed exact replyCodex CLI · GPT-6 Luna · nonematched552.8 s3.2 s18,8597
Fixed exact replyCodex CLI · GPT-6.1 Sol · lowinstalled662.6 s3.2 s10,9937
Fixed exact replyCodex CLI · GPT-6.1 Sol · lowfollowup-stream12122.8 s3.4 s8,6787
Fixed exact replyCodex CLI · GPT-6.1 Sol · lowtuning663.5 s3.9 s10,9907
Fixed exact replyCodex CLI · GPT-6.1 Sol · lowmatched553.8 s4.2 s19,5517
Fixed exact replyCodex CLI · GPT-6.1 Sol · highmatched553.8 s4.2 s19,5557
mergeRanges coding repairOpenAI API · GPT-6.1 Sol · low · flexflex-initial113.3 s5.6 s206289
mergeRanges coding repairOpenAI API · GPT-6.1 Sol · lowflex-initial113.2 s6 s206300
mergeRanges coding repairCodex CLI · GPT-6.1 Sol · lowfollowup-minimal553.3 s6.1 s4,872241
mergeRanges coding repairCodex CLI · GPT-6.1 Sol · lowfollowup-stream663.3 s9.2 s6,549256
PlanJobs scheduler repairClaude Code CLI · Sonnet 5.5 · mediumclaude-scheduler337.6 s15 s22,227
PlanJobs scheduler repairOpenAI API · GPT-6.1 Sol · mediumscheduler337.5 s17.3 s9,5631,313
PlanJobs scheduler repairCodex CLI · GPT-6.1 Sol · mediumscheduler3315.6 s61.2 s9,5631,181
Scratch file edit probeCodex CLI · GPT-6.1 Sol · lowedit-check11—8.3 s35,281323
Scratch file edit probeCodex CLI · GPT-6.1 Sol · lowinstalled-live22—14 s18,144348
Scratch file edit probeCodex CLI · GPT-6.1 Sol · lownative22—15.7 s18,211407
Synthetic coding task (API prompt)OpenAI API · GPT-6.1 Sol · lowparity330.8 s3.6 s363238
Synthetic coding task (API prompt)OpenAI API · GPT-6 Luna · nonematched330.7 s4 s382226
Synthetic coding task (API prompt)OpenAI API · GPT-6.1 Sol · lowwire-matched331.5 s4 s8,496225
Synthetic coding task (API prompt)OpenAI API · GPT-6.1 Sol · lowmatched331.1 s6 s382238
Synthetic coding task (API prompt)OpenAI API · GPT-6.1 Sol · highmatched335.3 s9.6 s382429
Synthetic coding task (CLI matched prompt)Codex CLI · GPT-6 Luna · nonematched338.7 s9.2 s38,110285
Synthetic coding task (CLI matched prompt)Codex CLI · GPT-6 Luna · nonetuning339.3 s9.9 s18,244261
Synthetic coding task (CLI matched prompt)Codex CLI · GPT-6.1 Sol · lowinstalled669.9 s11 s22,391251
Synthetic coding task (CLI matched prompt)Codex CLI · GPT-6.1 Sol · lowtuning6613.1 s13.7 s22,398260
Synthetic coding task (CLI matched prompt)Codex CLI · GPT-6.1 Sol · lowmatched3313.6 s14.2 s39,529266
Synthetic coding task (CLI matched prompt)Codex CLI · GPT-6.1 Sol · highmatched3317.3 s17.9 s39,527405
Synthetic coding task (CLI parity prompt)Codex CLI · GPT-6.1 Sol · lowwire-matched333.8 s4.1 s8,496210
Synthetic coding task (CLI parity prompt)Codex CLI · GPT-6.1 Sol · lowparity667.1 s7.5 s8,495210

Method

  1. Receipts from the provider explorer: each run records the route (CLI or API), the model, the effort, timings, reported tokens and a deterministic validator result.
  2. First useful output is the first streamed text that belongs to the answer. Total time runs from launch to exit, including CLI start-up.
  3. Only receipts classified "evaluated" count. Excluded runs (context mismatch, pilot runs, unsupported controls) and diagnostics are kept in the raw file but not charted.
  4. The matched cohort ran every configuration back to back on the same host with the same prompt.

Caveats

  • Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  • The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
  • The Claude CLI receipts record only the uncached remainder of the input (2 tokens), so the Claude input column is not comparable.
  • All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
  • API costs in the raw file are list-price estimates; CLI runs are subscription calls with no per-call price.

Sources

  • Provider explorer receipts: CLI vs API

    Our recorded runs ·

    230 imported receipts for short fixed tasks over Claude Code CLI, Codex CLI and the OpenAI API, with time to first useful output, total time, tokens and validation.

    Raw data: provider-explorer/runs.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Claude Code CLI vs Codex CLI vs the API: latency and tokens”, updated October 5, 2026, https://agent.sasid.ai/benchmarks/cli-model-latency-tokens.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

More comparisons based on this study (3)

These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.

More write-ups that cite this study (2)

Models and comparisons in this study

More studies

All benchmarks
Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.