8 measured metrics · 1 study
GPT-6.1 Sol (Codex CLI)vsGPT-6.1 Sol (OpenAI API)
GPT-6.1 Sol (OpenAI API) ahead on 2; 6 unclear. A side is ahead only where the intervals or ranges do not overlap.
The verdict
GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) share 8 measured metrics from one study. GPT-6.1 Sol (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 1.52 s vs 4.19 s; CLI vs API: time for a one-line answer (First useful output), 1.34 s vs 3.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 6 unclear; each row says why. Every row ran the two sides through different routes (for example Codex CLI vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | GPT-6.1 Sol (Codex CLI) | GPT-6.1 Sol (OpenAI API) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| CLI vs API: time for a one-line answer (Total time) | 4.19 sCodex CLI · effort high · fixed exact reply, 5 runs | 1.52 sOpenAI API · effort high · fixed exact reply, 5 runs | 5 | range: 3.8 s–4.7 s vs 1.4 s–2.2 s | GPT-6.1 Sol (OpenAI API) ahead | The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s). A range is not a confidence interval. Samples are small (5 runs per side). | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a one-line answer (First useful output) | 3.79 sCodex CLI · effort high · fixed exact reply, 5 runs | 1.34 sOpenAI API · effort high · fixed exact reply, 5 runs | 5 | range: 3.4 s–4.3 s vs 1.3 s–2.1 s | GPT-6.1 Sol (OpenAI API) ahead | The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s). A range is not a confidence interval. Samples are small (5 runs per side). | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Hidden prompt: input tokens for the same one-line request | 19,555Codex CLI · effort high · short fixed tasks | 17OpenAI API · effort high · short fixed tasks | 5 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Output tokens to repair the scheduler (Output tokens) | 1,181Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs | 1,313OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs | 3 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
Marks: fastest–slowest run ranges (not intervals)n is shown per side on every row
4 headline metrics as the ratio of the two values. 2 of them separate the sides in the data. Widest ratio: Hidden prompt: input tokens for the same one-line request, 1.2 thousand x (GPT-6.1 Sol (Codex CLI) larger).
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- GPT-6.1 Sol (Codex CLI)
- GPT-6.1 Sol (OpenAI API)
- fastest–slowest run (not an interval)
- where the two overlap
Claude Code CLI vs Codex CLI vs the API: latency and tokens
- CLI vs API: time for a one-line answer (Total time)4.19 sn 51.52 sn 5GPT-6.1 Sol (OpenAI API) aheadCLI vs API: time for a one-line answer (Total time): GPT-6.1 Sol (Codex CLI) 4.19 s (n 5, run range 3.8 s–4.7 s); GPT-6.1 Sol (OpenAI API) 1.52 s (n 5, run range 1.4 s–2.2 s). GPT-6.1 Sol (OpenAI API) ahead.
- CLI vs API: time for a one-line answer (First useful output)3.79 sn 51.34 sn 5GPT-6.1 Sol (OpenAI API) aheadCLI vs API: time for a one-line answer (First useful output): GPT-6.1 Sol (Codex CLI) 3.79 s (n 5, run range 3.4 s–4.3 s); GPT-6.1 Sol (OpenAI API) 1.34 s (n 5, run range 1.3 s–2.1 s). GPT-6.1 Sol (OpenAI API) ahead.
- CLI vs API: time for a small coding task (Total time)17.9 sn 39.56 sn 3UnclearCLI vs API: time for a small coding task (Total time): GPT-6.1 Sol (Codex CLI) 17.9 s (n 3, run range 17.7 s–22.4 s); GPT-6.1 Sol (OpenAI API) 9.56 s (n 3, run range 9.4 s–10.9 s). Unclear.
- CLI vs API: time for a small coding task (First useful output)17.3 sn 35.31 sn 3UnclearCLI vs API: time for a small coding task (First useful output): GPT-6.1 Sol (Codex CLI) 17.3 s (n 3, run range 17.1 s–21.9 s); GPT-6.1 Sol (OpenAI API) 5.31 s (n 3, run range 5 s–6.4 s). Unclear.
- Hidden prompt: input tokens for the same one-line request19,555n 517n 5UnclearHidden prompt: input tokens for the same one-line request: GPT-6.1 Sol (Codex CLI) 19,555 (n 5); GPT-6.1 Sol (OpenAI API) 17 (n 5). Unclear.
- Repairing a scheduler: Claude Code vs Codex vs API (Total time)61.2 sn 317.3 sn 3UnclearRepairing a scheduler: Claude Code vs Codex vs API (Total time): GPT-6.1 Sol (Codex CLI) 61.2 s (n 3, run range 59.9 s–69.5 s); GPT-6.1 Sol (OpenAI API) 17.3 s (n 3, run range 16.3 s–18.6 s). Unclear.
- Repairing a scheduler: Claude Code vs Codex vs API (First useful output)15.6 sn 37.46 sn 3UnclearRepairing a scheduler: Claude Code vs Codex vs API (First useful output): GPT-6.1 Sol (Codex CLI) 15.6 s (n 3, run range 13.7 s–23 s); GPT-6.1 Sol (OpenAI API) 7.46 s (n 3, run range 6.7 s–9.1 s). Unclear.
- Output tokens to repair the scheduler (Output tokens)1,181n 31,313n 3UnclearOutput tokens to repair the scheduler (Output tokens): GPT-6.1 Sol (Codex CLI) 1,181 (n 3); GPT-6.1 Sol (OpenAI API) 1,313 (n 3). Unclear.
| Metric | GPT-6.1 Sol (Codex CLI) | GPT-6.1 Sol (OpenAI API) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| CLI vs API: time for a one-line answer (Total time) | 4.19 sCodex CLI · effort high · fixed exact reply, 5 runs | 1.52 sOpenAI API · effort high · fixed exact reply, 5 runs | 5 | range: 3.8 s–4.7 s vs 1.4 s–2.2 s | GPT-6.1 Sol (OpenAI API) ahead | The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s). A range is not a confidence interval. Samples are small (5 runs per side). | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a one-line answer (First useful output) | 3.79 sCodex CLI · effort high · fixed exact reply, 5 runs | 1.34 sOpenAI API · effort high · fixed exact reply, 5 runs | 5 | range: 3.4 s–4.3 s vs 1.3 s–2.1 s | GPT-6.1 Sol (OpenAI API) ahead | The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s). A range is not a confidence interval. Samples are small (5 runs per side). | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a small coding task (Total time) | 17.9 sCodex CLI · effort high · small coding task, 3 runs | 9.56 sOpenAI API · effort high · small coding task, 3 runs | 3 | range: 17.7 s–22.4 s vs 9.4 s–10.9 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6.1 Sol (OpenAI API) 9.44 s to 10.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a small coding task (First useful output) | 17.3 sCodex CLI · effort high · small coding task, 3 runs | 5.31 sOpenAI API · effort high · small coding task, 3 runs | 3 | range: 17.1 s–21.9 s vs 5 s–6.4 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6.1 Sol (OpenAI API) 4.99 s to 6.42 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Hidden prompt: input tokens for the same one-line request | 19,555Codex CLI · effort high · short fixed tasks | 17OpenAI API · effort high · short fixed tasks | 5 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Repairing a scheduler: Claude Code vs Codex vs API (Total time) | 61.2 sCodex CLI · effort medium · scheduler repair, 296 checks, 3 runs | 17.3 sOpenAI API · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 59.9 s–69.5 s vs 16.3 s–18.6 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 59.9 s to 69.5 s; GPT-6.1 Sol (OpenAI API) 16.3 s to 18.6 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Repairing a scheduler: Claude Code vs Codex vs API (First useful output) | 15.6 sCodex CLI · effort medium · scheduler repair, 296 checks, 3 runs | 7.46 sOpenAI API · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 13.7 s–23 s vs 6.7 s–9.1 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 13.7 s to 23.0 s; GPT-6.1 Sol (OpenAI API) 6.68 s to 9.05 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Output tokens to repair the scheduler (Output tokens) | 1,181Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs | 1,313OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs | 3 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
Marks: fastest–slowest run ranges (not intervals)n is shown per side on every row
8 rows from 1 study. GPT-6.1 Sol (OpenAI API) ahead on 2; 6 unclear. A side is ahead only where the intervals or ranges do not overlap.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick GPT-6.1 Sol (Codex CLI)
No row in this data puts GPT-6.1 Sol (Codex CLI) ahead of GPT-6.1 Sol (OpenAI API). Pick on other grounds (price, access, the tasks you run), or measure your own workload.
When to pick GPT-6.1 Sol (OpenAI API)
- CLI vs API: time for a one-line answer (Total time): 1.52 s vs 4.19 s. The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s). A range is not a confidence interval. Samples are small (5 runs per side).
- CLI vs API: time for a one-line answer (First useful output): 1.34 s vs 3.79 s. The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s). A range is not a confidence interval. Samples are small (5 runs per side).
Side by side
The study charts, showing only these two. Open a study for every configuration.
- Total time
- First useful output
| Item | Total time | First useful output | Range (lowest–highest run) | n |
|---|---|---|---|---|
| OpenAI API · GPT-6.1 Sol · high | 1.5 s | 1.3 s | Total time: 1.4 s–2.2 s; First useful output: 1.3 s–2.1 s | 5 |
| Codex CLI · GPT-6.1 Sol · high | 4.2 s | 3.8 s | Total time: 3.8 s–4.7 s; First useful output: 3.4 s–4.3 s | 5 |
2 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 4.2 s (range 3.8 s–4.7 s, n 5). Fastest OpenAI API · GPT-6.1 Sol · high 1.5 s (range 1.4 s–2.2 s, n 5). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 3.8 s (range 3.4 s–4.3 s, n 5). Fastest OpenAI API · GPT-6.1 Sol · high 1.3 s (range 1.3 s–2.1 s, n 5). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 5 per row
Matched cohort, fixed exact reply, 5 runs per configuration
Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.
Source: Provider explorer receipts: CLI vs API
- Total time
- First useful output
| Item | Total time | First useful output | Range (lowest–highest run) | n |
|---|---|---|---|---|
| OpenAI API · GPT-6.1 Sol · high | 9.6 s | 5.3 s | Total time: 9.4 s–10.9 s; First useful output: 5 s–6.4 s | 3 |
| Codex CLI · GPT-6.1 Sol · high | 17.9 s | 17.3 s | Total time: 17.7 s–22.4 s; First useful output: 17.1 s–21.9 s | 3 |
2 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 17.9 s (range 17.7 s–22.4 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · high 9.6 s (range 9.4 s–10.9 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 17.3 s (range 17.1 s–21.9 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · high 5.3 s (range 5 s–6.4 s, n 3). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Matched cohort, small coding task, 3 runs per configuration
Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.
Source: Provider explorer receipts: CLI vs API
| Item | Input tokens | n |
|---|---|---|
| OpenAI API · GPT-6.1 Sol · low | 17 | 5 |
| Codex CLI · GPT-6.1 Sol · high | 19,555 | 5 |
2 rows. Highest Codex CLI · GPT-6.1 Sol · high 19,555 (n 5). Lowest OpenAI API · GPT-6.1 Sol · low 17 (n 5).
Notesn = 5 per row
Reported input tokens, matched cohort
The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.
Source: Provider explorer receipts: CLI vs API
- Total time
- First useful output
| Item | Total time | First useful output | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Codex CLI · GPT-6.1 Sol · medium | 61.2 s | 15.6 s | Total time: 59.9 s–69.5 s; First useful output: 13.7 s–23 s | 3 |
| OpenAI API · GPT-6.1 Sol · medium | 17.3 s | 7.5 s | Total time: 16.3 s–18.6 s; First useful output: 6.7 s–9.1 s | 3 |
2 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · medium 61.2 s (range 59.9 s–69.5 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · medium 17.3 s (range 16.3 s–18.6 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · medium 15.6 s (range 13.7 s–23 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · medium 7.5 s (range 6.7 s–9.1 s, n 3). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Same prompt, medium effort, 296 behavioral checks, 3 runs each
All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.
Source: Provider explorer receipts: CLI vs API
- Output tokens
- of which reasoning tokens (reported) (inner bar)
| Item | Output tokens | Reasoning tokens (reported) | n |
|---|---|---|---|
| Codex CLI · GPT-6.1 Sol · medium | 1,181 | 156 | 3 |
| OpenAI API · GPT-6.1 Sol · medium | 1,313 | 267 | 3 |
2 rows, 2 series: Output tokens, Reasoning tokens (reported). Output tokens: highest OpenAI API · GPT-6.1 Sol · medium 1,313 (n 3). Lowest Codex CLI · GPT-6.1 Sol · medium 1,181 (n 3). Reasoning tokens (reported): highest OpenAI API · GPT-6.1 Sol · medium 267 (n 3). Lowest Codex CLI · GPT-6.1 Sol · medium 156 (n 3).
Notesn = 3 per row
Median per run; reasoning tokens shown separately where reported
The Claude CLI does not report reasoning tokens separately; 0 there means "not reported", not "none".
Source: Provider explorer receipts: CLI vs API
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, GPT-6.1 Sol (Codex CLI) or GPT-6.1 Sol (OpenAI API)?
- GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) share 8 measured metrics from one study. GPT-6.1 Sol (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 1.52 s vs 4.19 s; CLI vs API: time for a one-line answer (First useful output), 1.34 s vs 3.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 6 unclear; each row says why. Every row ran the two sides through different routes (for example Codex CLI vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
- How were GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) measured?
- They share 8 measured metrics from 1 public study: Claude Code CLI vs Codex CLI vs the API: latency and tokens. Every row names its configuration, its sample size and its interval or range.
- How do GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) compare on cLI vs API: time for a one-line answer (Total time)?
- GPT-6.1 Sol (Codex CLI): 4.19 s (Codex CLI · effort high · fixed exact reply, 5 runs; n = 5; run range 3.8 s to 4.7 s). GPT-6.1 Sol (OpenAI API): 1.52 s (OpenAI API · effort high · fixed exact reply, 5 runs; n = 5; run range 1.4 s to 2.2 s). The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s). A range is not a confidence interval. Samples are small (5 runs per side).
- How do GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) compare on cLI vs API: time for a one-line answer (First useful output)?
- GPT-6.1 Sol (Codex CLI): 3.79 s (Codex CLI · effort high · fixed exact reply, 5 runs; n = 5; run range 3.4 s to 4.3 s). GPT-6.1 Sol (OpenAI API): 1.34 s (OpenAI API · effort high · fixed exact reply, 5 runs; n = 5; run range 1.3 s to 2.1 s). The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s). A range is not a confidence interval. Samples are small (5 runs per side).
- How do GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) compare on cLI vs API: time for a small coding task (Total time)?
- GPT-6.1 Sol (Codex CLI): 17.9 s (Codex CLI · effort high · small coding task, 3 runs; n = 3; run range 17.7 s to 22.4 s). GPT-6.1 Sol (OpenAI API): 9.56 s (OpenAI API · effort high · small coding task, 3 runs; n = 3; run range 9.4 s to 10.9 s). Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6.1 Sol (OpenAI API) 9.44 s to 10.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.
- How do GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) compare on cLI vs API: time for a small coding task (First useful output)?
- GPT-6.1 Sol (Codex CLI): 17.3 s (Codex CLI · effort high · small coding task, 3 runs; n = 3; run range 17.1 s to 21.9 s). GPT-6.1 Sol (OpenAI API): 5.31 s (OpenAI API · effort high · small coding task, 3 runs; n = 3; run range 5 s to 6.4 s). Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6.1 Sol (OpenAI API) 4.99 s to 6.42 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.
- How do GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) compare on repairing a scheduler: Claude Code vs Codex vs API (Total time)?
- GPT-6.1 Sol (Codex CLI): 61.2 s (Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs; n = 3; run range 59.9 s to 69.5 s). GPT-6.1 Sol (OpenAI API): 17.3 s (OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs; n = 3; run range 16.3 s to 18.6 s). Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 59.9 s to 69.5 s; GPT-6.1 Sol (OpenAI API) 16.3 s to 18.6 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.
The studies behind this page
Claude Code CLI vs Codex CLI vs the API: latency and tokens
194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.