6 measured metrics · 2 calculated · 2 studies
GPT-6.1 Sol (Codex CLI)vsGPT-6 Luna (Codex CLI)
No row separates them: 8 unclear.
The verdict
GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) share 6 measured metrics and 2 list-price calculations from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 8 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | GPT-6.1 Sol (Codex CLI) | GPT-6 Luna (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| CLI vs API: time for a one-line answer (Total time) | 4.19 sCodex CLI · effort high · fixed exact reply, 5 runs | 3.19 sCodex CLI · effort none · fixed exact reply, 5 runs | 5 | range: 3.8 s–4.7 s vs 2.9 s–3.8 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6 Luna (Codex CLI) 2.88 s to 3.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a one-line answer (First useful output) | 3.79 sCodex CLI · effort high · fixed exact reply, 5 runs | 2.79 sCodex CLI · effort none · fixed exact reply, 5 runs | 5 | range: 3.4 s–4.3 s vs 2.5 s–3.4 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6 Luna (Codex CLI) 2.46 s to 3.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Hidden prompt: input tokens for the same one-line request | 19,555Codex CLI · effort high · short fixed tasks | 18,859Codex CLI · effort none · short fixed tasks | 5 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
Marks: fastest–slowest run ranges (not intervals)n is shown per side on every row
3 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: CLI vs API: time for a one-line answer (First useful output), 1.4x (GPT-6.1 Sol (Codex CLI) larger).
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- GPT-6.1 Sol (Codex CLI)
- GPT-6 Luna (Codex CLI)
- fastest–slowest run (not an interval)
- where the two overlap
- hollow: list-price calculation
Claude Code CLI vs Codex CLI vs the API: latency and tokens
- CLI vs API: time for a one-line answer (Total time)4.19 sn 53.19 sn 5UnclearCLI vs API: time for a one-line answer (Total time): GPT-6.1 Sol (Codex CLI) 4.19 s (n 5, run range 3.8 s–4.7 s); GPT-6 Luna (Codex CLI) 3.19 s (n 5, run range 2.9 s–3.8 s). Unclear.
- CLI vs API: time for a one-line answer (First useful output)3.79 sn 52.79 sn 5UnclearCLI vs API: time for a one-line answer (First useful output): GPT-6.1 Sol (Codex CLI) 3.79 s (n 5, run range 3.4 s–4.3 s); GPT-6 Luna (Codex CLI) 2.79 s (n 5, run range 2.5 s–3.4 s). Unclear.
- CLI vs API: time for a small coding task (Total time)17.9 sn 39.23 sn 3UnclearCLI vs API: time for a small coding task (Total time): GPT-6.1 Sol (Codex CLI) 17.9 s (n 3, run range 17.7 s–22.4 s); GPT-6 Luna (Codex CLI) 9.23 s (n 3, run range 9 s–11.7 s). Unclear.
- CLI vs API: time for a small coding task (First useful output)17.3 sn 38.68 sn 3UnclearCLI vs API: time for a small coding task (First useful output): GPT-6.1 Sol (Codex CLI) 17.3 s (n 3, run range 17.1 s–21.9 s); GPT-6 Luna (Codex CLI) 8.68 s (n 3, run range 8.3 s–11 s). Unclear.
- Hidden prompt: input tokens for the same one-line request19,555n 518,859n 5UnclearHidden prompt: input tokens for the same one-line request: GPT-6.1 Sol (Codex CLI) 19,555 (n 5); GPT-6 Luna (Codex CLI) 18,859 (n 5). Unclear.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
- Time to first text: a 250-line answer, six models3.52 sn 43.30 sn 4UnclearTime to first text: a 250-line answer, six models: GPT-6.1 Sol (Codex CLI) 3.52 s (n 4, run range 2.8 s–4.4 s); GPT-6 Luna (Codex CLI) 3.30 s (n 4, run range 3.2 s–3.5 s). Unclear.
- Output speed after the first text: visible tokens per second (calculation)Calculation80n 4129n 4UnclearOutput speed after the first text: visible tokens per second (calculation), calculation: GPT-6.1 Sol (Codex CLI) 80 (n 4, run range 72–81); GPT-6 Luna (Codex CLI) 129 (n 4, run range 56–259). Unclear.
- Output speed in characters per second after the first text (calculation)Calculation323n 4524n 4UnclearOutput speed in characters per second after the first text (calculation), calculation: GPT-6.1 Sol (Codex CLI) 323 (n 4, run range 291–327); GPT-6 Luna (Codex CLI) 524 (n 4, run range 225–1,052). Unclear.
| Metric | GPT-6.1 Sol (Codex CLI) | GPT-6 Luna (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| CLI vs API: time for a one-line answer (Total time) | 4.19 sCodex CLI · effort high · fixed exact reply, 5 runs | 3.19 sCodex CLI · effort none · fixed exact reply, 5 runs | 5 | range: 3.8 s–4.7 s vs 2.9 s–3.8 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6 Luna (Codex CLI) 2.88 s to 3.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a one-line answer (First useful output) | 3.79 sCodex CLI · effort high · fixed exact reply, 5 runs | 2.79 sCodex CLI · effort none · fixed exact reply, 5 runs | 5 | range: 3.4 s–4.3 s vs 2.5 s–3.4 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6 Luna (Codex CLI) 2.46 s to 3.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a small coding task (Total time) | 17.9 sCodex CLI · effort high · small coding task, 3 runs | 9.23 sCodex CLI · effort none · small coding task, 3 runs | 3 | range: 17.7 s–22.4 s vs 9 s–11.7 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6 Luna (Codex CLI) 8.99 s to 11.7 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a small coding task (First useful output) | 17.3 sCodex CLI · effort high · small coding task, 3 runs | 8.68 sCodex CLI · effort none · small coding task, 3 runs | 3 | range: 17.1 s–21.9 s vs 8.3 s–11 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6 Luna (Codex CLI) 8.27 s to 11.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Hidden prompt: input tokens for the same one-line request | 19,555Codex CLI · effort high · short fixed tasks | 18,859Codex CLI · effort none · short fixed tasks | 5 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Time to first text: a 250-line answer, six models | 3.52 sCodex CLI · effort low | 3.30 sCodex CLI · effort low | 4 | range: 2.8 s–4.4 s vs 3.2 s–3.5 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 80Codex CLI · effort low | 129Codex CLI · effort low | 4 | range: 72–81 vs 56–259 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 323Codex CLI · effort low | 524Codex CLI · effort low | 4 | range: 291–327 vs 225–1,052 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
Marks: fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
8 rows from 2 studies. No row separates them: 8 unclear.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick GPT-6.1 Sol (Codex CLI)
No row in this data puts GPT-6.1 Sol (Codex CLI) ahead of GPT-6 Luna (Codex CLI). Pick on other grounds (price, access, the tasks you run), or measure your own workload.
When to pick GPT-6 Luna (Codex CLI)
No row in this data puts GPT-6 Luna (Codex CLI) ahead of GPT-6.1 Sol (Codex CLI). Pick on other grounds (price, access, the tasks you run), or measure your own workload.
Side by side
The study charts, showing only these two. Open a study for every configuration.
- Total time
- First useful output
| Item | Total time | First useful output | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Codex CLI · GPT-6 Luna · none | 3.2 s | 2.8 s | Total time: 2.9 s–3.8 s; First useful output: 2.5 s–3.4 s | 5 |
| Codex CLI · GPT-6.1 Sol · high | 4.2 s | 3.8 s | Total time: 3.8 s–4.7 s; First useful output: 3.4 s–4.3 s | 5 |
2 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 4.2 s (range 3.8 s–4.7 s, n 5). Fastest Codex CLI · GPT-6 Luna · none 3.2 s (range 2.9 s–3.8 s, n 5). All run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 3.8 s (range 3.4 s–4.3 s, n 5). Fastest Codex CLI · GPT-6 Luna · none 2.8 s (range 2.5 s–3.4 s, n 5). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 5 per row
Matched cohort, fixed exact reply, 5 runs per configuration
Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.
Source: Provider explorer receipts: CLI vs API
- Total time
- First useful output
| Item | Total time | First useful output | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Codex CLI · GPT-6 Luna · none | 9.2 s | 8.7 s | Total time: 9 s–11.7 s; First useful output: 8.3 s–11 s | 3 |
| Codex CLI · GPT-6.1 Sol · high | 17.9 s | 17.3 s | Total time: 17.7 s–22.4 s; First useful output: 17.1 s–21.9 s | 3 |
2 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 17.9 s (range 17.7 s–22.4 s, n 3). Fastest Codex CLI · GPT-6 Luna · none 9.2 s (range 9 s–11.7 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 17.3 s (range 17.1 s–21.9 s, n 3). Fastest Codex CLI · GPT-6 Luna · none 8.7 s (range 8.3 s–11 s, n 3). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Matched cohort, small coding task, 3 runs per configuration
Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.
Source: Provider explorer receipts: CLI vs API
| Item | Input tokens | n |
|---|---|---|
| Codex CLI · GPT-6 Luna · none | 18,859 | 5 |
| Codex CLI · GPT-6.1 Sol · high | 19,555 | 5 |
2 rows. Highest Codex CLI · GPT-6.1 Sol · high 19,555 (n 5). Lowest Codex CLI · GPT-6 Luna · none 18,859 (n 5).
Notesn = 5 per row
Reported input tokens, matched cohort
The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.
Source: Provider explorer receipts: CLI vs API
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (low) · Codex CLI | 3.5 s | 2.8 s–4.4 s | 4 |
| GPT-6 Luna (low) · Codex CLI | 3.3 s | 3.2 s–3.5 s | 4 |
2 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 3.5 s (range 2.8 s–4.4 s, n 4). Fastest GPT-6 Luna (low) · Codex CLI 3.3 s (range 3.2 s–3.5 s, n 4). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
| Item | Visible tokens per second | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (low) · Codex CLI | 80 | 72–81 | 4 |
| GPT-6 Luna (low) · Codex CLI | 129 | 56–259 | 4 |
List-price calculation, not a run. 2 rows. Highest GPT-6 Luna (low) · Codex CLI 129 (range 56–259, n 4). Lowest GPT-6.1 Sol (low) · Codex CLI 80 (range 72–81, n 4). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = slowest and fastest call
Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.
Source: LLM speed anatomy
| Item | Characters per second | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (low) · Codex CLI | 323 | 291–327 | 4 |
| GPT-6 Luna (low) · Codex CLI | 524 | 225–1,052 | 4 |
List-price calculation, not a run. 2 rows. Highest GPT-6 Luna (low) · Codex CLI 524 (range 225–1,052, n 4). Lowest GPT-6.1 Sol (low) · Codex CLI 323 (range 291–327, n 4). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call
Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, GPT-6.1 Sol (Codex CLI) or GPT-6 Luna (Codex CLI)?
- GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) share 6 measured metrics and 2 list-price calculations from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 8 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).
- How were GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) measured?
- They share 6 measured metrics and 2 list-price calculations from 2 public studies: Claude Code CLI vs Codex CLI vs the API: latency and tokens; Where the seconds go: first text, output speed and prompt size for 6 LLMs. Every row names its configuration, its sample size and its interval or range.
- How do GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) compare on cLI vs API: time for a one-line answer (Total time)?
- GPT-6.1 Sol (Codex CLI): 4.19 s (Codex CLI · effort high · fixed exact reply, 5 runs; n = 5; run range 3.8 s to 4.7 s). GPT-6 Luna (Codex CLI): 3.19 s (Codex CLI · effort none · fixed exact reply, 5 runs; n = 5; run range 2.9 s to 3.8 s). The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6 Luna (Codex CLI) 2.88 s to 3.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How do GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) compare on cLI vs API: time for a one-line answer (First useful output)?
- GPT-6.1 Sol (Codex CLI): 3.79 s (Codex CLI · effort high · fixed exact reply, 5 runs; n = 5; run range 3.4 s to 4.3 s). GPT-6 Luna (Codex CLI): 2.79 s (Codex CLI · effort none · fixed exact reply, 5 runs; n = 5; run range 2.5 s to 3.4 s). The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6 Luna (Codex CLI) 2.46 s to 3.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How do GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) compare on cLI vs API: time for a small coding task (Total time)?
- GPT-6.1 Sol (Codex CLI): 17.9 s (Codex CLI · effort high · small coding task, 3 runs; n = 3; run range 17.7 s to 22.4 s). GPT-6 Luna (Codex CLI): 9.23 s (Codex CLI · effort none · small coding task, 3 runs; n = 3; run range 9 s to 11.7 s). Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6 Luna (Codex CLI) 8.99 s to 11.7 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.
- How do GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) compare on cLI vs API: time for a small coding task (First useful output)?
- GPT-6.1 Sol (Codex CLI): 17.3 s (Codex CLI · effort high · small coding task, 3 runs; n = 3; run range 17.1 s to 21.9 s). GPT-6 Luna (Codex CLI): 8.68 s (Codex CLI · effort none · small coding task, 3 runs; n = 3; run range 8.3 s to 11 s). Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6 Luna (Codex CLI) 8.27 s to 11.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.
- How do GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) compare on time to first text: a 250-line answer, six models?
- GPT-6.1 Sol (Codex CLI): 3.52 s (Codex CLI · effort low; n = 4; run range 2.8 s to 4.4 s). GPT-6 Luna (Codex CLI): 3.30 s (Codex CLI · effort low; n = 4; run range 3.2 s to 3.5 s). The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
The studies behind this page
Claude Code CLI vs Codex CLI vs the API: latency and tokens
194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.