5 measured metrics · 1 study
GPT-6.1 Sol (OpenAI API)lowvshigheffort
No row separates them: 1 tie, 4 unclear.
The verdict
GPT-6.1 Sol (OpenAI API) at low effort and GPT-6.1 Sol (OpenAI API) at high effort share 5 measured metrics from one study. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 4 unclear; each row says why. Some rows rest on small samples (n = 3 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | GPT-6.1 Sol (OpenAI API) at low effort | GPT-6.1 Sol (OpenAI API) at high effort | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| CLI vs API: time for a one-line answer (Total time) | 1.02 sOpenAI API · effort low · fixed exact reply, 5 runs | 1.52 sOpenAI API · effort high · fixed exact reply, 5 runs | 5 | range: 1 s–1.9 s vs 1.4 s–2.2 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.96 s to 1.87 s; GPT-6.1 Sol (OpenAI API) at high effort 1.35 s to 2.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a one-line answer (First useful output) | 0.87 sOpenAI API · effort low · fixed exact reply, 5 runs | 1.34 sOpenAI API · effort high · fixed exact reply, 5 runs | 5 | range: 0.8 s–1.7 s vs 1.3 s–2.1 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.84 s to 1.74 s; GPT-6.1 Sol (OpenAI API) at high effort 1.26 s to 2.12 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Hidden prompt: input tokens for the same one-line request | 17OpenAI API · effort low · short fixed tasks | 17OpenAI API · effort high · short fixed tasks | 5 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
Marks: fastest–slowest run ranges (not intervals)n is shown per side on every row
3 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: CLI vs API: time for a one-line answer (First useful output), 1.5x (High effort larger).
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Low effort
- High effort
- fastest–slowest run (not an interval)
- where the two overlap
Claude Code CLI vs Codex CLI vs the API: latency and tokens
- CLI vs API: time for a one-line answer (Total time)1.02 sn 51.52 sn 5UnclearCLI vs API: time for a one-line answer (Total time): Low effort 1.02 s (n 5, run range 1 s–1.9 s); High effort 1.52 s (n 5, run range 1.4 s–2.2 s). Unclear.
- CLI vs API: time for a one-line answer (First useful output)0.87 sn 51.34 sn 5UnclearCLI vs API: time for a one-line answer (First useful output): Low effort 0.87 s (n 5, run range 0.8 s–1.7 s); High effort 1.34 s (n 5, run range 1.3 s–2.1 s). Unclear.
- CLI vs API: time for a small coding task (Total time)6.00 sn 39.56 sn 3UnclearCLI vs API: time for a small coding task (Total time): Low effort 6.00 s (n 3, run range 5.4 s–6.2 s); High effort 9.56 s (n 3, run range 9.4 s–10.9 s). Unclear.
- CLI vs API: time for a small coding task (First useful output)1.05 sn 35.31 sn 3UnclearCLI vs API: time for a small coding task (First useful output): Low effort 1.05 s (n 3, run range 1 s–1.4 s); High effort 5.31 s (n 3, run range 5 s–6.4 s). Unclear.
- Hidden prompt: input tokens for the same one-line request17n 517n 5TieHidden prompt: input tokens for the same one-line request: Low effort 17 (n 5); High effort 17 (n 5). Tie.
| Metric | GPT-6.1 Sol (OpenAI API) at low effort | GPT-6.1 Sol (OpenAI API) at high effort | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| CLI vs API: time for a one-line answer (Total time) | 1.02 sOpenAI API · effort low · fixed exact reply, 5 runs | 1.52 sOpenAI API · effort high · fixed exact reply, 5 runs | 5 | range: 1 s–1.9 s vs 1.4 s–2.2 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.96 s to 1.87 s; GPT-6.1 Sol (OpenAI API) at high effort 1.35 s to 2.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a one-line answer (First useful output) | 0.87 sOpenAI API · effort low · fixed exact reply, 5 runs | 1.34 sOpenAI API · effort high · fixed exact reply, 5 runs | 5 | range: 0.8 s–1.7 s vs 1.3 s–2.1 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.84 s to 1.74 s; GPT-6.1 Sol (OpenAI API) at high effort 1.26 s to 2.12 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a small coding task (Total time) | 6.00 sOpenAI API · effort low · small coding task, 3 runs | 9.56 sOpenAI API · effort high · small coding task, 3 runs | 3 | range: 5.4 s–6.2 s vs 9.4 s–10.9 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) at low effort 5.44 s to 6.20 s; GPT-6.1 Sol (OpenAI API) at high effort 9.44 s to 10.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| CLI vs API: time for a small coding task (First useful output) | 1.05 sOpenAI API · effort low · small coding task, 3 runs | 5.31 sOpenAI API · effort high · small coding task, 3 runs | 3 | range: 1 s–1.4 s vs 5 s–6.4 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.97 s to 1.40 s; GPT-6.1 Sol (OpenAI API) at high effort 4.99 s to 6.42 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Hidden prompt: input tokens for the same one-line request | 17OpenAI API · effort low · short fixed tasks | 17OpenAI API · effort high · short fixed tasks | 5 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
Marks: fastest–slowest run ranges (not intervals)n is shown per side on every row
5 rows from 1 study. No row separates them: 1 tie, 4 unclear.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick GPT-6.1 Sol (OpenAI API) at low effort
No row in this data puts GPT-6.1 Sol (OpenAI API) at low effort ahead of GPT-6.1 Sol (OpenAI API) at high effort. Pick on other grounds (price, access, the tasks you run), or measure your own workload.
When to pick GPT-6.1 Sol (OpenAI API) at high effort
No row in this data puts GPT-6.1 Sol (OpenAI API) at high effort ahead of GPT-6.1 Sol (OpenAI API) at low effort. Pick on other grounds (price, access, the tasks you run), or measure your own workload.
Side by side
The study charts, showing only GPT-6.1 Sol (OpenAI API) at each effort it ran; the two compared settings are highlighted. Open a study for every configuration.
- Total time
- First useful output
| Item | Total time | First useful output | Range (lowest–highest run) | n |
|---|---|---|---|---|
| OpenAI API · GPT-6.1 Sol · low | 1 s | 0.9 s | Total time: 1 s–1.9 s; First useful output: 0.8 s–1.7 s | 5 |
| OpenAI API · GPT-6.1 Sol · high | 1.5 s | 1.3 s | Total time: 1.4 s–2.2 s; First useful output: 1.3 s–2.1 s | 5 |
2 rows, 2 series: Total time, First useful output. Total time: slowest OpenAI API · GPT-6.1 Sol · high 1.5 s (range 1.4 s–2.2 s, n 5). Fastest OpenAI API · GPT-6.1 Sol · low 1 s (range 1 s–1.9 s, n 5). All run ranges overlap. First useful output: slowest OpenAI API · GPT-6.1 Sol · high 1.3 s (range 1.3 s–2.1 s, n 5). Fastest OpenAI API · GPT-6.1 Sol · low 0.9 s (range 0.8 s–1.7 s, n 5). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 5 per row
Matched cohort, fixed exact reply, 5 runs per configuration
Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.
Source: Provider explorer receipts: CLI vs API
- Total time
- First useful output
| Item | Total time | First useful output | Range (lowest–highest run) | n |
|---|---|---|---|---|
| OpenAI API · GPT-6.1 Sol · low | 6 s | 1.1 s | Total time: 5.4 s–6.2 s; First useful output: 1 s–1.4 s | 3 |
| OpenAI API · GPT-6.1 Sol · high | 9.6 s | 5.3 s | Total time: 9.4 s–10.9 s; First useful output: 5 s–6.4 s | 3 |
2 rows, 2 series: Total time, First useful output. Total time: slowest OpenAI API · GPT-6.1 Sol · high 9.6 s (range 9.4 s–10.9 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · low 6 s (range 5.4 s–6.2 s, n 3). Not all run ranges overlap. First useful output: slowest OpenAI API · GPT-6.1 Sol · high 5.3 s (range 5 s–6.4 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · low 1.1 s (range 1 s–1.4 s, n 3). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Matched cohort, small coding task, 3 runs per configuration
Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.
Source: Provider explorer receipts: CLI vs API
| Item | Input tokens | n |
|---|---|---|
| OpenAI API · GPT-6.1 Sol · low | 17 | 5 |
| OpenAI API · GPT-6.1 Sol · high | 17 | 5 |
2 rows. All at 17.
Notesn = 5 per row
Reported input tokens, matched cohort
The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.
Source: Provider explorer receipts: CLI vs API
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Is GPT-6.1 Sol (OpenAI API) better at low effort or high effort?
- GPT-6.1 Sol (OpenAI API) at low effort and GPT-6.1 Sol (OpenAI API) at high effort share 5 measured metrics from one study. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 4 unclear; each row says why. Some rows rest on small samples (n = 3 at the smallest).
- How were the two effort settings measured?
- They share 5 measured metrics from 1 public study: Claude Code CLI vs Codex CLI vs the API: latency and tokens. Only the effort setting differs between the two sides of a row.
- How does cLI vs API: time for a one-line answer (Total time) change from low effort to high effort?
- GPT-6.1 Sol (OpenAI API) at low effort: 1.02 s (n = 5, run range 1 s to 1.9 s). GPT-6.1 Sol (OpenAI API) at high effort: 1.52 s (n = 5, run range 1.4 s to 2.2 s). The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.96 s to 1.87 s; GPT-6.1 Sol (OpenAI API) at high effort 1.35 s to 2.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How does cLI vs API: time for a one-line answer (First useful output) change from low effort to high effort?
- GPT-6.1 Sol (OpenAI API) at low effort: 0.87 s (n = 5, run range 0.8 s to 1.7 s). GPT-6.1 Sol (OpenAI API) at high effort: 1.34 s (n = 5, run range 1.3 s to 2.1 s). The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.84 s to 1.74 s; GPT-6.1 Sol (OpenAI API) at high effort 1.26 s to 2.12 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How does cLI vs API: time for a small coding task (Total time) change from low effort to high effort?
- GPT-6.1 Sol (OpenAI API) at low effort: 6.00 s (n = 3, run range 5.4 s to 6.2 s). GPT-6.1 Sol (OpenAI API) at high effort: 9.56 s (n = 3, run range 9.4 s to 10.9 s). Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) at low effort 5.44 s to 6.20 s; GPT-6.1 Sol (OpenAI API) at high effort 9.44 s to 10.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.
- How does cLI vs API: time for a small coding task (First useful output) change from low effort to high effort?
- GPT-6.1 Sol (OpenAI API) at low effort: 1.05 s (n = 3, run range 1 s to 1.4 s). GPT-6.1 Sol (OpenAI API) at high effort: 5.31 s (n = 3, run range 5 s to 6.4 s). Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.97 s to 1.40 s; GPT-6.1 Sol (OpenAI API) at high effort 4.99 s to 6.42 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.
The studies behind this page
Claude Code CLI vs Codex CLI vs the API: latency and tokens
194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.