3 measured metrics · 1 study

Claude Sonnet 5.5vsGPT-6.1 Sol (OpenAI API)

No row separates them: 3 unclear.

The verdict

Claude Sonnet 5.5 and GPT-6.1 Sol (OpenAI API) share 3 measured metrics from one study. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example Claude Code vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: fastest–slowest run ranges (not intervals)n is shown per side on every row

3 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Output tokens to repair the scheduler (Output tokens), 1.7x (Claude Sonnet 5.5 larger).

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Sonnet 5.5
  • GPT-6.1 Sol (OpenAI API)
  • fastest–slowest run (not an interval)
  • where the two overlap

Claude Code CLI vs Codex CLI vs the API: latency and tokens

3 unclear
  • Repairing a scheduler: Claude Code vs Codex vs API (Total time): Claude Sonnet 5.5 15.0 s (n 3, run range 13.9 s–15.9 s); GPT-6.1 Sol (OpenAI API) 17.3 s (n 3, run range 16.3 s–18.6 s). Unclear.
  • Repairing a scheduler: Claude Code vs Codex vs API (First useful output): Claude Sonnet 5.5 7.55 s (n 3, run range 6.8 s–7.6 s); GPT-6.1 Sol (OpenAI API) 7.46 s (n 3, run range 6.7 s–9.1 s). Unclear.
  • Output tokens to repair the scheduler (Output tokens): Claude Sonnet 5.5 2,227 (n 3); GPT-6.1 Sol (OpenAI API) 1,313 (n 3). Unclear.

Marks: fastest–slowest run ranges (not intervals)n is shown per side on every row

3 rows from 1 study. No row separates them: 3 unclear.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Sonnet 5.5

No row in this data puts Claude Sonnet 5.5 ahead of GPT-6.1 Sol (OpenAI API). Pick on other grounds (price, access, the tasks you run), or measure your own workload.

When to pick GPT-6.1 Sol (OpenAI API)

No row in this data puts GPT-6.1 Sol (OpenAI API) ahead of Claude Sonnet 5.5. Pick on other grounds (price, access, the tasks you run), or measure your own workload.

Side by side

The study charts, showing only these two. Open a study for every configuration.

  • Total time
  • First useful output
Entrance: medians race at 12× real timeMotion reduced: press Replay to animateThe slowest median is 17.3 s. The clock runs at the recorded speed.
Claude Code CLI · Sonnet 5.5 · medium
OpenAI API · GPT-6.1 Sol · medium

2 rows, 2 series: Total time, First useful output. Total time: slowest OpenAI API · GPT-6.1 Sol · medium 17.3 s (range 16.3 s–18.6 s, n 3). Fastest Claude Code CLI · Sonnet 5.5 · medium 15 s (range 13.9 s–15.9 s, n 3). Not all run ranges overlap. First useful output: slowest Claude Code CLI · Sonnet 5.5 · medium 7.6 s (range 6.8 s–7.6 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · medium 7.5 s (range 6.7 s–9.1 s, n 3). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Same prompt, medium effort, 296 behavioral checks, 3 runs each

All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.

Source: Provider explorer receipts: CLI vs API

  • Output tokens
  • of which reasoning tokens (reported) (inner bar)
Claude Code CLI · Sonnet 5.5 · medium
OpenAI API · GPT-6.1 Sol · medium

2 rows, 2 series: Output tokens, Reasoning tokens (reported). Output tokens: highest Claude Code CLI · Sonnet 5.5 · medium 2,227 (n 3). Lowest OpenAI API · GPT-6.1 Sol · medium 1,313 (n 3). Reasoning tokens (reported): highest OpenAI API · GPT-6.1 Sol · medium 267 (n 3). Lowest Claude Code CLI · Sonnet 5.5 · medium 0 (n 0).

Notesn 0–3 per row

Median per run; reasoning tokens shown separately where reported

The Claude CLI does not report reasoning tokens separately; 0 there means "not reported", not "none".

Source: Provider explorer receipts: CLI vs API

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Sonnet 5.5 or GPT-6.1 Sol (OpenAI API)?
Claude Sonnet 5.5 and GPT-6.1 Sol (OpenAI API) share 3 measured metrics from one study. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example Claude Code vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
How were Claude Sonnet 5.5 and GPT-6.1 Sol (OpenAI API) measured?
They share 3 measured metrics from 1 public study: Claude Code CLI vs Codex CLI vs the API: latency and tokens. Every row names its configuration, its sample size and its interval or range.
How do Claude Sonnet 5.5 and GPT-6.1 Sol (OpenAI API) compare on repairing a scheduler: Claude Code vs Codex vs API (Total time)?
Claude Sonnet 5.5: 15.0 s (Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs; n = 3; run range 13.9 s to 15.9 s). GPT-6.1 Sol (OpenAI API): 17.3 s (OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs; n = 3; run range 16.3 s to 18.6 s). Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 13.9 s to 15.9 s; GPT-6.1 Sol (OpenAI API) 16.3 s to 18.6 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.
How do Claude Sonnet 5.5 and GPT-6.1 Sol (OpenAI API) compare on repairing a scheduler: Claude Code vs Codex vs API (First useful output)?
Claude Sonnet 5.5: 7.55 s (Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs; n = 3; run range 6.8 s to 7.6 s). GPT-6.1 Sol (OpenAI API): 7.46 s (OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs; n = 3; run range 6.7 s to 9.1 s). The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 6.77 s to 7.63 s; GPT-6.1 Sol (OpenAI API) 6.68 s to 9.05 s); the medians alone do not show a reliable difference. A range is not a confidence interval.

The studies behind this page

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.