Calculation on recorded times · updated October 6, 2026

What does waiting for an AI coding agent cost?

Pick a task set. Pick two to four configurations. The page multiplies the recorded time of one task by the number of tasks a person runs each day. This is arithmetic on measured times, not a new run. It does not say which configuration does the work well.

Configurations (choose 2 to 4)

The first two you choose are compared in the sentences under the results. Keep 2 to 4 selected.

On: the run time counts as waiting. Off: nobody waits, so the waiting cost is 0.

On: one correct answer takes 1 ÷ p attempts, where p is the recorded pass rate.

Hours waited per month

Calculation

Claude Sonnet 5.5 · Claude Code

0.90

hours waited per month

2.58 min a day, one person

Per task: 7.8 s (recorded median, n = 24)

Fastest to slowest run: 0.26 to 4.06 hours a month. This is a range of runs, not an interval.

Calculation

GPT-6.1 Sol (medium) · Codex CLI

1.53

hours waited per month

4.37 min a day, one person

Per task: 13.1 s (recorded median, n = 16)

Fastest to slowest run: 1.00 to 7.19 hours a month. This is a range of runs, not an interval.

Calculation. Hours waited per month. Claude Sonnet 5.5 · Claude Code: 0.90 hours a month. GPT-6.1 Sol (medium) · Codex CLI: 1.53 hours a month.

The medians of Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI differ by a factor of 1.7 (calculation: larger median ÷ smaller median).

Their ranges overlap: in your runs the gap can be smaller or reversed.

The wait for one task, as recorded

These lanes use the recorded times. They are not scaled to your day. Press Replay in real time to wait as long as one task takes.

Entrance: medians race at 9.4× real timeMotion reduced: press Replay to animateThe slowest median is 13.1 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 13.1 s (range 8.5 s–61.6 s, n 16). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 16–24 per row

Median for each configuration you chose. Lines show the fastest and the slowest recorded run.

One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Fastest run, median and slowest run

The two configurations at three points of their recorded runs.

  • Claude Sonnet 5.5 · Claude Code
  • GPT-6.1 Sol (medium) · Codex CLI (square)
In chart order.
Fastest run
Median
Slowest run

Gap labels, GPT-6.1 Sol (medium) · Codex CLI vs Claude Sonnet 5.5 · Claude Code: GPT-6.1 Sol (medium) · Codex CLI is x% higher (+) or lower (−) than Claude Sonnet 5.5 · Claude Code, calculated from the two values shown (the change counted from Claude Sonnet 5.5 · Claude Code’s value).

3 rows, 2 series: Claude Sonnet 5.5 · Claude Code, GPT-6.1 Sol (medium) · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest Slowest run 34.8 s (n 24). Fastest Fastest run 2.3 s (n 24). GPT-6.1 Sol (medium) · Codex CLI: slowest Slowest run 61.6 s (n 16). Fastest Fastest run 8.5 s (n 16).

Notesn 16–24 per row

Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI, for one call. The two ends are single recorded runs, not an interval.

These are recorded values of the two configurations you chose. They do not include retries or the routing add-on.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Every number

The same results as a table. Each row is a calculation on the recorded values shown in it.

ConfigurationRecorded mediannRouting add-on (upper bound)AttemptsCalculated time per taskPerson wait per dayTeam hours per monthBand (hours per month)Money per month (your currency)
Claude Sonnet 5.5 · Claude Code7.8 s240 s17.8 s155 s0.900.26 to 4.06 (run-range)Not entered
GPT-6.1 Sol (medium) · Codex CLI13.1 s160 s113.1 s262.2 s1.531.00 to 7.19 (run-range)Not entered

How the page calculates

  1. Time per task = the recorded median time, plus the routing add-on, times the attempts.
  2. Waiting per person per day = tasks per day × time per task.
  3. Hours per month = waiting per person per day × people × working days ÷ 3,600.
  4. Money per month = hours per month × your hourly cost.
  5. Attempts = 1 ÷ p, where p is the recorded pass rate. The range uses the 95% Wilson interval of p.
  6. Routing add-on = decisions per task × the recorded median decision time.

What to know before you use these numbers

  • This is a calculation, not a run. We multiply recorded times. We did not time your tasks.
  • The time per task is the recorded median. A real day mixes quick and slow runs. The band uses the fastest and the slowest recorded run. It is a range of runs, not an interval.
  • Every study ran on one host and one network. Its own caveats are listed under the task set you chose. Read them before you use a time.
  • The money line is hours × your hourly cost. It assumes the person does nothing else while they wait.
  • Retries assume that every attempt takes the median time and that attempts are independent. The range leaves out the spread of run times.
  • The routing add-on is an upper bound. It assumes each decision waits for the one before.
  • The page adds no queue time, rate-limit wait or network delay beyond what the runs recorded.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks: study caveats
  • 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  • Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  • Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  • The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.
  • Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  • CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.
  • Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  • List-price costs are calculations; the calls used a flat subscription.

Where the times come from

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

Questions

How much time does waiting for an AI coding agent cost?

It depends on the task and the configuration. In the task set “Hard coding tasks: Claude Code and Codex CLI” (Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks), one person who runs 20 tasks a day waits 2.58 min a day with Claude Sonnet 5.5 · Claude Code (median 7.75 s, n 24) and 4.37 min a day with GPT-6.1 Sol (medium) · Codex CLI (median 13.11 s, n 16). At 21 working days a month, that is 0.90 and 1.53 hours (calculation). The medians of Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI differ by a factor of 1.7 (calculation: larger median ÷ smaller median). Their ranges overlap: in your runs the gap can be smaller or reversed.

How do AI coding agents compare in speed?

In the task set “Hard coding tasks: Claude Code and Codex CLI”, the median time for one call runs from 7.75 s (Claude Sonnet 5.5 · Claude Code) to 39.01 s (Claude Haiku 4.5 · Claude Code). The fastest-to-slowest run ranges of all configurations overlap, so these medians describe this run and do not rank the configurations.

How is the cost of waiting for AI calculated?

Hours waited a month = tasks per person per day × time per task × people × working days ÷ 3,600. The time per task is the recorded median, plus an optional routing delay, times 1 ÷ p attempts when you count retries. Money is the hours × the hourly cost you type. Every number is a calculation on recorded times.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.