Calculation
Claude Sonnet 5.5 · Claude Code
0.90
hours waited per month
2.58 min a day, one person
Per task: 7.8 s (recorded median, n = 24)
Fastest to slowest run: 0.26 to 4.06 hours a month. This is a range of runs, not an interval.
Calculation on recorded times · updated October 6, 2026
Pick a task set. Pick two to four configurations. The page multiplies the recorded time of one task by the number of tasks a person runs each day. This is arithmetic on measured times, not a new run. It does not say which configuration does the work well.
Calculation
0.90
hours waited per month
2.58 min a day, one person
Per task: 7.8 s (recorded median, n = 24)
Fastest to slowest run: 0.26 to 4.06 hours a month. This is a range of runs, not an interval.
Calculation
1.53
hours waited per month
4.37 min a day, one person
Per task: 13.1 s (recorded median, n = 16)
Fastest to slowest run: 1.00 to 7.19 hours a month. This is a range of runs, not an interval.
Calculation. Hours waited per month. Claude Sonnet 5.5 · Claude Code: 0.90 hours a month. GPT-6.1 Sol (medium) · Codex CLI: 1.53 hours a month.
The medians of Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI differ by a factor of 1.7 (calculation: larger median ÷ smaller median).
Their ranges overlap: in your runs the gap can be smaller or reversed.
These lanes use the recorded times. They are not scaled to your day. Press Replay in real time to wait as long as one task takes.
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call on hard tasks (separate batches) | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 7.8 s | 2.3 s–34.8 s | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 8.5 s–61.6 s | 16 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 13.1 s (range 8.5 s–61.6 s, n 16). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.
Median for each configuration you chose. Lines show the fastest and the slowest recorded run.
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
The two configurations at three points of their recorded runs.
Gap labels, GPT-6.1 Sol (medium) · Codex CLI vs Claude Sonnet 5.5 · Claude Code: GPT-6.1 Sol (medium) · Codex CLI is x% higher (+) or lower (−) than Claude Sonnet 5.5 · Claude Code, calculated from the two values shown (the change counted from Claude Sonnet 5.5 · Claude Code’s value).
| Item | Claude Sonnet 5.5 · Claude Code | GPT-6.1 Sol (medium) · Codex CLI | n |
|---|---|---|---|
| Fastest run | 2.3 s | 8.5 s | 24 |
| Median | 7.8 s | 13.1 s | 24 |
| Slowest run | 34.8 s | 61.6 s | 24 |
3 rows, 2 series: Claude Sonnet 5.5 · Claude Code, GPT-6.1 Sol (medium) · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest Slowest run 34.8 s (n 24). Fastest Fastest run 2.3 s (n 24). GPT-6.1 Sol (medium) · Codex CLI: slowest Slowest run 61.6 s (n 16). Fastest Fastest run 8.5 s (n 16).
Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI, for one call. The two ends are single recorded runs, not an interval.
These are recorded values of the two configurations you chose. They do not include retries or the routing add-on.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
The same results as a table. Each row is a calculation on the recorded values shown in it.
| Configuration | Recorded median | n | Routing add-on (upper bound) | Attempts | Calculated time per task | Person wait per day | Team hours per month | Band (hours per month) | Money per month (your currency) |
|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 7.8 s | 24 | 0 s | 1 | 7.8 s | 155 s | 0.90 | 0.26 to 4.06 (run-range) | Not entered |
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 16 | 0 s | 1 | 13.1 s | 262.2 s | 1.53 | 1.00 to 7.19 (run-range) | Not entered |
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
It depends on the task and the configuration. In the task set “Hard coding tasks: Claude Code and Codex CLI” (Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks), one person who runs 20 tasks a day waits 2.58 min a day with Claude Sonnet 5.5 · Claude Code (median 7.75 s, n 24) and 4.37 min a day with GPT-6.1 Sol (medium) · Codex CLI (median 13.11 s, n 16). At 21 working days a month, that is 0.90 and 1.53 hours (calculation). The medians of Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI differ by a factor of 1.7 (calculation: larger median ÷ smaller median). Their ranges overlap: in your runs the gap can be smaller or reversed.
In the task set “Hard coding tasks: Claude Code and Codex CLI”, the median time for one call runs from 7.75 s (Claude Sonnet 5.5 · Claude Code) to 39.01 s (Claude Haiku 4.5 · Claude Code). The fastest-to-slowest run ranges of all configurations overlap, so these medians describe this run and do not rank the configurations.
Hours waited a month = tasks per person per day × time per task × people × working days ÷ 3,600. The time per task is the recorded median, plus an optional routing delay, times 1 ÷ p attempts when you count retries. Money is the hours × the hourly cost you type. Every number is a calculation on recorded times.
Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.