How long does an AI coding agent take per task? Minutes, calls and where the time goes
Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.
TL;DR
- A whole task takes minutes. One call takes seconds. Agent's full pipeline took a median 9.6 minutes per SWE-bench Verified instance (range 1.6 to 54.1, n = 33). It made 49.5 model calls per attempt on average (calculation; range 13 to 71, n = 33). One hard single-turn call to Claude Sonnet 5.5 took a median 7.75 s (range 2.26 to 34.79, n = 24).
- More calls came with a higher median time. Attempts with 48 calls or fewer took a median 7.2 minutes (range 1.6 to 38.2, n = 17). Attempts with more took 13.27 minutes (range 9.1 to 54.1, n = 16). Both are calculations. The ranges overlap, and the split does not prove cause.
- The tail is long. 7 of 33 attempts took more than 20 minutes.
- Our reading: plan for minutes per task, set a limit for the tail, and measure calls and tool time before changing either.
Rungs use different tasks and builds, so do not subtract one from another. All ranges below are observed run ranges, not confidence intervals.
| Rung | Unit of work | Time | n |
|---|---|---|---|
| 1 | One hard single-turn call, Sonnet 5.5 | Median 7.75 s (range 2.26 to 34.79 s) | 24 calls |
| 2 | One scheduler repair, Sonnet 5.5 | Median 15.0 s (range 13.89 to 15.89 s) | 3 runs |
| 3 | One small task with tools, Sonnet 5.5 | Median per session 18.0 to 27.3 s, by memory setup (pooled range 10.3 to 44.8 s) | 15 per setup |
| 4 | One real upstream issue, Agent pipeline | 7.2 to 22.7 min across attempts | 12 attempts |
| 5 | One SWE-bench Verified instance, Agent pipeline | Median 9.6 min (range 1.6 to 54.1 min) | 33 attempts |
| 6 | One recorded bench task, Agent pipeline | Median 10.3 min (range 1.3 to 40.8 min) | 48 runs |
1. One hard call: 7.75 seconds
We gave Claude Sonnet 5.5, through Claude Code, 8 hard tasks that a validator grades. Each call had tools off, an empty folder and one turn. The median call took 7.75 s (range 2.26 to 34.79 s, n = 24). All 24 passed (95% interval 86% to 100%). This Sonnet cell hit a ceiling, so it cannot show which tasks would break it. Study: hard model head-to-head.
2. One repair prompt: 15 seconds
Next task: repair a job scheduler that 296 behavioral checks grade. Claude Code with Sonnet 5.5 took a median 15.0 s (13.89 to 15.89 s, n = 3). All 9 evaluated runs in the study passed all 296 checks (95% interval 70% to 100%, n = 9). This task hit a ceiling and cannot rank quality. Excluded and diagnostic receipts do not count.
GPT-6.1 Sol took a median 17.3 s through the OpenAI API (16.28 to 18.61 s, n = 3). It took 61.2 s through the Codex CLI (59.9 to 69.51 s, n = 3).
That is one model on two routes. The ranges do not overlap, and the CLI took 3.5 times as long (calculation: ratio of medians). Three runs only show a direction. Study: CLI vs API latency.
3. One small task with tools: 18 to 27 seconds
With tools, the agent reads, edits and runs tests. In the memory study, Claude Code ran Sonnet 5.5 on 5 small tasks in one Node.js repository. Each session did one task under one of 8 memory setups. Median wall time per session ran from 18.0 s to 27.3 s (n = 15 each). Haiku 4.5 ran from 49.9 s to 68.5 s (n = 10 each).
Across all setups, run times ranged from 10.3 to 44.8 s for Sonnet and 22.5 to 95.4 s for Haiku. These are pooled ranges, not ranges for each setup.
Up to four sessions ran at once on one machine. Some sessions without memory stopped to ask or skipped the changelog. A shorter session need not mean a completed task. Study: agent memory.
4. One real upstream issue: 7 to 23 minutes
We gave the Agent pipeline three real issues (fastify/session, h3js/h3 and Kludex/uvicorn) on four platform builds. Each attempt took 7.2 to 22.7 minutes (12 attempts, one per task and build; median 13.0, a calculation). The times include onboarding.
The fastest attempt, 7.2 minutes, ended with an empty patch and a hand-off to a person. Caps stopped 3 of 12 attempts, and the four builds differ. On the latest build, 1 of 3 tasks reached verified delivery (95% Wilson interval 6% to 79%). One attempt per task finds defects; it does not estimate a stable delivery rate. Study: coding calibration.
5. A full SWE-bench Verified attempt: 9.6 minutes
Agent runs a full pipeline on Sonnet 5.5: onboard, research, plan, act, verify and review. It attempted 33 SWE-bench Verified instances, one attempt each. It resolved 25 (76%, 95% interval 59% to 87%).
The median attempt took 9.6 minutes (1.6 to 54.1). The mean was 14.5 minutes (calculation, n = 33; same range). Each attempt made 49.5 model calls on average (median 48, range 13 to 71; calculation).
The 8 compiled-extension attempts (astropy, matplotlib, scikit-learn) took a median 28.0 minutes (8.6 to 54.1). The other 25 took a median 8.0 (range 1.6 to 38.2). The ranges overlap. They used different builds, so build and repository effects mix. Study: SWE-bench Verified.
The 11 public bash agents made a mean of 20.8 to 88.2 calls per instance (calculation over the same 33 instances per agent). Each panel model call produces a bash step; Agent calls also cover planning and review. Counts differ in kind.
6. Recorded bench tasks: 10.3 minutes
The median across 48 non-empty recorded Agent bench runs was 10.3 minutes (range 1.3 to 40.8) and 49.5 model calls (range 13 to 73). Two empty runs were excluded. These recorded runs are not 48 independent new tasks or an extra validation set. Source: the routing-overhead extract.
Where the time goes
Across the whole attempt. Divide total wall time by total model calls. SWE-bench gives 17.6 s per call (478.45 minutes over 1,632 calls, n = 33 attempts). The calibration set gives 15.8 s (170.4 minutes over 649 calls, n = 12 attempts). Both are calculations of wall time per call, not measured call latency.
These figures include model calls, tool runs, tests, builds and waiting. Our data does not split the time. The hard single-turn tasks differ, so their times cannot reveal Agent’s overhead.
Inside a call. A Sonnet 5.5 routing call (Claude Code, low effort) took a median 2,597 ms (range 1,993 to 5,583 ms, n = 82). The CLI-reported API-time median was 1,596 ms (range 1,060 to 4,757). CLI and harness time had a median of 973 ms (range 826 to 3,673). The routing summary uses the lower middle value for even samples. Medians of parts do not add up. On a one-word answer, Claude Code with Haiku 4.5 spent a median 1,690 ms outside the model (range 1,533 to 1,811 ms, n = 5).
Routing. Routing 49.5 calls in series through Sonnet 5.5 at low effort projects 128.6 s per task. Calculation: 49.5 calls × its median 2.597 s. That is about a fifth of a 10.3-minute task (calculation). It is a serial-delay scenario, not a measured delay or a maximum. The recorded bench runs had model routing off.
The rules hot path took a median 1.42 µs (p95 2.33 µs, n = 20,000 timed decisions). This median-to-p95 span is not a confidence interval. Calculation: 49.5 decisions at the median add about 70 µs. The rules timing excludes database reads and writes. The Sonnet timing includes CLI time, so this compares deployed routes, not equal model conditions. Study: routing overhead.
Stages. A full pipeline onboards, plans, verifies and reviews. Those stages do more work than a bare bash agent. The panel’s times are not in our data, so we cannot measure a time gap or attribute it to those stages. See harness vs model and the Agent harness page.
How to make it faster (our reading)
We did not test these changes end to end. Each rests on a number above.
- Inspect high-call attempts. The split above has medians of 7.2 and 13.27 minutes, with overlapping ranges. It does not show that cutting calls saves time or keeps quality.
- Try lower effort where the task allows. Sonnet 5.5 took a median 5.82 s at low effort (range 2.78 to 19.96 s). At high effort it took 8.81 s (range 2.93 to 35.81 s, n = 16 each). Both passed 16 of 16 (95% interval 81% to 100%). Both cells hit a ceiling on quality. The timing ranges overlap, so this does not establish a faster setting. Source: effort ladder.
- Use rules when they can answer the decision. The serial calculation above projects 128.6 s for Sonnet versus about 70 µs for the rules hot path. It does not test routing quality on your tasks.
- Start fewer CLI processes. The routing calls had a median 973 ms of CLI and harness time (range 826 to 3,673 ms, n = 82). A one-word answer spent a median 1,690 ms outside the model (range 1,533 to 1,811 ms, Haiku 4.5, n = 5). We did not test process reuse.
- Set a limit for the tail. 5 of 33 attempts took more than 30 minutes.
How we measured
- Rungs 1 to 3: validated tasks, one host. Rung 2 ran on 2026-10-03.
- Rungs 4 to 6: Agent’s recorded wall time. Rung 4 repeats three tasks across four builds. Rung 5 keeps one attempt per instance. Rung 6 pools recorded runs and excludes two empty runs.
- Calculations use the published raw extracts and tables. SWE-bench totals use the extracts’ two-decimal minute values. Rung 4 excludes the separate first calibration run on 2026-10-02. A range is not an interval.
Caveats
- Different tasks on every rung. The ladder is not one experiment.
- Small samples. Rung 2 has 3 runs per cell. Rung 4 has one attempt per cell.
- Agent is our product. We have no panel times to compare.
- Contamination. SWE-bench issues are public and likely in training data.
What to read next
- Agent vs 11 public models on SWE-bench Verified
- What a resolved SWE-bench task costs
- Claude Code vs Codex CLI vs the API: latency
Time your own agent
Agent records the time, the calls and the result of every task. Try Agent to see where your minutes go.