• Latency
  • Coding Agent
  • SWE-bench
  • Claude Code

How long does an AI coding agent take per task? Minutes, calls and where the time goes

Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.

TL;DR

  • A whole task takes minutes. One call takes seconds. Agent's full pipeline took a median 9.6 minutes per SWE-bench Verified instance (range 1.6 to 54.1, n = 33). It made 49.5 model calls per attempt on average (calculation; range 13 to 71, n = 33). One hard single-turn call to Claude Sonnet 5.5 took a median 7.75 s (range 2.26 to 34.79, n = 24).
  • More calls came with a higher median time. Attempts with 48 calls or fewer took a median 7.2 minutes (range 1.6 to 38.2, n = 17). Attempts with more took 13.27 minutes (range 9.1 to 54.1, n = 16). Both are calculations. The ranges overlap, and the split does not prove cause.
  • The tail is long. 7 of 33 attempts took more than 20 minutes.
  • Our reading: plan for minutes per task, set a limit for the tail, and measure calls and tool time before changing either.

Rungs use different tasks and builds, so do not subtract one from another. All ranges below are observed run ranges, not confidence intervals.

RungUnit of workTimen
1One hard single-turn call, Sonnet 5.5Median 7.75 s (range 2.26 to 34.79 s)24 calls
2One scheduler repair, Sonnet 5.5Median 15.0 s (range 13.89 to 15.89 s)3 runs
3One small task with tools, Sonnet 5.5Median per session 18.0 to 27.3 s, by memory setup (pooled range 10.3 to 44.8 s)15 per setup
4One real upstream issue, Agent pipeline7.2 to 22.7 min across attempts12 attempts
5One SWE-bench Verified instance, Agent pipelineMedian 9.6 min (range 1.6 to 54.1 min)33 attempts
6One recorded bench task, Agent pipelineMedian 10.3 min (range 1.3 to 40.8 min)48 runs

1. One hard call: 7.75 seconds

We gave Claude Sonnet 5.5, through Claude Code, 8 hard tasks that a validator grades. Each call had tools off, an empty folder and one turn. The median call took 7.75 s (range 2.26 to 34.79 s, n = 24). All 24 passed (95% interval 86% to 100%). This Sonnet cell hit a ceiling, so it cannot show which tasks would break it. Study: hard model head-to-head.

2. One repair prompt: 15 seconds

Next task: repair a job scheduler that 296 behavioral checks grade. Claude Code with Sonnet 5.5 took a median 15.0 s (13.89 to 15.89 s, n = 3). All 9 evaluated runs in the study passed all 296 checks (95% interval 70% to 100%, n = 9). This task hit a ceiling and cannot rank quality. Excluded and diagnostic receipts do not count.

GPT-6.1 Sol took a median 17.3 s through the OpenAI API (16.28 to 18.61 s, n = 3). It took 61.2 s through the Codex CLI (59.9 to 69.51 s, n = 3).

That is one model on two routes. The ranges do not overlap, and the CLI took 3.5 times as long (calculation: ratio of medians). Three runs only show a direction. Study: CLI vs API latency.

3. One small task with tools: 18 to 27 seconds

With tools, the agent reads, edits and runs tests. In the memory study, Claude Code ran Sonnet 5.5 on 5 small tasks in one Node.js repository. Each session did one task under one of 8 memory setups. Median wall time per session ran from 18.0 s to 27.3 s (n = 15 each). Haiku 4.5 ran from 49.9 s to 68.5 s (n = 10 each).

Across all setups, run times ranged from 10.3 to 44.8 s for Sonnet and 22.5 to 95.4 s for Haiku. These are pooled ranges, not ranges for each setup.

Up to four sessions ran at once on one machine. Some sessions without memory stopped to ask or skipped the changelog. A shorter session need not mean a completed task. Study: agent memory.

4. One real upstream issue: 7 to 23 minutes

We gave the Agent pipeline three real issues (fastify/session, h3js/h3 and Kludex/uvicorn) on four platform builds. Each attempt took 7.2 to 22.7 minutes (12 attempts, one per task and build; median 13.0, a calculation). The times include onboarding.

The fastest attempt, 7.2 minutes, ended with an empty patch and a hand-off to a person. Caps stopped 3 of 12 attempts, and the four builds differ. On the latest build, 1 of 3 tasks reached verified delivery (95% Wilson interval 6% to 79%). One attempt per task finds defects; it does not estimate a stable delivery rate. Study: coding calibration.

Baseline (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Fix wave 1 (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build f0ac3a8a

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build 236c0d3f

fastify/session
h3js/h3
Kludex/uvicorn

One panel per series, all on the same axis.

3 rows, 4 series: Baseline (capped), Fix wave 1 (capped), Uncapped, build f0ac3a8a, Uncapped, build 236c0d3f. Baseline (capped): slowest Kludex/uvicorn 17.9 min. Fastest fastify/session 9 min. Fix wave 1 (capped): slowest Kludex/uvicorn 20.3 min. Fastest h3js/h3 13.5 min.

Notes

Onboarding included. uvicorn needed more than the old 20 minute cap once the caps were removed.

Source: Coding calibration: fastify/session, h3, uvicorn

5. A full SWE-bench Verified attempt: 9.6 minutes

Agent runs a full pipeline on Sonnet 5.5: onboard, research, plan, act, verify and review. It attempted 33 SWE-bench Verified instances, one attempt each. It resolved 25 (76%, 95% interval 59% to 87%).

The median attempt took 9.6 minutes (1.6 to 54.1). The mean was 14.5 minutes (calculation, n = 33; same range). Each attempt made 49.5 model calls on average (median 48, range 13 to 71; calculation).

The 8 compiled-extension attempts (astropy, matplotlib, scikit-learn) took a median 28.0 minutes (8.6 to 54.1). The other 25 took a median 8.0 (range 1.6 to 38.2). The ranges overlap. They used different builds, so build and repository effects mix. Study: SWE-bench Verified.

Model calls per instance

Mean over the same 33 instances

DeepSeek V3.2 (high)
GLM 5 (high)
Claude 4.5 Haiku (high)
MiniMax M2.5 (high)
Kimi K2.5 (high)
Gemini 3 Flash (high)
Claude 4.5 Sonnet (high)
Agent
Claude 4.5 Opus (high)
GPT 5.2 (high)
Claude 4.6 Opus
GPT 5 mini

Hover or focus a bar for its ratio to Agent (the highlighted row): a ratio of the two values shown, not a measurement.

12 rows. Highest DeepSeek V3.2 (high) 88.2 (n 33). Lowest GPT 5 mini 20.8 (n 33).

Notesn = 33 per row

A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review).

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

The 11 public bash agents made a mean of 20.8 to 88.2 calls per instance (calculation over the same 33 instances per agent). Each panel model call produces a bash step; Agent calls also cover planning and review. Counts differ in kind.

6. Recorded bench tasks: 10.3 minutes

The median across 48 non-empty recorded Agent bench runs was 10.3 minutes (range 1.3 to 40.8) and 49.5 model calls (range 13 to 73). Two empty runs were excluded. These recorded runs are not 48 independent new tasks or an extra validation set. Source: the routing-overhead extract.

Where the time goes

Across the whole attempt. Divide total wall time by total model calls. SWE-bench gives 17.6 s per call (478.45 minutes over 1,632 calls, n = 33 attempts). The calibration set gives 15.8 s (170.4 minutes over 649 calls, n = 12 attempts). Both are calculations of wall time per call, not measured call latency.

These figures include model calls, tool runs, tests, builds and waiting. Our data does not split the time. The hard single-turn tasks differ, so their times cannot reveal Agent’s overhead.

Inside a call. A Sonnet 5.5 routing call (Claude Code, low effort) took a median 2,597 ms (range 1,993 to 5,583 ms, n = 82). The CLI-reported API-time median was 1,596 ms (range 1,060 to 4,757). CLI and harness time had a median of 973 ms (range 826 to 3,673). The routing summary uses the lower middle value for even samples. Medians of parts do not add up. On a one-word answer, Claude Code with Haiku 4.5 spent a median 1,690 ms outside the model (range 1,533 to 1,811 ms, n = 5).

Routing. Routing 49.5 calls in series through Sonnet 5.5 at low effort projects 128.6 s per task. Calculation: 49.5 calls × its median 2.597 s. That is about a fifth of a 10.3-minute task (calculation). It is a serial-delay scenario, not a measured delay or a maximum. The recorded bench runs had model routing off.

The rules hot path took a median 1.42 µs (p95 2.33 µs, n = 20,000 timed decisions). This median-to-p95 span is not a confidence interval. Calculation: 49.5 decisions at the median add about 70 µs. The rules timing excludes database reads and writes. The Sonnet timing includes CLI time, so this compares deployed routes, not equal model conditions. Study: routing overhead.

Stages. A full pipeline onboards, plans, verifies and reviews. Those stages do more work than a bare bash agent. The panel’s times are not in our data, so we cannot measure a time gap or attribute it to those stages. See harness vs model and the Agent harness page.

How to make it faster (our reading)

We did not test these changes end to end. Each rests on a number above.

  1. Inspect high-call attempts. The split above has medians of 7.2 and 13.27 minutes, with overlapping ranges. It does not show that cutting calls saves time or keeps quality.
  2. Try lower effort where the task allows. Sonnet 5.5 took a median 5.82 s at low effort (range 2.78 to 19.96 s). At high effort it took 8.81 s (range 2.93 to 35.81 s, n = 16 each). Both passed 16 of 16 (95% interval 81% to 100%). Both cells hit a ceiling on quality. The timing ranges overlap, so this does not establish a faster setting. Source: effort ladder.
  3. Use rules when they can answer the decision. The serial calculation above projects 128.6 s for Sonnet versus about 70 µs for the rules hot path. It does not test routing quality on your tasks.
  4. Start fewer CLI processes. The routing calls had a median 973 ms of CLI and harness time (range 826 to 3,673 ms, n = 82). A one-word answer spent a median 1,690 ms outside the model (range 1,533 to 1,811 ms, Haiku 4.5, n = 5). We did not test process reuse.
  5. Set a limit for the tail. 5 of 33 attempts took more than 30 minutes.

How we measured

  • Rungs 1 to 3: validated tasks, one host. Rung 2 ran on 2026-10-03.
  • Rungs 4 to 6: Agent’s recorded wall time. Rung 4 repeats three tasks across four builds. Rung 5 keeps one attempt per instance. Rung 6 pools recorded runs and excludes two empty runs.
  • Calculations use the published raw extracts and tables. SWE-bench totals use the extracts’ two-decimal minute values. Rung 4 excludes the separate first calibration run on 2026-10-02. A range is not an interval.

Caveats

  • Different tasks on every rung. The ladder is not one experiment.
  • Small samples. Rung 2 has 3 runs per cell. Rung 4 has one attempt per cell.
  • Agent is our product. We have no panel times to compare.
  • Contamination. SWE-bench issues are public and likely in training data.

Time your own agent

Agent records the time, the calls and the result of every task. Try Agent to see where your minutes go.

The data behind this post

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

  • Calibration
  • Coding Agents

Coding calibration: what broke on three real pull requests

One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.

1 of 3Verified deliveries, latest build · n = 3

4 chartsUpdated October 5, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.