What does an agent loop cost? 1.3 to 8.7 times the median tokens (calculation)
Agent loop cost on 8 hard tasks: 1.3–8.7 times the median tokens and up to 2.1 times the price per strict pass (calculations), with no clear pass-rate gain.
TL;DR
- Question: what does an agent loop cost next to one model call, and what does the extra spend buy? We ran the same 8 hard tasks both ways for three models, with the same strict validators.
- Answer: the loop used more tokens in every pair. It cost more per strict pass for two of three models. It gave no clear gain in passes. Haiku's median time went from 39.0 s to 56.8 s (ranges 15.3–75.1 / 24.5–223.7 s; n = 24 each). The ranges overlap.
- Tokens: median total tokens per attempt, single call to loop. Claude Haiku 4.5 9,038 to 78,432 (8.7 times). Claude Sonnet 5.5 3,380 to 10,483 (3.1 times). GPT-6 Luna in the Codex CLI 11,954 to 16,058 (1.3 times). The ratios are calculations. In every pair the fewest-to-most ranges do not overlap.
- Price per strict pass, at list price (a calculation): Haiku $0.0672 to $0.1422, Sonnet $0.0143 to $0.0275, GPT-6 Luna $0.0012 to $0.0010. The calls ran on flat subscriptions, so none of this is a bill.
- What the spend bought: Haiku passed 11 of 24 as one call and 13 of 24 as a loop (95% intervals 28% to 65% and 35% to 72%). Sonnet passed 24 of 24 and 16 of 16 (95% intervals 86% to 100% and 81% to 100%), a ceiling. GPT-6 Luna passed 10 of 16 and 12 of 14 (39% to 82% and 60% to 96%). Every pair of intervals overlaps, so no loop is ahead.
- Tokens rose even when no tool ran. Sonnet ran no tool in 13 of 16 loop attempts. Even so, every Sonnet loop attempt sent at least 9,398 input tokens. No Sonnet single call sent more than 2,669.
- A bigger model's single call had a lower calculated price in this sample. Per strict pass, Sonnet as one call cost $0.0143. Haiku as a loop cost $0.1422 (9.9 times, a calculation). Each row is a CLI + model pair.
Read the full study, with charts and data. This post is about cost. For the pass-rate question, read Does letting an AI run code help?
What we tested
We used the eight hard tasks of our hard model head-to-head. Each task has a sandboxed validator. A strict pass needs the whole final message to pass the validator as written. A right answer in the wrong wrapping is a format miss. A format miss does not count as a pass.
- Single call. The task prompt, tools off, one turn. We had no paid API key, so the CLI with tools off stands in for an API call.
- Agent loop. The same prompt plus one paragraph that lets the model create and run files in the current folder. The sandbox gave it one empty folder per attempt and no network. Claude Code had 30 turns at most. Each session had 10 minutes.
| Cell | Source | Attempts |
|---|---|---|
| Haiku 4.5, single call | reference cell, hard head-to-head | 24 (8 tasks × 3) |
| Haiku 4.5, agent loop | new | 24 (8 tasks × 3) |
| Sonnet 5.5, single call | reference cell, hard head-to-head | 24 (8 tasks × 3) |
| Sonnet 5.5, agent loop | new | 16 (8 tasks × 2) |
| GPT-6 Luna, single call | new | 16 (8 tasks × 2) |
| GPT-6 Luna, agent loop | new | 16 (8 tasks × 2), 2 left out |
An attempt is one call or one session. We kept every attempt, failures included. Two GPT-6 Luna loop attempts read a file outside the work folder. We left them out of every scored metric, including time, tokens and cost. Both were on the DST task, so its loop covers seven tasks (n = 14) and its single-call cell covers eight (n = 16).
Tokens: how many more does an agent loop use?
Median total tokens per attempt (input with cache reads and writes, plus output). The fewest and most in the cell are in brackets. These ranges are not confidence intervals.
| Model (n: single / loop) | Single call | Agent loop | Ratio (calculation) |
|---|---|---|---|
| Haiku 4.5 (24 / 24) | 9,038 (6,075 to 13,248) | 78,432 (44,273 to 536,960) | 8.7 times |
| Sonnet 5.5 (24 / 16) | 3,380 (2,412 to 6,154) | 10,483 (9,623 to 36,283) | 3.1 times |
| GPT-6 Luna (16 / 14) | 11,954 (11,616 to 12,314) | 16,058 (15,534 to 39,292) | 1.3 times |
The ranges do not overlap in any row. A range is not a confidence interval. GPT-6 Luna used different Codex routes for the single call and the loop. Part of its gap may come from the route and not from the loop.
Where do the extra tokens come from? The receipts show repeated input and long attempts. Prompt and tool-definition changes may also contribute, but we did not isolate their effect.
1. A loop sends its context again on every turn. Haiku's 24 loop attempts took 128 turns (3 to 19 per attempt). A loop turn carried about 20,100 input tokens on average. A single call carried about 4,000 (input tokens divided by turns, a calculation).
Of Haiku's loop input, 89% came from cache reads (2,275,566 of 2,569,125 tokens, a calculation). The cache makes the repeated context cheap. It does not make it free.
2. Input rose even when no tool ran. Sonnet ran no tool in 13 of 16 loop attempts, and those 13 took one turn each. Still, every Sonnet loop attempt sent at least 9,398 input tokens. No Sonnet single call sent more than 2,669.
Most likely, Claude Code sends its tool definitions with every request once tools are on. We did not test that. The two Claude runners also differ in their flags.
3. A few attempts run long. On the DST day-length fix, three Haiku attempts took 9, 12 and 19 turns. Together they used 1,114,624 tokens, which is 40% of the tokens of all 24 attempts (a calculation). The longest used 536,960 tokens, 6.8 times the cell median (calculation), and it failed.
Sonnet's three tool-using attempts used 42% of its loop tokens (98,831 of 232,786, a calculation). A median hides this tail. The cost follows the sum.
Time: is an agent loop slower?
Median total time per attempt, single call to loop. Brackets show fastest-to-slowest ranges, not confidence intervals. Ratios are calculations.
| Model (n: single / loop) | Single call | Agent loop | Ratio (calculation) |
|---|---|---|---|
| Haiku 4.5 (24 / 24) | 39.0 s (15.3 to 75.1) | 56.8 s (24.5 to 223.7) | 1.5 times |
| Sonnet 5.5 (24 / 16) | 7.7 s (2.3 to 34.8) | 7.4 s (2.7 to 24.2) | 0.96 times |
| GPT-6 Luna (16 / 14) | 5.2 s (3.6 to 11.3) | 9.3 s (3.8 to 15.9) | 1.8 times |
In all three pairs the ranges overlap, so we do not call the loop slower.
Input tokens rose far more than time did. Haiku's median input rose 18.2 times (3,941 to 71,691), a calculation. Input ranges were 3,879–4,221 / 41,732–516,306 tokens (n = 24 each). Its median time rose 1.5 times (calculation). The tail is the exception. Haiku's 95th percentile went from 68.5 s to 176.3 s, and its slowest loop attempt took 223.7 s.
Two limits apply to every time in this post. The Mac that ran the sessions also ran other local jobs, so wall times may carry contention. Token counts do not. The Claude single-call cells also ran in an earlier batch.
Price: what does a strict pass cost?
Cost per strict pass divides the list-price cost of every attempt in a cell, failures included, by its strict passes. Single call to loop. Both columns are calculations.
| Model (n: single / loop) | USD per strict pass | Tokens per strict pass (calculation) |
|---|---|---|
| Haiku 4.5 (24 / 24) | $0.0672 to $0.1422 (2.1 times) | 20,163 to 213,557 (10.6 times) |
| Sonnet 5.5 (24 / 16) | $0.0143 to $0.0275 (1.9 times) | 3,391 to 14,549 (4.3 times) |
| GPT-6 Luna (16 / 14) | $0.0012 to $0.0010 (0.85 times) | 19,044 to 22,625 (1.2 times) |
Tokens per strict pass is the total tokens of the cell divided by its strict passes. Haiku's loop used 10.6 times the tokens for each strict pass in this sample (calculation). Its token use rose far more than its pass count.
Where the money goes. Share of list-price cost by token type (calculation: reported tokens times list price). The last column is the share of tokens that are output. Cell sizes are n = 24 / 24 for Haiku, 24 / 16 for Sonnet, and 16 / 14 for Luna. These calculations have no tested uncertainty interval.
| Cell | Output | Cache writes | Cache reads | Other input | Output share of tokens |
|---|---|---|---|---|---|
| Haiku 4.5, single call | 85.4% | 3.4% | 0.0% | 11.2% | 56.9% |
| Haiku 4.5, agent loop | 56.0% | 31.6% | 12.3% | 0.1% | 7.5% |
| Sonnet 5.5, single call | 72.0% | 25.9% | 2.0% | 0.0% | 30.5% |
| Sonnet 5.5, agent loop | 35.0% | 57.9% | 7.0% | 0.0% | 6.6% |
| GPT-6 Luna, single call | 19.2% | 0.0% | 8.8% | 72.0% | 2.4% |
| GPT-6 Luna, agent loop | 29.6% | 0.0% | 16.9% | 53.6% | 2.6% |
Output is 7.5% of Haiku's loop tokens and 56.0% of its cost. Cache writes, the price of storing the growing context, add another 31.6%. In a loop you pay for context and for output. GPT-6 Luna reports no cache writes.
Can any pass rate make the loop pay? For Haiku on this set, no. A Haiku loop attempt cost $0.0770. A single call cost $0.0308. The loop cost 2.5 times as much per attempt (a calculation).
Cost per pass cannot fall below cost per attempt. Even a perfect 24 of 24 would cost $0.0770 per pass at the same spend. The single call cost $0.0672 per pass.
Sonnet: the single call already passed 24 of 24. The loop nearly doubled the calculated cost per pass (1.9 times). It also passed every attempt: 16 of 16.
GPT-6 Luna is the one pair where the price per pass fell, from $0.0012 to $0.0010. Its price per attempt rose 1.2 times (calculation). Its observed pass count went from 10 of 16 (95% interval 39–82%) to 12 of 14 (60–96%). The intervals overlap, and the loop omits DST. The two routes differ, and 12 of 14 loop attempts ran no tool. We do not credit the code.
A bigger model, one call. Sonnet 5.5 as one call passed 24 of 24 at $0.0143 per pass. Haiku 4.5 as a loop passed 13 of 24 at $0.1422 per pass, which is 9.9 times as much (a calculation). Even Haiku's single call cost 4.7 times as much per pass as Sonnet's (calculation) ($0.0672 against $0.0143). Cost per call is not cost per answer. Our post on hard tasks across Claude models has the single-call rows.
What did the extra spend buy?
| Pair | Single call | Agent loop | Verdict |
|---|---|---|---|
| Haiku 4.5 | 11/24 (28% to 65%) | 13/24 (35% to 72%) | Intervals overlap: no clear difference |
| Sonnet 5.5 | 24/24 (86% to 100%) | 16/16 (81% to 100%) | Both at the ceiling: this set cannot separate them |
| GPT-6 Luna | 10/16 (39% to 82%) | 12/14 (60% to 96%) | Intervals overlap: no clear difference |
Our rule: a side is ahead only when the 95% Wilson intervals do not overlap. None did. For Haiku the gain is 2 attempts in 24 (+8 points, a calculation).
The spend did not fall evenly across tasks. This table shows Haiku's loop, task by task. Each task has 3 attempts per arm. These rows show where the money went; they do not rank tasks. All token shares and list-price costs are calculations. The 95% Wilson intervals are 0–56% for 0/3, 6–79% for 1/3, 21–94% for 2/3, and 44–100% for 3/3. Every task pair overlaps.
| Task | Strict passes, single call to loop | Turns per attempt | Loop tokens (share) | Loop list-price cost (share) |
|---|---|---|---|---|
| Interval merge fix | 3/3 to 2/3 | 3 to 5 | 179,915 (6%) | $0.1409 (8%) |
| DST day-length fix | 1/3 to 2/3 | 9 to 19 | 1,114,624 (40%) | $0.5326 (29%) |
| CSV parser | 2/3 to 2/3 | 3 to 4 | 224,815 (8%) | $0.2181 (12%) |
| Event-loop order | 0/3 to 3/3 | 3 to 3 | 165,125 (6%) | $0.1627 (9%) |
| Room schedule | 0/3 to 2/3 | 3 to 6 | 299,719 (11%) | $0.2818 (15%) |
| SemVer regex | 3/3 to 2/3 | 4 to 5 | 212,785 (8%) | $0.1286 (7%) |
| Money refactor | 2/3 to 0/3 | 3 to 6 | 243,736 (9%) | $0.1563 (8%) |
| SQL report | 0/3 to 0/3 | 3 to 7 | 335,524 (12%) | $0.2283 (12%) |
- Where the loop had more strict passes: Event-loop order went from 0 of 3 to 3 of 3 for $0.1627 in total. That is $0.0542 per pass (a calculation), the lowest of the six tasks that had a loop pass. The single call had no pass to divide by.
- Where it cost most: the DST day-length fix took 29% of the cost and gained 1 pass in 3. It was the task that ran longest.
- Where it had no strict pass: the money refactor and the SQL report passed 0 of 3 each in the loop. They cost $0.3846 together, 21% of the cell (calculation). In 4 of those 6 attempts, the answer was right and the wrapping was wrong.
That last point is general. Haiku's 7 format-miss attempts cost $0.4771, which is 26% of the loop's spend (a calculation). They gave no strict pass.
What to do
These steps are our advice. Each one has the number behind it.
- Run the single call first. If it already passes, skip the loop. Sonnet passed 24 of 24 as one call. The loop cost 1.9 times as much per pass (calculation). It also passed all 16 attempts; this set hit a ceiling.
- Turn tools on per task, not per account. Input rose even without tool use: Sonnet ranges were 2,234–2,669 (single, n = 24) / 9,398–33,040 (loop, n = 16). The cause was not isolated. Send tasks that have a check to run, such as a program's output, to the loop. Send the rest to a plain call.
- Compare cost per strict pass, not cost per call. Haiku's loop cost 2.5 times as much per attempt and 2.1 times as much per pass (calculations).
- Cap turns and tokens. Log every attempt that reaches the cap. Three of 24 Haiku attempts used 40% of the tokens. Our 30-turn cap never applied, because the longest attempt used 19 turns.
- End with only the answer. Say in the prompt what the last message must hold. Report any grading after prose removal separately from the strict result. A format miss costs the full price of an attempt.
- Keep the start of the prompt stable. That lets the cache work. Without a cache, the same Haiku loop tokens would cost $3.60 at list price instead of $1.85, which is 1.9 times as much (a calculation).
Questions people ask
How many more tokens does an agent loop use than a single call? In this test, a median of 1.3 to 8.7 times as many, depending on the model. The ranges do not overlap in any pair.
Why do AI agents use so many tokens? We saw repeated input on each turn and a few long attempts. Input also rose when no tool ran. Prompt and tool definitions may contribute, but we did not isolate their effect.
Is an agent loop slower than a single call? Not clearly. The time ranges overlap in all three pairs. The tail is longer: Haiku's slowest loop attempt took 223.7 s.
Is an agent loop worth the cost? For Haiku and Sonnet here, it cost 2.1 and 1.9 times as much per strict pass. The data cannot separate the pass gain from noise. Sonnet had no gain to make. Only GPT-6 Luna got cheaper per pass, and we do not credit the code.
Is it cheaper to use a bigger model than to loop a smaller one? On this set, yes. Sonnet as one call cost $0.0143 per pass. Haiku as a loop cost $0.1422. Each row is a CLI + model pair, and the Claude single-call cells came from an earlier batch.
Does the answer change if a format miss counts as right? No. On the lenient reading, Haiku's calculated cost per correct answer went from $0.0462 (16 of 24, 95% interval 47–82%) to $0.0925 (20 of 24, 64–93%). That is 2.0 times (a calculation). The intervals overlap.
How we measured
- Protocol. The file was created on 2026-10-06 at 21:30 UTC, before the new calls but after the reused Claude reference calls. It was modified on 2026-10-07 at 00:37 UTC, after the runs. File times do not prove that all current text existed before inference. Later changes appear as dated amendments. 40 Claude sessions and 1 probe; 32 Codex calls and 2 probes. We retried nothing, trimmed nothing and hit no usage limit.
- Tasks and grading. The same prompts, validators, strict grading and lenient extractor as the hard head-to-head. Controls ran first: 8 of 8 reference answers pass and 26 of 26 planted wrong answers fail.
- Sandbox. Claude Code ran with its sandbox on, no network and no MCP servers. Codex ran
codex execwith the workspace-write sandbox. - Audit. We audited all 56 loop transcripts. 0 edits landed outside the work folder. 2 GPT-6 Luna attempts read a notes folder outside it. We kept those 2 in the receipts and left them out of every rate and every cost.
- Cost. We compute cost as reported tokens times list price. Cache reads use the cache-read price. Cache writes use the one-hour write price.
- Prices per million tokens, for input, output, cache read and cache write: Haiku 4.5 $1, $5, $0.10 and $2. Sonnet 5.5 $2, $10, $0.20 and $4. GPT-6 Luna $0.10, $0.50, $0.01 and $0.10. Anthropic prices are the list of 2026-09-21. OpenAI prices are the list of 2026-10-03.
- Calculation marks a number we derived from reported tokens, times or list prices. It is not a measured value and it has no interval. Recompute the new cells from the loop and Luna receipts. The two reused Claude single-call cells come from the hard head-to-head receipts.
- Intervals are 95% Wilson intervals for rates. Times and tokens carry fastest-to-slowest or fewest-to-most ranges. The rate intervals treat attempts as independent trials. Repeats on these fixed tasks do not measure success across coding work.
- Medians and sums. Ratios of tokens and time compare medians. Costs and per-pass figures use sums, so a few long attempts move them more.
Caveats
- Different hours. The two Claude single-call cells ran earlier on 2026-10-06 (03:23 to 04:02 UTC), on the same Claude Code version, tasks and validators, but through the other subscription account. The Claude loops ran from 22:06 to 22:39 UTC. Provider load can differ by hour.
- Runners differ beyond tools. The loop runner uses other flags and one more paragraph. Prompt and tool-definition changes may contribute to the token gap. We did not isolate their effect.
- CLI and model pairs. Claude Code and the Codex CLI add their own prompts and tool schemas. The Codex CLI also loads the account's instruction file. A gap between a Claude row and a GPT-6 Luna row is partly the CLI.
- Two Luna routes. The single call and the loop used different Codex routes. Also, 12 of 14 loop attempts ran no tool. Both DST loop attempts were excluded, so the scored task sets differ.
- Contention. Other local jobs ran on the same Mac. Wall times may run high. Token counts are not affected.
- Small samples. 14 to 24 attempts per cell and 2 or 3 per task. The per-task rows find failures. They do not rank tasks.
- A ceiling. Sonnet passed everything both ways, so this set cannot show whether a loop helps a model that already passes.
- Costs are calculations, at list price. The calls ran on flat subscriptions, so no cost here is a bill.
- Disclosure. I build Agent, a product that runs agent loops and routes work between models. I have an interest in loops looking good, and in cheap routes looking good. The numbers above favour neither, and we publish them as measured. The receipts and the method are public.
What to read next
- Does letting an AI run code help? A single call vs an agent loop on 8 hard tasks
- Hard tasks: Claude Haiku vs Sonnet vs Opus vs Fable
- How much does prompt caching save?
- Claude Code vs Codex CLI: the context tax
- Claude Haiku vs Claude Sonnet: every measured row
See what your loop costs
Agent records tokens, cache hits and cost for every stage of every task. Try Agent. Compare a single call with a loop on your own work.