• Agent Loop
  • Token Cost
  • Cost Per Pass
  • Single Call

What does an agent loop cost? 1.3 to 8.7 times the median tokens (calculation)

Agent loop cost on 8 hard tasks: 1.3–8.7 times the median tokens and up to 2.1 times the price per strict pass (calculations), with no clear pass-rate gain.

TL;DR

  • Question: what does an agent loop cost next to one model call, and what does the extra spend buy? We ran the same 8 hard tasks both ways for three models, with the same strict validators.
  • Answer: the loop used more tokens in every pair. It cost more per strict pass for two of three models. It gave no clear gain in passes. Haiku's median time went from 39.0 s to 56.8 s (ranges 15.3–75.1 / 24.5–223.7 s; n = 24 each). The ranges overlap.
  • Tokens: median total tokens per attempt, single call to loop. Claude Haiku 4.5 9,038 to 78,432 (8.7 times). Claude Sonnet 5.5 3,380 to 10,483 (3.1 times). GPT-6 Luna in the Codex CLI 11,954 to 16,058 (1.3 times). The ratios are calculations. In every pair the fewest-to-most ranges do not overlap.
  • Price per strict pass, at list price (a calculation): Haiku $0.0672 to $0.1422, Sonnet $0.0143 to $0.0275, GPT-6 Luna $0.0012 to $0.0010. The calls ran on flat subscriptions, so none of this is a bill.
  • What the spend bought: Haiku passed 11 of 24 as one call and 13 of 24 as a loop (95% intervals 28% to 65% and 35% to 72%). Sonnet passed 24 of 24 and 16 of 16 (95% intervals 86% to 100% and 81% to 100%), a ceiling. GPT-6 Luna passed 10 of 16 and 12 of 14 (39% to 82% and 60% to 96%). Every pair of intervals overlaps, so no loop is ahead.
  • Tokens rose even when no tool ran. Sonnet ran no tool in 13 of 16 loop attempts. Even so, every Sonnet loop attempt sent at least 9,398 input tokens. No Sonnet single call sent more than 2,669.
  • A bigger model's single call had a lower calculated price in this sample. Per strict pass, Sonnet as one call cost $0.0143. Haiku as a loop cost $0.1422 (9.9 times, a calculation). Each row is a CLI + model pair.

Read the full study, with charts and data. This post is about cost. For the pass-rate question, read Does letting an AI run code help?

What we tested

We used the eight hard tasks of our hard model head-to-head. Each task has a sandboxed validator. A strict pass needs the whole final message to pass the validator as written. A right answer in the wrong wrapping is a format miss. A format miss does not count as a pass.

  • Single call. The task prompt, tools off, one turn. We had no paid API key, so the CLI with tools off stands in for an API call.
  • Agent loop. The same prompt plus one paragraph that lets the model create and run files in the current folder. The sandbox gave it one empty folder per attempt and no network. Claude Code had 30 turns at most. Each session had 10 minutes.
CellSourceAttempts
Haiku 4.5, single callreference cell, hard head-to-head24 (8 tasks × 3)
Haiku 4.5, agent loopnew24 (8 tasks × 3)
Sonnet 5.5, single callreference cell, hard head-to-head24 (8 tasks × 3)
Sonnet 5.5, agent loopnew16 (8 tasks × 2)
GPT-6 Luna, single callnew16 (8 tasks × 2)
GPT-6 Luna, agent loopnew16 (8 tasks × 2), 2 left out

An attempt is one call or one session. We kept every attempt, failures included. Two GPT-6 Luna loop attempts read a file outside the work folder. We left them out of every scored metric, including time, tokens and cost. Both were on the DST task, so its loop covers seven tasks (n = 14) and its single-call cell covers eight (n = 16).

Tokens: how many more does an agent loop use?

  • Input tokens (cache reads included)
  • Output tokens
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI

6 rows, 2 series: Input tokens (cache reads included), Output tokens. Input tokens (cache reads included): highest Claude Haiku 4.5 (agent loop) · Claude Code 71,691 (range 41,732–516,306, n 24). Lowest Claude Sonnet 5.5 (single call) · Claude Code 2,281 (range 2,234–2,669, n 24). Not all run ranges overlap. Output tokens: highest Claude Haiku 4.5 (agent loop) · Claude Code 7,912 (range 2,541–20,654, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 345 (range 36–634, n 16). Not all run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 14–24 per row

Median per configuration; whiskers = fewest and most

Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Median total tokens per attempt (input with cache reads and writes, plus output). The fewest and most in the cell are in brackets. These ranges are not confidence intervals.

Model (n: single / loop)Single callAgent loopRatio (calculation)
Haiku 4.5 (24 / 24)9,038 (6,075 to 13,248)78,432 (44,273 to 536,960)8.7 times
Sonnet 5.5 (24 / 16)3,380 (2,412 to 6,154)10,483 (9,623 to 36,283)3.1 times
GPT-6 Luna (16 / 14)11,954 (11,616 to 12,314)16,058 (15,534 to 39,292)1.3 times

The ranges do not overlap in any row. A range is not a confidence interval. GPT-6 Luna used different Codex routes for the single call and the loop. Part of its gap may come from the route and not from the loop.

Where do the extra tokens come from? The receipts show repeated input and long attempts. Prompt and tool-definition changes may also contribute, but we did not isolate their effect.

1. A loop sends its context again on every turn. Haiku's 24 loop attempts took 128 turns (3 to 19 per attempt). A loop turn carried about 20,100 input tokens on average. A single call carried about 4,000 (input tokens divided by turns, a calculation).

Of Haiku's loop input, 89% came from cache reads (2,275,566 of 2,569,125 tokens, a calculation). The cache makes the repeated context cheap. It does not make it free.

2. Input rose even when no tool ran. Sonnet ran no tool in 13 of 16 loop attempts, and those 13 took one turn each. Still, every Sonnet loop attempt sent at least 9,398 input tokens. No Sonnet single call sent more than 2,669.

Most likely, Claude Code sends its tool definitions with every request once tools are on. We did not test that. The two Claude runners also differ in their flags.

3. A few attempts run long. On the DST day-length fix, three Haiku attempts took 9, 12 and 19 turns. Together they used 1,114,624 tokens, which is 40% of the tokens of all 24 attempts (a calculation). The longest used 536,960 tokens, 6.8 times the cell median (calculation), and it failed.

Sonnet's three tool-using attempts used 42% of its loop tokens (98,831 of 232,786, a calculation). A median hides this tail. The cost follows the sum.

Time: is an agent loop slower?

Entrance: medians race at 41× real time
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

6 rows. Slowest Claude Haiku 4.5 (agent loop) · Claude Code 56.8 s (range 24.5 s–224 s, n 24). Fastest GPT-6 Luna (single call) · Codex CLI 5.2 s (range 3.6 s–11.3 s, n 16). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 14–24 per row

Median per configuration; whiskers = fastest and slowest attempt

Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Median total time per attempt, single call to loop. Brackets show fastest-to-slowest ranges, not confidence intervals. Ratios are calculations.

Model (n: single / loop)Single callAgent loopRatio (calculation)
Haiku 4.5 (24 / 24)39.0 s (15.3 to 75.1)56.8 s (24.5 to 223.7)1.5 times
Sonnet 5.5 (24 / 16)7.7 s (2.3 to 34.8)7.4 s (2.7 to 24.2)0.96 times
GPT-6 Luna (16 / 14)5.2 s (3.6 to 11.3)9.3 s (3.8 to 15.9)1.8 times

In all three pairs the ranges overlap, so we do not call the loop slower.

Input tokens rose far more than time did. Haiku's median input rose 18.2 times (3,941 to 71,691), a calculation. Input ranges were 3,879–4,221 / 41,732–516,306 tokens (n = 24 each). Its median time rose 1.5 times (calculation). The tail is the exception. Haiku's 95th percentile went from 68.5 s to 176.3 s, and its slowest loop attempt took 223.7 s.

Two limits apply to every time in this post. The Mac that ran the sessions also ran other local jobs, so wall times may carry contention. Token counts do not. The Claude single-call cells also ran in an earlier batch.

Price: what does a strict pass cost?

Calculation
Largest value is 140x the smallest; Log shows the small bars.
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI

Hover or focus a bar for its ratio to GPT-6 Luna (agent loop) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 6 rows. Highest Claude Haiku 4.5 (agent loop) · Claude Code $0.14 (n 24). Lowest GPT-6 Luna (agent loop) · Codex CLI $0.00099 (n 14).

Notesn 14–24 per row

All attempts in a configuration divided by its strict passes

Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Cost per strict pass divides the list-price cost of every attempt in a cell, failures included, by its strict passes. Single call to loop. Both columns are calculations.

Model (n: single / loop)USD per strict passTokens per strict pass (calculation)
Haiku 4.5 (24 / 24)$0.0672 to $0.1422 (2.1 times)20,163 to 213,557 (10.6 times)
Sonnet 5.5 (24 / 16)$0.0143 to $0.0275 (1.9 times)3,391 to 14,549 (4.3 times)
GPT-6 Luna (16 / 14)$0.0012 to $0.0010 (0.85 times)19,044 to 22,625 (1.2 times)

Tokens per strict pass is the total tokens of the cell divided by its strict passes. Haiku's loop used 10.6 times the tokens for each strict pass in this sample (calculation). Its token use rose far more than its pass count.

Where the money goes. Share of list-price cost by token type (calculation: reported tokens times list price). The last column is the share of tokens that are output. Cell sizes are n = 24 / 24 for Haiku, 24 / 16 for Sonnet, and 16 / 14 for Luna. These calculations have no tested uncertainty interval.

CellOutputCache writesCache readsOther inputOutput share of tokens
Haiku 4.5, single call85.4%3.4%0.0%11.2%56.9%
Haiku 4.5, agent loop56.0%31.6%12.3%0.1%7.5%
Sonnet 5.5, single call72.0%25.9%2.0%0.0%30.5%
Sonnet 5.5, agent loop35.0%57.9%7.0%0.0%6.6%
GPT-6 Luna, single call19.2%0.0%8.8%72.0%2.4%
GPT-6 Luna, agent loop29.6%0.0%16.9%53.6%2.6%

Output is 7.5% of Haiku's loop tokens and 56.0% of its cost. Cache writes, the price of storing the growing context, add another 31.6%. In a loop you pay for context and for output. GPT-6 Luna reports no cache writes.

Can any pass rate make the loop pay? For Haiku on this set, no. A Haiku loop attempt cost $0.0770. A single call cost $0.0308. The loop cost 2.5 times as much per attempt (a calculation).

Cost per pass cannot fall below cost per attempt. Even a perfect 24 of 24 would cost $0.0770 per pass at the same spend. The single call cost $0.0672 per pass.

Sonnet: the single call already passed 24 of 24. The loop nearly doubled the calculated cost per pass (1.9 times). It also passed every attempt: 16 of 16.

GPT-6 Luna is the one pair where the price per pass fell, from $0.0012 to $0.0010. Its price per attempt rose 1.2 times (calculation). Its observed pass count went from 10 of 16 (95% interval 39–82%) to 12 of 14 (60–96%). The intervals overlap, and the loop omits DST. The two routes differ, and 12 of 14 loop attempts ran no tool. We do not credit the code.

A bigger model, one call. Sonnet 5.5 as one call passed 24 of 24 at $0.0143 per pass. Haiku 4.5 as a loop passed 13 of 24 at $0.1422 per pass, which is 9.9 times as much (a calculation). Even Haiku's single call cost 4.7 times as much per pass as Sonnet's (calculation) ($0.0672 against $0.0143). Cost per call is not cost per answer. Our post on hard tasks across Claude models has the single-call rows.

What did the extra spend buy?

Claude Haiku 4.5 (single call)
Claude Haiku 4.5 (agent loop)
Claude Sonnet 5.5 (single cal…
Claude Sonnet 5.5 (agent loop)
GPT-6 Luna (single call)
GPT-6 Luna (agent loop)

6 rows. Highest Claude Sonnet 5.5 (single call) · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 (single call) · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 14–24 per row

Same tasks and validators. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

PairSingle callAgent loopVerdict
Haiku 4.511/24 (28% to 65%)13/24 (35% to 72%)Intervals overlap: no clear difference
Sonnet 5.524/24 (86% to 100%)16/16 (81% to 100%)Both at the ceiling: this set cannot separate them
GPT-6 Luna10/16 (39% to 82%)12/14 (60% to 96%)Intervals overlap: no clear difference

Our rule: a side is ahead only when the 95% Wilson intervals do not overlap. None did. For Haiku the gain is 2 attempts in 24 (+8 points, a calculation).

The spend did not fall evenly across tasks. This table shows Haiku's loop, task by task. Each task has 3 attempts per arm. These rows show where the money went; they do not rank tasks. All token shares and list-price costs are calculations. The 95% Wilson intervals are 0–56% for 0/3, 6–79% for 1/3, 21–94% for 2/3, and 44–100% for 3/3. Every task pair overlaps.

TaskStrict passes, single call to loopTurns per attemptLoop tokens (share)Loop list-price cost (share)
Interval merge fix3/3 to 2/33 to 5179,915 (6%)$0.1409 (8%)
DST day-length fix1/3 to 2/39 to 191,114,624 (40%)$0.5326 (29%)
CSV parser2/3 to 2/33 to 4224,815 (8%)$0.2181 (12%)
Event-loop order0/3 to 3/33 to 3165,125 (6%)$0.1627 (9%)
Room schedule0/3 to 2/33 to 6299,719 (11%)$0.2818 (15%)
SemVer regex3/3 to 2/34 to 5212,785 (8%)$0.1286 (7%)
Money refactor2/3 to 0/33 to 6243,736 (9%)$0.1563 (8%)
SQL report0/3 to 0/33 to 7335,524 (12%)$0.2283 (12%)
  • Where the loop had more strict passes: Event-loop order went from 0 of 3 to 3 of 3 for $0.1627 in total. That is $0.0542 per pass (a calculation), the lowest of the six tasks that had a loop pass. The single call had no pass to divide by.
  • Where it cost most: the DST day-length fix took 29% of the cost and gained 1 pass in 3. It was the task that ran longest.
  • Where it had no strict pass: the money refactor and the SQL report passed 0 of 3 each in the loop. They cost $0.3846 together, 21% of the cell (calculation). In 4 of those 6 attempts, the answer was right and the wrapping was wrong.

That last point is general. Haiku's 7 format-miss attempts cost $0.4771, which is 26% of the loop's spend (a calculation). They gave no strict pass.

What to do

These steps are our advice. Each one has the number behind it.

  1. Run the single call first. If it already passes, skip the loop. Sonnet passed 24 of 24 as one call. The loop cost 1.9 times as much per pass (calculation). It also passed all 16 attempts; this set hit a ceiling.
  2. Turn tools on per task, not per account. Input rose even without tool use: Sonnet ranges were 2,234–2,669 (single, n = 24) / 9,398–33,040 (loop, n = 16). The cause was not isolated. Send tasks that have a check to run, such as a program's output, to the loop. Send the rest to a plain call.
  3. Compare cost per strict pass, not cost per call. Haiku's loop cost 2.5 times as much per attempt and 2.1 times as much per pass (calculations).
  4. Cap turns and tokens. Log every attempt that reaches the cap. Three of 24 Haiku attempts used 40% of the tokens. Our 30-turn cap never applied, because the longest attempt used 19 turns.
  5. End with only the answer. Say in the prompt what the last message must hold. Report any grading after prose removal separately from the strict result. A format miss costs the full price of an attempt.
  6. Keep the start of the prompt stable. That lets the cache work. Without a cache, the same Haiku loop tokens would cost $3.60 at list price instead of $1.85, which is 1.9 times as much (a calculation).

Questions people ask

How many more tokens does an agent loop use than a single call? In this test, a median of 1.3 to 8.7 times as many, depending on the model. The ranges do not overlap in any pair.

Why do AI agents use so many tokens? We saw repeated input on each turn and a few long attempts. Input also rose when no tool ran. Prompt and tool definitions may contribute, but we did not isolate their effect.

Is an agent loop slower than a single call? Not clearly. The time ranges overlap in all three pairs. The tail is longer: Haiku's slowest loop attempt took 223.7 s.

Is an agent loop worth the cost? For Haiku and Sonnet here, it cost 2.1 and 1.9 times as much per strict pass. The data cannot separate the pass gain from noise. Sonnet had no gain to make. Only GPT-6 Luna got cheaper per pass, and we do not credit the code.

Is it cheaper to use a bigger model than to loop a smaller one? On this set, yes. Sonnet as one call cost $0.0143 per pass. Haiku as a loop cost $0.1422. Each row is a CLI + model pair, and the Claude single-call cells came from an earlier batch.

Does the answer change if a format miss counts as right? No. On the lenient reading, Haiku's calculated cost per correct answer went from $0.0462 (16 of 24, 95% interval 47–82%) to $0.0925 (20 of 24, 64–93%). That is 2.0 times (a calculation). The intervals overlap.

How we measured

  • Protocol. The file was created on 2026-10-06 at 21:30 UTC, before the new calls but after the reused Claude reference calls. It was modified on 2026-10-07 at 00:37 UTC, after the runs. File times do not prove that all current text existed before inference. Later changes appear as dated amendments. 40 Claude sessions and 1 probe; 32 Codex calls and 2 probes. We retried nothing, trimmed nothing and hit no usage limit.
  • Tasks and grading. The same prompts, validators, strict grading and lenient extractor as the hard head-to-head. Controls ran first: 8 of 8 reference answers pass and 26 of 26 planted wrong answers fail.
  • Sandbox. Claude Code ran with its sandbox on, no network and no MCP servers. Codex ran codex exec with the workspace-write sandbox.
  • Audit. We audited all 56 loop transcripts. 0 edits landed outside the work folder. 2 GPT-6 Luna attempts read a notes folder outside it. We kept those 2 in the receipts and left them out of every rate and every cost.
  • Cost. We compute cost as reported tokens times list price. Cache reads use the cache-read price. Cache writes use the one-hour write price.
  • Prices per million tokens, for input, output, cache read and cache write: Haiku 4.5 $1, $5, $0.10 and $2. Sonnet 5.5 $2, $10, $0.20 and $4. GPT-6 Luna $0.10, $0.50, $0.01 and $0.10. Anthropic prices are the list of 2026-09-21. OpenAI prices are the list of 2026-10-03.
  • Calculation marks a number we derived from reported tokens, times or list prices. It is not a measured value and it has no interval. Recompute the new cells from the loop and Luna receipts. The two reused Claude single-call cells come from the hard head-to-head receipts.
  • Intervals are 95% Wilson intervals for rates. Times and tokens carry fastest-to-slowest or fewest-to-most ranges. The rate intervals treat attempts as independent trials. Repeats on these fixed tasks do not measure success across coding work.
  • Medians and sums. Ratios of tokens and time compare medians. Costs and per-pass figures use sums, so a few long attempts move them more.

Caveats

  • Different hours. The two Claude single-call cells ran earlier on 2026-10-06 (03:23 to 04:02 UTC), on the same Claude Code version, tasks and validators, but through the other subscription account. The Claude loops ran from 22:06 to 22:39 UTC. Provider load can differ by hour.
  • Runners differ beyond tools. The loop runner uses other flags and one more paragraph. Prompt and tool-definition changes may contribute to the token gap. We did not isolate their effect.
  • CLI and model pairs. Claude Code and the Codex CLI add their own prompts and tool schemas. The Codex CLI also loads the account's instruction file. A gap between a Claude row and a GPT-6 Luna row is partly the CLI.
  • Two Luna routes. The single call and the loop used different Codex routes. Also, 12 of 14 loop attempts ran no tool. Both DST loop attempts were excluded, so the scored task sets differ.
  • Contention. Other local jobs ran on the same Mac. Wall times may run high. Token counts are not affected.
  • Small samples. 14 to 24 attempts per cell and 2 or 3 per task. The per-task rows find failures. They do not rank tasks.
  • A ceiling. Sonnet passed everything both ways, so this set cannot show whether a loop helps a model that already passes.
  • Costs are calculations, at list price. The calls ran on flat subscriptions, so no cost here is a bill.
  • Disclosure. I build Agent, a product that runs agent loops and routes work between models. I have an interest in loops looking good, and in cheap routes looking good. The numbers above favour neither, and we publish them as measured. The receipts and the method are public.

See what your loop costs

Agent records tokens, cache hits and cost for every stage of every task. Try Agent. Compare a single call with a loop on your own work.

The data behind this post

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.