• Agent Loop
  • Tool Use
  • Single Call
  • Hard Tasks

Does letting an AI run code help? A single call vs an agent loop on 8 hard tasks

Does tool use improve LLM accuracy? 8 hard tasks, 120 attempts (118 scored): single call vs agent loop for Haiku, Sonnet and GPT-6 Luna. No clear gain.

TL;DR

  • Question: does an agent loop that can write and run code pass more hard tasks than one call with tools off? We ran the same 8 hard tasks, with the same strict validators, both ways.
  • Answer: not clearly. No model's loop was ahead of its single call, because every pair of 95% intervals overlaps. Claude Haiku 4.5: 11 of 24 as one call, 13 of 24 as a loop (28% to 65% against 35% to 72%). Claude Sonnet 5.5: 24 of 24 and 16 of 16 (86% to 100% against 81% to 100%), a ceiling. GPT-6 Luna in the Codex CLI: 10 of 16 and 12 of 14 (39% to 82% against 60% to 96%).
  • The model, not the harness, decided whether to run code. Haiku ran code in 24 of 24 loop attempts. Sonnet did in 3 of 16. GPT-6 Luna did in 2 of 14. For the last two, the "loop" was mostly a single call with tools in reach.
  • The loop used more tokens; time ranges overlap. For Haiku, median total tokens went from 9,038 to 78,432 (8.7 times, a calculation). Median time went from 39.0 s to 56.8 s (the ranges overlap). List-price cost per strict pass went from $0.067 to $0.142 (a calculation).
  • Code ran on the largest observed task change. Haiku's event-loop prediction went from 0 of 3 to 3 of 3. It wrote the program and ran it. Its strict count fell on the money refactor (2 of 3 to 0 of 3). Three attempts per task is too few to rank tasks.
  • No outside edits landed. 56 loop attempts, 0 edits landed outside the work folder. Two attempts read a notes folder outside it and are left out of every rate.

Full study, charts and data: the single call vs agent loop study.

What we tested

A common claim is that an agent harness beats a plain model call. A harness lets the model write code, run it and fix its answer. We wanted a number for that claim.

We used the eight hard tasks of our hard model head-to-head. Each task has a sandboxed validator. The tasks are an interval-merge fix, a DST day-length fix, a CSV parser and an event-loop output prediction. The other four are a room schedule, a strict SemVer regex, a money refactor and a SQL report.

  • Single call. The task prompt, tools off, one turn. This is our stand-in for an API call. We used no paid API key, so the CLI with tools off is the "API call" arm.
  • Agent loop. The same prompt plus one paragraph: "You may create and run files in the current folder to test your answer. Your final message must be only the answer, in the format stated above." We graded the final message exactly like a single call.
  • Six cells. Claude Haiku 4.5 and Claude Sonnet 5.5 in Claude Code, and GPT-6 Luna in the Codex CLI at medium effort. Each model ran both ways. The two Claude single-call cells are reference cells from the hard head-to-head. We reused them and did not rerun them.
CellCell typeAttempts
Haiku 4.5, single callreference24 (8 tasks × 3)
Haiku 4.5, agent loopnew24 (8 tasks × 3)
Sonnet 5.5, single callreference24 (8 tasks × 3)
Sonnet 5.5, agent loopnew16 (8 tasks × 2)
GPT-6 Luna, single callnew16 (8 tasks × 2)
GPT-6 Luna, agent loopnew16 (8 tasks × 2), 2 left out

Did the loop raise the pass rate? Not clearly

Claude Haiku 4.5 (single call)
Claude Haiku 4.5 (agent loop)
Claude Sonnet 5.5 (single cal…
Claude Sonnet 5.5 (agent loop)
GPT-6 Luna (single call)
GPT-6 Luna (agent loop)

6 rows. Highest Claude Sonnet 5.5 (single call) · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 (single call) · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 14–24 per row

Same tasks and validators. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

A strict pass needs the whole final message to pass the validator as written. A right answer in the wrong wrapping is a format miss. Examples are a code fence or a sentence before the code. A format miss does not count.

PairSingle callAgent loopVerdict
Haiku 4.511/24 (28% to 65%)13/24 (35% to 72%)Intervals overlap: no clear difference
Sonnet 5.524/24 (86% to 100%)16/16 (81% to 100%)Both at the ceiling: the set cannot separate them
GPT-6 Luna10/16 (39% to 82%)12/14 (60% to 96%)Intervals overlap: no clear difference

Our rule: a side is ahead only when the 95% Wilson intervals do not overlap. None did here. For Haiku, the gain is 2 attempts out of 24 (+8 points, a calculation). The overlapping intervals do not establish a gain.

The lenient reading counts format misses as right answers. Haiku went from 16 of 24 (47% to 82%) as a single call to 20 of 24 (64% to 93%) as a loop. Those intervals also overlap. Haiku had 8 wrong answers as a single call and 4 as a loop. Its format misses went from 5 to 7. The final message was not "only the answer" in those 7 attempts. Read these as counts from small samples.

Tool use: the model chose whether to run code

Claude Haiku 4.5 (agent loop)
Claude Sonnet 5.5 (agent loop)
GPT-6 Luna (agent loop)

3 rows. Highest Claude Haiku 4.5 (agent loop) · Claude Code 3 (range 2–18, n 24). Lowest GPT-6 Luna (agent loop) · Codex CLI 0 (range 0–1, n 14). Not all run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 14–24 per row

Median per configuration; whiskers = fewest and most. A single call makes none

Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.

Source: Single call vs agent loop

The prompt allowed code. It did not require it. The models chose very differently.

  • Haiku 4.5 ran at least one tool in 24 of 24 attempts. The median was 3 tool calls (2 to 18). It made 49 shell calls and 45 Write calls; 2 Write calls were refused.
  • Sonnet 5.5 ran a tool in 3 of 16 attempts: the DST fix once and the event-loop prediction twice. It answered the other 13 directly.
  • GPT-6 Luna ran one shell command in 2 of 14 scored attempts, the event-loop prediction and the room schedule, both in repetition 2.

So the Sonnet and Luna loop cells are mostly single calls that had tools in reach. This matters for Luna in particular. Its pass rate rose from 10 of 16 to 12 of 14, and 12 of the 14 loop attempts ran no code at all. We cannot credit the rise to code. Luna's scored loop covers seven tasks; the single-call cell covers all eight. Both excluded loop attempts were on the DST task. The two Luna routes also differ in other ways. The single call used our runner's route. The loop used codex exec. The CLI prompt and its length differ too.

Where pass counts rose, and where they fell

Claude Haiku 4.5 (single call) · Claude Code

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

Claude Haiku 4.5 (agent loop) · Claude Code

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

Claude Sonnet 5.5 (single call) · Claude Code

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

Claude Sonnet 5.5 (agent loop) · Claude Code

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

GPT-6 Luna (single call) · Codex CLI

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

GPT-6 Luna (agent loop) · Codex CLI

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 6 series: Claude Haiku 4.5 (single call) · Claude Code, Claude Haiku 4.5 (agent loop) · Claude Code, Claude Sonnet 5.5 (single call) · Claude Code, Claude Sonnet 5.5 (agent loop) · Claude Code, GPT-6 Luna (single call) · Codex CLI, GPT-6 Luna (agent loop) · Codex CLI. Claude Haiku 4.5 (single call) · Claude Code: highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest SQL report 0% (95% interval 0%–56%, n 3). All intervals overlap. Claude Haiku 4.5 (agent loop) · Claude Code: highest Event-loop order 100% (95% interval 44%–100%, n 3). Lowest SQL report 0% (95% interval 0%–56%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 2–3 per row

Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration

Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Each bar rests on 2 or 3 attempts, so every whisker is wide. For Haiku, the 95% Wilson intervals are 0–56% for 0/3, 6–79% for 1/3, 21–94% for 2/3 and 44–100% for 3/3. Every task pair overlaps. Use this chart to find failures, not to rank tasks. For Haiku:

  • Up: event-loop prediction 0/3 to 3/3. In all three attempts it wrote the program into a file, ran it with node and copied the output. Running the program turns a prediction into a measurement. This was the largest observed strict-pass change. The intervals still overlap.
  • Up: room schedule 0/3 to 2/3, and DST fix 1/3 to 2/3. For the schedule, Haiku wrote a script and ran it in all three attempts.
  • Down: money refactor 2/3 to 0/3, interval merge 3/3 to 2/3 and SemVer regex 3/3 to 2/3. Most of these were right answers with prose around them.
  • Same: CSV parser 2/3 both ways. SQL report 0/3 both ways, each with 2 right answers in the wrong wrapping.

In one DST attempt, Haiku used 19 turns and 18 tool calls. It returned prose followed by code. Strict grading rejected the leading prose, and lenient extraction found no passing answer. This attempt failed under both readings.

What the loop costs

Entrance: medians race at 41× real time
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

6 rows. Slowest Claude Haiku 4.5 (agent loop) · Claude Code 56.8 s (range 24.5 s–224 s, n 24). Fastest GPT-6 Luna (single call) · Codex CLI 5.2 s (range 3.6 s–11.3 s, n 16). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 14–24 per row

Median per configuration; whiskers = fastest and slowest attempt

Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

All ranges below show the smallest and largest scored attempt, not a confidence interval.

Modeln, single / loopTime range, single / loop (s)Total-token range, single / loop
Haiku 4.524 / 2415.3–75.1 / 24.5–223.76,075–13,248 / 44,273–536,960
Sonnet 5.524 / 162.3–34.8 / 2.7–24.22,412–6,154 / 9,623–36,283
GPT-6 Luna16 / 143.6–11.3 / 3.8–15.911,616–12,314 / 15,534–39,292

Median total time per attempt, single call to loop: Haiku 39.0 s to 56.8 s (1.5 times, a calculation). Sonnet 7.7 s to 7.4 s (about the same). GPT-6 Luna 5.2 s to 9.3 s (1.8 times, a calculation). All three pairs of fastest-to-slowest ranges overlap. The median does not show the full spread. Haiku's slowest loop attempt took 223.7 s. Its 95th percentile went from 68.5 s to 176.3 s (n = 24 each; calculation by linear interpolation).

  • Input tokens (cache reads included)
  • Output tokens
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI

6 rows, 2 series: Input tokens (cache reads included), Output tokens. Input tokens (cache reads included): highest Claude Haiku 4.5 (agent loop) · Claude Code 71,691 (range 41,732–516,306, n 24). Lowest Claude Sonnet 5.5 (single call) · Claude Code 2,281 (range 2,234–2,669, n 24). Not all run ranges overlap. Output tokens: highest Claude Haiku 4.5 (agent loop) · Claude Code 7,912 (range 2,541–20,654, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 345 (range 36–634, n 16). Not all run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 14–24 per row

Median per configuration; whiskers = fewest and most

Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Tokens are where the loop is clearly different. A loop sends its growing context again on every turn. Median total tokens (input with cache reads, plus output), single call to loop: Haiku 9,038 to 78,432. Sonnet 3,380 to 10,483. GPT-6 Luna 11,954 to 16,058. Part of the rise has nothing to do with running code. Sonnet ran no tool in 13 of 16 attempts, yet its median input rose from 2,281 to 9,550 tokens (ranges 2,234–2,669 and 9,398–33,040). The CLI prompt and tool definitions may explain part of this rise. We did not test the cause. A token count is not good or bad by itself. It is the price of the loop.

Calculation
Largest value is 140x the smallest; Log shows the small bars.
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI

Hover or focus a bar for its ratio to GPT-6 Luna (agent loop) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 6 rows. Highest Claude Haiku 4.5 (agent loop) · Claude Code $0.14 (n 24). Lowest GPT-6 Luna (agent loop) · Codex CLI $0.00099 (n 14).

Notesn 14–24 per row

All attempts in a configuration divided by its strict passes

Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

At list price, cost per strict pass, single call to loop, is a calculation. It divides the cost of every attempt in the cell, failures included, by its strict passes:

ModelSingle callAgent loopChange (calculation)
Haiku 4.5$0.0672$0.14222.1 times
Sonnet 5.5$0.0143$0.02751.9 times
GPT-6 Luna$0.0012$0.00100.85 times

For Haiku, 24 loop attempts cost about $1.85 at list price and 24 single calls about $0.74. The calls ran on flat subscriptions, so none of this is a bill. The costs have no interval, so the Luna change is an untested difference.

Did the sandbox hold?

We ran the loops with write access to one empty folder and no network. Before each CLI's loop batch, we ran one uncounted sandbox probe. Both CLIs blocked the network command, refused a write to the parent folder and created a file in the work folder.

Then we audited all 56 loop transcripts for file edits and for paths in shell commands.

  • 0 edits outside the work folder ran. In 2 Haiku attempts the model tried to write a test file under a temp path outside the folder. Claude Code or the sandbox refused every one of those writes, and no file landed.
  • 2 attempts are contaminated and left out. GPT-6 Luna searched a notes folder outside the work folder on the DST task, in both repetitions. Both attempts passed. The Codex CLI loads the account's own instruction file, but we did not test whether those instructions caused the search. We kept them in the data and left them out of all scored metrics, including time, tokens and cost. The protocol states this exclusion rule, but its file times do not prove when the rule was written.
  • Reference answers were locked. The file that holds the planted answers had no read access while any agent ran.

If a loop reads outside its folder, its score may not be its own work. That is why we audit.

Questions people ask

Does tool use improve LLM accuracy? In this test, not clearly. Haiku gained 2 attempts out of 24 with overlapping intervals. Sonnet was already at 24 of 24 as a single call (95% interval 86–100%). The largest observed task change was on program-output prediction; its task intervals still overlap. It did not fix the final message: the loop raised Haiku's format misses from 5 to 7.

What is the difference between an agent harness and a single call? A single call sends one prompt and takes one reply. A harness adds tools, a loop and a way to read the results. The model can then test and revise. It can add turns, tokens and time. Here, the time ranges overlap.

Will my model use the tools I give it? It depends on the model. Haiku 4.5 used them every time. Sonnet 5.5 and GPT-6 Luna used them rarely and only on some tasks. Check the tool-call count in your own runs.

Is the loop worth the cost? For Haiku here, the loop cost about 2.1 times as much per strict pass (a calculation). The data cannot separate its gain from noise. For Sonnet, which already passed, it cost about 1.9 times as much per strict pass (a calculation). Both Sonnet cells passed every attempt.

What to do

  1. Measure pass, time and tokens together. A loop that adds 2 passes and 8.7 times the median tokens (a calculation) is a trade, not a free gain.
  2. Use a loop where a check can run. Predicting program output and searching a small space were the two tasks where Haiku moved most. Check your own tasks the same way.
  3. Fix the last message. Tell the model to end with only the answer. If you strip prose, report that as lenient grading; keep the strict score. 7 of 24 Haiku loop attempts had the right answer in the wrong wrapping.
  4. Check whether the loop adds value when a single call already passes. Sonnet passed every attempt both ways on this set. The loop used more tokens and cost more per pass (calculation). This ceiling does not show that loops have no value on other tasks.
  5. Sandbox it and audit the transcripts. Give the loop one folder, no network and a list of the paths it may touch. Then read what it did.

How we measured

  • Protocol timing. The file was created at 21:30 UTC on 2026-10-06, before the new calls started at 21:41 UTC. The reused Claude reference calls ran earlier. The file was modified after the runs and includes dated amendments and an outcome log. Its file times do not prove that all current text existed before inference. 40 Claude sessions and 1 probe; 32 Codex calls and 2 probes. Nothing was retried or trimmed, and no usage limit was hit.
  • Tasks and grading. The same prompts, validators, strict grading and lenient extractor as the hard head-to-head. Controls ran before the new inference batch. 8 of 8 reference answers pass. 26 of 26 planted wrong answers fail. 8 of 8 wrapped references are flagged as format misses.
  • Claude Code loop. Tools: shell, read, edit, write, glob and grep. The Claude Code sandbox was on, with no network and no unsandboxed commands. No MCP servers, no user settings, at most 30 turns and the same 16,000-token output cap per response as the single calls.
  • Codex CLI loop. codex exec with the workspace-write sandbox, network off and approvals never. 10 minutes per loop session in both CLIs; single calls had a 300-second timeout.
  • No answer is a fail. An error, a timeout or a turn limit counts as a failed attempt. None happened.
  • Cost is reported tokens times list price. A calculation.
  • Intervals are 95% Wilson intervals for rates and fastest-to-slowest ranges for times, tokens and tool calls. A range is not a confidence interval.

Public receipts: new calls and audit, reference calls, and study data.

Caveats

  • Shared host. Other local jobs may affect wall time. These ranges describe the observed runs, not isolated model speed.
  • Different hours. The Claude single-call cells ran earlier on 2026-10-06 (03:23 to 04:02 UTC), on the same Claude Code version, tasks and validators. Provider load can differ by hour.
  • CLI and model pairs. Claude Code and the Codex CLI add their own system prompts and tool schemas. A gap between Claude and GPT-6 Luna rows is partly the CLI.
  • Two Luna routes. The single call and the loop used different Codex routes, so the Luna comparison mixes tools with route.
  • Small samples. 2 or 3 attempts per task and cell. Wilson intervals treat attempts as independent trials. Repeated attempts on eight fixed tasks do not measure success across coding work. Read the intervals.
  • A ceiling. Sonnet passed everything both ways. This set cannot show whether a loop helps a model that already passes.
  • Two attempts left out. GPT-6 Luna has no scored DST result in the loop cell, so that cell has n = 14. Its single-call cell includes both DST attempts. The scored task mix differs. All scored metrics exclude the two contaminated attempts.
  • Costs are calculations, not bills.
  • Disclosure. I build Agent, a product that runs agent loops, so I have an interest in loops looking good. The result does not show that, and we publish it as measured. The receipts and the method are public.

See what your agent loop does

Agent records every step of a loop with its model, its cost and its time. Try Agent and compare a single call with a loop on your own work.

The data behind this post

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.