• Agent Loop
  • Tool Use
  • Single Call
  • Hard Tasks
  • Claude Haiku
  • Claude Sonnet
  • Gpt 6 Luna
  • Claude Code
  • Codex CLI
  • Sandbox

Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks

On 8 hard tasks with strict validators, does an agent loop that may write and run code in a sandbox pass more often than one call with tools off, and what does the loop cost in time, tokens, tool calls and list price per pass?

Published · 6 charts · Download the data or a carousel

54%

95% CI 35%–72% · n = 24

13/24 · Claude Haiku 4.5 strict pass rate, agent loop

Single call: 11/24 (28% to 65%); the intervals overlap, so there is no clear difference.

The answer

Not clearly, on this set: no model's agent loop is ahead of its single call by the 95% intervals. Haiku 4.5: single call 11/24 (28% to 65%), agent loop 13/24 (35% to 72%); the intervals overlap, so there is no clear difference. Sonnet 5.5: single call 24/24 (86% to 100%), agent loop 16/16 (81% to 100%); both at the ceiling, so the set cannot separate them. GPT-6 Luna (Codex CLI): single call 10/16 (39% to 82%), agent loop 12/14 (60% to 96%); the intervals overlap, so there is no clear difference. Tools were optional. Scored agent-loop attempts that ran at least one tool: Haiku 4.5 24 of 24, Sonnet 5.5 3 of 16 and GPT-6 Luna (Codex CLI) 2 of 14. Median total time per attempt, single call to agent loop. Haiku 4.5: 39.0 s to 56.8 s (1.5×; the ranges overlap). Sonnet 5.5: 7.7 s to 7.4 s (about the same; the ranges overlap). GPT-6 Luna (Codex CLI): 5.2 s to 9.3 s (1.8×; the ranges overlap). Median tokens per attempt (input with cache reads, plus output), single call to agent loop. Haiku 4.5: 9,038 to 78,432. Sonnet 5.5: 3,380 to 10,483. GPT-6 Luna (Codex CLI): 11,954 to 16,058. List-price cost per strict pass, single call to agent loop. This is a calculation; the calls ran on subscriptions. Haiku 4.5: $0.0672 to $0.1422. Sonnet 5.5: $0.0143 to $0.0275. GPT-6 Luna (Codex CLI): $0.0012 to $0.0010. 2 of 56 agent-loop attempts read a file outside their work folder and are left out; 0 edits landed outside it (the CLI refused 2 outside edit or read attempts before they ran).

Key numbers

100% (16/16)

Claude Sonnet 5.5 strict pass rate, agent loop

95% CI 81%–100% · n = 16

86% (12/14)

GPT-6 Luna strict pass rate, agent loop (Codex CLI)

95% CI 60%–96% · n = 14

+8 points

Change in strict pass rate, agent loop minus single call, Claude Haiku 4.5 (calculation)

n = 48

29 of 54

Scored agent-loop attempts that ran at least one tool

n = 54

0

Edits outside the work folder that ran, in agent-loop attempts

in 56 attempts; 2 outside attempts were refused before they ran · n = 56

2 of 56

Agent-loop attempts left out for reading outside the work folder

n = 56

1.5

Median tool calls per agent-loop attempt

n = 54

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Claude Haiku 4.5 (single call)
Claude Haiku 4.5 (agent loop)
Claude Sonnet 5.5 (single cal…
Claude Sonnet 5.5 (agent loop)
GPT-6 Luna (single call)
GPT-6 Luna (agent loop)

6 rows. Highest Claude Sonnet 5.5 (single call) · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 (single call) · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 14–24 per row

Same tasks and validators. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)

Claude Haiku 4.5 (single call) · Claude Code

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

Claude Haiku 4.5 (agent loop) · Claude Code

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

Claude Sonnet 5.5 (single call) · Claude Code

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

Claude Sonnet 5.5 (agent loop) · Claude Code

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

GPT-6 Luna (single call) · Codex CLI

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

GPT-6 Luna (agent loop) · Codex CLI

Interval merge fix
DST day-length fix
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQL report

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 6 series: Claude Haiku 4.5 (single call) · Claude Code, Claude Haiku 4.5 (agent loop) · Claude Code, Claude Sonnet 5.5 (single call) · Claude Code, Claude Sonnet 5.5 (agent loop) · Claude Code, GPT-6 Luna (single call) · Codex CLI, GPT-6 Luna (agent loop) · Codex CLI. Claude Haiku 4.5 (single call) · Claude Code: highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest SQL report 0% (95% interval 0%–56%, n 3). All intervals overlap. Claude Haiku 4.5 (agent loop) · Claude Code: highest Event-loop order 100% (95% interval 44%–100%, n 3). Lowest SQL report 0% (95% interval 0%–56%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 2–3 per row

Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration

Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
Entrance: medians race at 41× real time
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

6 rows. Slowest Claude Haiku 4.5 (agent loop) · Claude Code 56.8 s (range 24.5 s–224 s, n 24). Fastest GPT-6 Luna (single call) · Codex CLI 5.2 s (range 3.6 s–11.3 s, n 16). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 14–24 per row

Median per configuration; whiskers = fastest and slowest attempt

Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
  • Input tokens (cache reads included)
  • Output tokens
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI

6 rows, 2 series: Input tokens (cache reads included), Output tokens. Input tokens (cache reads included): highest Claude Haiku 4.5 (agent loop) · Claude Code 71,691 (range 41,732–516,306, n 24). Lowest Claude Sonnet 5.5 (single call) · Claude Code 2,281 (range 2,234–2,669, n 24). Not all run ranges overlap. Output tokens: highest Claude Haiku 4.5 (agent loop) · Claude Code 7,912 (range 2,541–20,654, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 345 (range 36–634, n 16). Not all run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 14–24 per row

Median per configuration; whiskers = fewest and most

Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
Claude Haiku 4.5 (agent loop)
Claude Sonnet 5.5 (agent loop)
GPT-6 Luna (agent loop)

3 rows. Highest Claude Haiku 4.5 (agent loop) · Claude Code 3 (range 2–18, n 24). Lowest GPT-6 Luna (agent loop) · Codex CLI 0 (range 0–1, n 14). Not all run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 14–24 per row

Median per configuration; whiskers = fewest and most. A single call makes none

Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.

Source: Single call vs agent loop

Share card (PNG)
Calculation
Largest value is 140x the smallest; Log shows the small bars.
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI

Hover or focus a bar for its ratio to GPT-6 Luna (agent loop) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 6 rows. Highest Claude Haiku 4.5 (agent loop) · Claude Code $0.14 (n 24). Lowest GPT-6 Luna (agent loop) · Codex CLI $0.00099 (n 14).

Notesn 14–24 per row

All attempts in a configuration divided by its strict passes

Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.

Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)

Tables

Every single-call and agent-loop cell

ConfigurationCellStrict passes95% intervalAnswer correct (lenient)Format missesWrong answersNo answer (error, time-out, turn limit)Median total (s)Fastest to slowest (s)Median tool callsAttempts that ran a toolMedian total tokensUSD per strict pass (calculation)Contaminated (left out)
Claude Haiku 4.5 (single call) · Claude Codereference (hard head-to-head)11/2428% to 65%16/2458039 s15.3 to 75.10tools off9,038$0.0670
Claude Haiku 4.5 (agent loop) · Claude Codenew run13/2435% to 72%20/2474056.8 s24.5 to 223.7324/2478,432$0.140
Claude Sonnet 5.5 (single call) · Claude Codereference (hard head-to-head)24/2486% to 100%24/240007.8 s2.3 to 34.80tools off3,380$0.0140
Claude Sonnet 5.5 (agent loop) · Claude Codenew run16/1681% to 100%16/160007.4 s2.7 to 24.203/1610,483$0.0270
GPT-6 Luna (single call) · Codex CLInew run10/1639% to 82%11/161505.2 s3.6 to 11.30tools off11,954$0.00120
GPT-6 Luna (agent loop) · Codex CLInew run12/1460% to 96%12/140209.3 s3.8 to 15.902/1416,058$0.000992

Sandbox audit of every agent-loop attempt

ConfigurationAttemptsEdits outside the work folder that ranOutside edit or read attempts refused before they ranContaminated attempts (read outside the work folder)Which (task, repetition)Time-outs (10 min)Stopped at 30 turns
Claude Haiku 4.5 (agent loop) · Claude Code24020none00
Claude Sonnet 5.5 (agent loop) · Claude Code16000none00
GPT-6 Luna (agent loop) · Codex CLI16002DST day-length fix r1, DST day-length fix r200

Method

  1. Protocol declared before the first counted call. The same 8 tasks, prompts, validators and strict grading as the hard head-to-head (/benchmarks/hard-model-head-to-head).
  2. Single call: the task prompt, tools off, one turn, 300 s timeout. Agent loop: the same prompt plus one paragraph ("You may create and run files in the current folder to test your answer. Your final message must be only the answer, in the format stated above."). The final message is graded exactly like a single call.
  3. New cells: Claude Haiku 4.5 (agent loop) · Claude Code (24); Claude Sonnet 5.5 (agent loop) · Claude Code (16); GPT-6 Luna (single call) · Codex CLI (16); GPT-6 Luna (agent loop) · Codex CLI (16). Reference cells: Claude Haiku 4.5 (single call) · Claude Code (24); Claude Sonnet 5.5 (single call) · Claude Code (24), from the hard head-to-head. One call or session at a time per account; order rep-major, then task, then configuration.
  4. Agent-loop sandbox: a fresh empty work folder per attempt, outside the temp folder. Claude Code: tools Bash, Read, Edit, Write, Glob and Grep only, the Claude Code sandbox on (writes only in the work folder, no network, no unsandboxed commands), no MCP servers, no user settings, at most 30 turns, the same 16,000-token output cap per response as the single calls. Codex CLI: exec with the workspace-write sandbox (network off), approvals never. 10-minute limit per session.
  5. Sandbox probes before the matrix (not counted): Claude Code: network blocked (the command was refused before it ran), write to the parent folder refused, file in the work folder created; Codex CLI: network blocked, write to the parent folder refused, file in the work folder created.
  6. Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.
  7. Audit of every agent-loop transcript: file-edit tool calls, read tool paths and paths in shell commands. An attempt that read a file outside its work folder is contaminated: kept in the raw extract, left out of the rates. Edits outside the work folder that ran must be 0; a call the CLI refused before it ran is counted as an attempt, not an access. The reference answers were locked (no read access) while agents ran.
  8. Stop rules: stop a route at the first usage-limit or rate-limit message; errors, time-outs and turn-limit stops count as fails. No batch stopped early and nothing was trimmed or retried.
  9. Cost per strict pass: list price × reported tokens for every attempt in the cell, divided by its strict passes. A calculation.

Caveats

  • The Claude single-call cells ran in another batch on 2026-10-06 (03:23 to 04:02 UTC), with the same CLI version, tasks and validators; provider load can differ by hour.
  • Each row is a CLI + model pair. Claude Code and Codex CLI add their own system prompts and tool schemas, and Codex CLI also loads the account’s user-level instruction file. A gap between Claude and GPT-6 Luna rows is partly the CLI.
  • The loop changes time and tokens as well as passes. A higher pass rate that costs several times the time and tokens is a trade, not a free gain.
  • Only 2 or 3 attempts per task and configuration (n = 14, 16 and 24 per cell). Read the intervals; per-task bars are for finding failures, not for ranking.
  • Claude Sonnet 5.5 (single call) · Claude Code and Claude Sonnet 5.5 (agent loop) · Claude Code passed every attempt: the set has a ceiling for these configurations, so it cannot show whether the loop helps a model that already passes.
  • List-price costs are calculations; the calls used flat subscriptions.
  • 2 agent-loop attempts read a file outside the work folder and are left out of every rate (GPT-6 Luna (agent loop) · Codex CLI: DST day-length fix r1, DST day-length fix r2), so that cell has fewer attempts and no result for those task repetitions. They stay in the raw extract. The cause was most likely the account’s user-level Codex instructions.

Sources

  • Single call vs agent loop

    Our recorded runs ·

    Receipts of the single call vs agent loop study (public-runs/single-call-vs-agent-loop). Every attempt is kept, failures and contaminated attempts included. Reference single-call cells (Claude Haiku 4.5 and Claude Sonnet 5.5, default effort) are read from raw/provider-h2h-hard.

    Raw data: agent-loop/receipts.json

  • Provider head-to-head, hard set: eight hard tasks with strict validators

    Our recorded runs ·

    Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.

    Raw data: provider-h2h-hard/receipts.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/single-call-vs-agent-loop.

Models and comparisons in this study

More studies

All benchmarks

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.