{"i":16,"study":{"slug":"single-call-vs-agent-loop","title":"Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks","seoTitle":"Agent loop vs single call: does tool use improve accuracy?","description":"118 attempts on 8 hard tasks: one call vs an agent loop that runs code in a sandbox. Pass rate, time, tokens and cost.","question":"On 8 hard tasks with strict validators, does an agent loop that may write and run code in a sandbox pass more often than one call with tools off, and what does the loop cost in time, tokens, tool calls and list price per pass?","answer":"Not clearly, on this set: no model's agent loop is ahead of its single call by the 95% intervals. Haiku 4.5: single call 11/24 (28% to 65%), agent loop 13/24 (35% to 72%); the intervals overlap, so there is no clear difference. Sonnet 5.5: single call 24/24 (86% to 100%), agent loop 16/16 (81% to 100%); both at the ceiling, so the set cannot separate them. GPT-6 Luna (Codex CLI): single call 10/16 (39% to 82%), agent loop 12/14 (60% to 96%); the intervals overlap, so there is no clear difference. Tools were optional. Scored agent-loop attempts that ran at least one tool: Haiku 4.5 24 of 24, Sonnet 5.5 3 of 16 and GPT-6 Luna (Codex CLI) 2 of 14. Median total time per attempt, single call to agent loop. Haiku 4.5: 39.0 s to 56.8 s (1.5×; the ranges overlap). Sonnet 5.5: 7.7 s to 7.4 s (about the same; the ranges overlap). GPT-6 Luna (Codex CLI): 5.2 s to 9.3 s (1.8×; the ranges overlap). Median tokens per attempt (input with cache reads, plus output), single call to agent loop. Haiku 4.5: 9,038 to 78,432. Sonnet 5.5: 3,380 to 10,483. GPT-6 Luna (Codex CLI): 11,954 to 16,058. List-price cost per strict pass, single call to agent loop. This is a calculation; the calls ran on subscriptions. Haiku 4.5: $0.0672 to $0.1422. Sonnet 5.5: $0.0143 to $0.0275. GPT-6 Luna (Codex CLI): $0.0012 to $0.0010. 2 of 56 agent-loop attempts read a file outside their work folder and are left out; 0 edits landed outside it (the CLI refused 2 outside edit or read attempts before they ran).","date":"2026-10-06","updated":"2026-10-06","tags":["agent-loop","tool-use","single-call","hard-tasks","claude-haiku","claude-sonnet","gpt-6-luna","claude-code","codex-cli","sandbox"],"caveats":["The Claude single-call cells ran in another batch on 2026-10-06 (03:23 to 04:02 UTC), with the same CLI version, tasks and validators; provider load can differ by hour.","Each row is a CLI + model pair. Claude Code and Codex CLI add their own system prompts and tool schemas, and Codex CLI also loads the account’s user-level instruction file. A gap between Claude and GPT-6 Luna rows is partly the CLI.","The loop changes time and tokens as well as passes. A higher pass rate that costs several times the time and tokens is a trade, not a free gain.","Only 2 or 3 attempts per task and configuration (n = 14, 16 and 24 per cell). Read the intervals; per-task bars are for finding failures, not for ranking.","Claude Sonnet 5.5 (single call) · Claude Code and Claude Sonnet 5.5 (agent loop) · Claude Code passed every attempt: the set has a ceiling for these configurations, so it cannot show whether the loop helps a model that already passes.","List-price costs are calculations; the calls used flat subscriptions.","2 agent-loop attempts read a file outside the work folder and are left out of every rate (GPT-6 Luna (agent loop) · Codex CLI: DST day-length fix r1, DST day-length fix r2), so that cell has fewer attempts and no result for those task repetitions. They stay in the raw extract. The cause was most likely the account’s user-level Codex instructions."],"sourceIds":["agent-agent-loop","agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["agent-loop-pass-haiku","Claude Haiku 4.5 strict pass rate, agent loop",0.5417,"rate","54% (13/24)",24,[0.3507,0.7211],"Single call: 11/24 (28% to 65%); the intervals overlap, so there is no clear difference."],["agent-loop-pass-sonnet","Claude Sonnet 5.5 strict pass rate, agent loop",1,"rate","100% (16/16)",16,[0.8064,1],"Single call: 24/24 (86% to 100%); both at the ceiling, so the set cannot separate them."],["agent-loop-pass-luna","GPT-6 Luna strict pass rate, agent loop (Codex CLI)",0.8571,"rate","86% (12/14)",14,[0.6006,0.9599],"Single call: 10/16 (39% to 82%); the intervals overlap, so there is no clear difference."],["agent-loop-haiku-gain","Change in strict pass rate, agent loop minus single call, Claude Haiku 4.5 (calculation)",0.0833,"rate","+8 points",48,"\u0001","11/24 to 13/24; the intervals overlap."],["agent-loop-used-tools","Scored agent-loop attempts that ran at least one tool",29,"count","29 of 54",54,"\u0001","Haiku 4.5 24 of 24; Sonnet 5.5 3 of 16; GPT-6 Luna (Codex CLI) 2 of 14"],["agent-loop-outside-edits","Edits outside the work folder that ran, in agent-loop attempts",0,"count","0 in 56 attempts; 2 outside attempts were refused before they ran",56,"\u0001","\u0001"],["agent-loop-contaminated","Agent-loop attempts left out for reading outside the work folder",2,"count","2 of 56",56,"\u0001","\u0001"],["agent-loop-median-tools","Median tool calls per agent-loop attempt",1.5,"count","1.5",54,"\u0001","Haiku 4.5 3; Sonnet 5.5 0; GPT-6 Luna (Codex CLI) 0"]]},"charts":{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","sourceIds"],"$r":[["agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks","Same tasks and validators. Whiskers are 95% Wilson intervals","dot-range","rate","higher","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",0.4583,0.2789,0.6493,24],["Claude Haiku 4.5 (agent loop) · Claude Code",0.5417,0.3507,0.7211,24],["Claude Sonnet 5.5 (single call) · Claude Code",1,0.862,1,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",1,0.8064,1,16],["GPT-6 Luna (single call) · Codex CLI",0.625,0.3864,0.8152,16],["GPT-6 Luna (agent loop) · Codex CLI",0.8571,0.6006,0.9599,14]]}}],"Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.","ci95",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-by-task","Strict passes per task: single call vs agent loop","Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration","grouped-bar","rate","higher","Passed",{"$k":["name","points"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.4385,1,3],["DST day-length fix",0.3333,0.0615,0.7923,3],["CSV parser",0.6667,0.2077,0.9385,3],["Event-loop order",0,0,0.5615,3],["Room schedule",0,0,0.5615,3],["SemVer regex",1,0.4385,1,3],["Money refactor",0.6667,0.2077,0.9385,3],["SQL report",0,0,0.5615,3]]}],["Claude Haiku 4.5 (agent loop) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",0.6667,0.2077,0.9385,3],["DST day-length fix",0.6667,0.2077,0.9385,3],["CSV parser",0.6667,0.2077,0.9385,3],["Event-loop order",1,0.4385,1,3],["Room schedule",0.6667,0.2077,0.9385,3],["SemVer regex",0.6667,0.2077,0.9385,3],["Money refactor",0,0,0.5615,3],["SQL report",0,0,0.5615,3]]}],["Claude Sonnet 5.5 (single call) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.4385,1,3],["DST day-length fix",1,0.4385,1,3],["CSV parser",1,0.4385,1,3],["Event-loop order",1,0.4385,1,3],["Room schedule",1,0.4385,1,3],["SemVer regex",1,0.4385,1,3],["Money refactor",1,0.4385,1,3],["SQL report",1,0.4385,1,3]]}],["Claude Sonnet 5.5 (agent loop) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.3424,1,2],["DST day-length fix",1,0.3424,1,2],["CSV parser",1,0.3424,1,2],["Event-loop order",1,0.3424,1,2],["Room schedule",1,0.3424,1,2],["SemVer regex",1,0.3424,1,2],["Money refactor",1,0.3424,1,2],["SQL report",1,0.3424,1,2]]}],["GPT-6 Luna (single call) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.3424,1,2],["DST day-length fix",1,0.3424,1,2],["CSV parser",1,0.3424,1,2],["Event-loop order",0,0,0.6576,2],["Room schedule",0.5,0.0945,0.9055,2],["SemVer regex",0.5,0.0945,0.9055,2],["Money refactor",0,0,0.6576,2],["SQL report",1,0.3424,1,2]]}],["GPT-6 Luna (agent loop) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.3424,1,2],["CSV parser",0.5,0.0945,0.9055,2],["Event-loop order",1,0.3424,1,2],["Room schedule",1,0.3424,1,2],["SemVer regex",1,0.3424,1,2],["Money refactor",0.5,0.0945,0.9055,2],["SQL report",1,0.3424,1,2]]}]]},"Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.","ci95",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-total-time","Total time per attempt: single call vs agent loop","Median per configuration; whiskers = fastest and slowest attempt","dot-range","seconds","\u0001","Seconds",[{"name":"Total time per attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",39.01,15.27,75.13,24],["Claude Haiku 4.5 (agent loop) · Claude Code",56.77,24.53,223.7,24],["Claude Sonnet 5.5 (single call) · Claude Code",7.75,2.26,34.79,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",7.41,2.75,24.19,16],["GPT-6 Luna (single call) · Codex CLI",5.16,3.59,11.32,16],["GPT-6 Luna (agent loop) · Codex CLI",9.32,3.78,15.89,14]]}}],"Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.","minmax",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-tokens","Tokens per attempt: single call vs agent loop","Median per configuration; whiskers = fewest and most","grouped-bar","tokens","\u0001","Tokens",[{"name":"Input tokens (cache reads included)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",3941,3879,4221,24],["Claude Haiku 4.5 (agent loop) · Claude Code",71691,41732,516306,24],["Claude Sonnet 5.5 (single call) · Claude Code",2281,2234,2669,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",9550,9398,33040,16],["GPT-6 Luna (single call) · Codex CLI",11582,11526,11818,16],["GPT-6 Luna (agent loop) · Codex CLI",15530,15391,39009,14]]}},{"name":"Output tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",5064,1899,9321,24],["Claude Haiku 4.5 (agent loop) · Claude Code",7912,2541,20654,24],["Claude Sonnet 5.5 (single call) · Claude Code",1050,176,3895,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",876,219,3243,16],["GPT-6 Luna (single call) · Codex CLI",345,36,634,16],["GPT-6 Luna (agent loop) · Codex CLI",480,143,858,14]]}}],"Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.","minmax",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-tool-calls","Tool calls per agent-loop attempt","Median per configuration; whiskers = fewest and most. A single call makes none","dot-range","count","\u0001","Tool calls",[{"name":"Tool calls per attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (agent loop) · Claude Code",3,2,18,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",0,0,3,16],["GPT-6 Luna (agent loop) · Codex CLI",0,0,1,14]]}}],"Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.","minmax",["agent-agent-loop"]],["agent-loop-cost-per-pass","List-price cost per strict pass: single call vs agent loop (calculation)","All attempts in a configuration divided by its strict passes","bar","usd","\u0001","USD per strict pass",[{"name":"Cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",0.0672,24],["Claude Haiku 4.5 (agent loop) · Claude Code",0.14225,24],["Claude Sonnet 5.5 (single call) · Claude Code",0.01435,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",0.02746,16],["GPT-6 Luna (single call) · Codex CLI",0.00116,16],["GPT-6 Luna (agent loop) · Codex CLI",0.00099,14]]}}],"Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.","\u0001",["agent-agent-loop","agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]]]},"related":["hard-model-head-to-head","cli-model-latency-tokens"]}}