{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/single-call-vs-agent-loop",
  "study": {
    "slug": "single-call-vs-agent-loop",
    "title": "Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks",
    "seoTitle": "Agent loop vs single call: does tool use improve accuracy?",
    "description": "118 attempts on 8 hard tasks: one call vs an agent loop that runs code in a sandbox. Pass rate, time, tokens and cost.",
    "question": "On 8 hard tasks with strict validators, does an agent loop that may write and run code in a sandbox pass more often than one call with tools off, and what does the loop cost in time, tokens, tool calls and list price per pass?",
    "answer": "Not clearly, on this set: no model's agent loop is ahead of its single call by the 95% intervals. Haiku 4.5: single call 11/24 (28% to 65%), agent loop 13/24 (35% to 72%); the intervals overlap, so there is no clear difference. Sonnet 5.5: single call 24/24 (86% to 100%), agent loop 16/16 (81% to 100%); both at the ceiling, so the set cannot separate them. GPT-6 Luna (Codex CLI): single call 10/16 (39% to 82%), agent loop 12/14 (60% to 96%); the intervals overlap, so there is no clear difference. Tools were optional. Scored agent-loop attempts that ran at least one tool: Haiku 4.5 24 of 24, Sonnet 5.5 3 of 16 and GPT-6 Luna (Codex CLI) 2 of 14. Median total time per attempt, single call to agent loop. Haiku 4.5: 39.0 s to 56.8 s (1.5×; the ranges overlap). Sonnet 5.5: 7.7 s to 7.4 s (about the same; the ranges overlap). GPT-6 Luna (Codex CLI): 5.2 s to 9.3 s (1.8×; the ranges overlap). Median tokens per attempt (input with cache reads, plus output), single call to agent loop. Haiku 4.5: 9,038 to 78,432. Sonnet 5.5: 3,380 to 10,483. GPT-6 Luna (Codex CLI): 11,954 to 16,058. List-price cost per strict pass, single call to agent loop. This is a calculation; the calls ran on subscriptions. Haiku 4.5: $0.0672 to $0.1422. Sonnet 5.5: $0.0143 to $0.0275. GPT-6 Luna (Codex CLI): $0.0012 to $0.0010. 2 of 56 agent-loop attempts read a file outside their work folder and are left out; 0 edits landed outside it (the CLI refused 2 outside edit or read attempts before they ran).",
    "date": "2026-10-06",
    "updated": "2026-10-06",
    "tags": [
      "agent-loop",
      "tool-use",
      "single-call",
      "hard-tasks",
      "claude-haiku",
      "claude-sonnet",
      "gpt-6-luna",
      "claude-code",
      "codex-cli",
      "sandbox"
    ],
    "method": [
      "Protocol declared before the first counted call. The same 8 tasks, prompts, validators and strict grading as the hard head-to-head (/benchmarks/hard-model-head-to-head).",
      "Single call: the task prompt, tools off, one turn, 300 s timeout. Agent loop: the same prompt plus one paragraph (\"You may create and run files in the current folder to test your answer. Your final message must be only the answer, in the format stated above.\"). The final message is graded exactly like a single call.",
      "New cells: Claude Haiku 4.5 (agent loop) · Claude Code (24); Claude Sonnet 5.5 (agent loop) · Claude Code (16); GPT-6 Luna (single call) · Codex CLI (16); GPT-6 Luna (agent loop) · Codex CLI (16). Reference cells: Claude Haiku 4.5 (single call) · Claude Code (24); Claude Sonnet 5.5 (single call) · Claude Code (24), from the hard head-to-head. One call or session at a time per account; order rep-major, then task, then configuration.",
      "Agent-loop sandbox: a fresh empty work folder per attempt, outside the temp folder. Claude Code: tools Bash, Read, Edit, Write, Glob and Grep only, the Claude Code sandbox on (writes only in the work folder, no network, no unsandboxed commands), no MCP servers, no user settings, at most 30 turns, the same 16,000-token output cap per response as the single calls. Codex CLI: exec with the workspace-write sandbox (network off), approvals never. 10-minute limit per session.",
      "Sandbox probes before the matrix (not counted): Claude Code: network blocked (the command was refused before it ran), write to the parent folder refused, file in the work folder created; Codex CLI: network blocked, write to the parent folder refused, file in the work folder created.",
      "Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.",
      "Audit of every agent-loop transcript: file-edit tool calls, read tool paths and paths in shell commands. An attempt that read a file outside its work folder is contaminated: kept in the raw extract, left out of the rates. Edits outside the work folder that ran must be 0; a call the CLI refused before it ran is counted as an attempt, not an access. The reference answers were locked (no read access) while agents ran.",
      "Stop rules: stop a route at the first usage-limit or rate-limit message; errors, time-outs and turn-limit stops count as fails. No batch stopped early and nothing was trimmed or retried.",
      "Cost per strict pass: list price × reported tokens for every attempt in the cell, divided by its strict passes. A calculation."
    ],
    "caveats": [
      "The Claude single-call cells ran in another batch on 2026-10-06 (03:23 to 04:02 UTC), with the same CLI version, tasks and validators; provider load can differ by hour.",
      "Each row is a CLI + model pair. Claude Code and Codex CLI add their own system prompts and tool schemas, and Codex CLI also loads the account’s user-level instruction file. A gap between Claude and GPT-6 Luna rows is partly the CLI.",
      "The loop changes time and tokens as well as passes. A higher pass rate that costs several times the time and tokens is a trade, not a free gain.",
      "Only 2 or 3 attempts per task and configuration (n = 14, 16 and 24 per cell). Read the intervals; per-task bars are for finding failures, not for ranking.",
      "Claude Sonnet 5.5 (single call) · Claude Code and Claude Sonnet 5.5 (agent loop) · Claude Code passed every attempt: the set has a ceiling for these configurations, so it cannot show whether the loop helps a model that already passes.",
      "List-price costs are calculations; the calls used flat subscriptions.",
      "2 agent-loop attempts read a file outside the work folder and are left out of every rate (GPT-6 Luna (agent loop) · Codex CLI: DST day-length fix r1, DST day-length fix r2), so that cell has fewer attempts and no result for those task repetitions. They stay in the raw extract. The cause was most likely the account’s user-level Codex instructions."
    ],
    "sourceIds": [
      "agent-agent-loop",
      "agent-provider-h2h-hard",
      "calc-repricing",
      "price-anthropic",
      "price-openai"
    ],
    "stats": [
      {
        "id": "agent-loop-pass-haiku",
        "label": "Claude Haiku 4.5 strict pass rate, agent loop",
        "value": 0.5417,
        "unit": "rate",
        "display": "54% (13/24)",
        "n": 24,
        "ci": [
          0.3507,
          0.7211
        ],
        "note": "Single call: 11/24 (28% to 65%); the intervals overlap, so there is no clear difference."
      },
      {
        "id": "agent-loop-pass-sonnet",
        "label": "Claude Sonnet 5.5 strict pass rate, agent loop",
        "value": 1,
        "unit": "rate",
        "display": "100% (16/16)",
        "n": 16,
        "ci": [
          0.8064,
          1
        ],
        "note": "Single call: 24/24 (86% to 100%); both at the ceiling, so the set cannot separate them."
      },
      {
        "id": "agent-loop-pass-luna",
        "label": "GPT-6 Luna strict pass rate, agent loop (Codex CLI)",
        "value": 0.8571,
        "unit": "rate",
        "display": "86% (12/14)",
        "n": 14,
        "ci": [
          0.6006,
          0.9599
        ],
        "note": "Single call: 10/16 (39% to 82%); the intervals overlap, so there is no clear difference."
      },
      {
        "id": "agent-loop-haiku-gain",
        "label": "Change in strict pass rate, agent loop minus single call, Claude Haiku 4.5 (calculation)",
        "value": 0.0833,
        "unit": "rate",
        "display": "+8 points",
        "note": "11/24 to 13/24; the intervals overlap.",
        "n": 48
      },
      {
        "id": "agent-loop-used-tools",
        "label": "Scored agent-loop attempts that ran at least one tool",
        "value": 29,
        "unit": "count",
        "display": "29 of 54",
        "note": "Haiku 4.5 24 of 24; Sonnet 5.5 3 of 16; GPT-6 Luna (Codex CLI) 2 of 14",
        "n": 54
      },
      {
        "id": "agent-loop-outside-edits",
        "label": "Edits outside the work folder that ran, in agent-loop attempts",
        "value": 0,
        "unit": "count",
        "display": "0 in 56 attempts; 2 outside attempts were refused before they ran",
        "n": 56
      },
      {
        "id": "agent-loop-contaminated",
        "label": "Agent-loop attempts left out for reading outside the work folder",
        "value": 2,
        "unit": "count",
        "display": "2 of 56",
        "n": 56
      },
      {
        "id": "agent-loop-median-tools",
        "label": "Median tool calls per agent-loop attempt",
        "value": 1.5,
        "unit": "count",
        "display": "1.5",
        "note": "Haiku 4.5 3; Sonnet 5.5 0; GPT-6 Luna (Codex CLI) 0",
        "n": 54
      }
    ],
    "charts": [
      {
        "id": "agent-loop-pass-rate",
        "title": "Strict pass rate: single call vs agent loop on eight hard tasks",
        "subtitle": "Same tasks and validators. Whiskers are 95% Wilson intervals",
        "kind": "dot-range",
        "unit": "rate",
        "polarity": "higher",
        "yLabel": "Passed",
        "series": [
          {
            "name": "Strict pass",
            "points": [
              {
                "label": "Claude Haiku 4.5 (single call) · Claude Code",
                "value": 0.4583,
                "lo": 0.2789,
                "hi": 0.6493,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (agent loop) · Claude Code",
                "value": 0.5417,
                "lo": 0.3507,
                "hi": 0.7211,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (single call) · Claude Code",
                "value": 1,
                "lo": 0.862,
                "hi": 1,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (agent loop) · Claude Code",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (single call) · Codex CLI",
                "value": 0.625,
                "lo": 0.3864,
                "hi": 0.8152,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (agent loop) · Codex CLI",
                "value": 0.8571,
                "lo": 0.6006,
                "hi": 0.9599,
                "n": 14
              }
            ]
          }
        ],
        "note": "Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.",
        "whisker": "ci95",
        "sourceIds": [
          "agent-agent-loop",
          "agent-provider-h2h-hard"
        ]
      },
      {
        "id": "agent-loop-by-task",
        "title": "Strict passes per task: single call vs agent loop",
        "subtitle": "Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration",
        "kind": "grouped-bar",
        "unit": "rate",
        "polarity": "higher",
        "yLabel": "Passed",
        "series": [
          {
            "name": "Claude Haiku 4.5 (single call) · Claude Code",
            "points": [
              {
                "label": "Interval merge fix",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "DST day-length fix",
                "value": 0.3333,
                "lo": 0.0615,
                "hi": 0.7923,
                "n": 3
              },
              {
                "label": "CSV parser",
                "value": 0.6667,
                "lo": 0.2077,
                "hi": 0.9385,
                "n": 3
              },
              {
                "label": "Event-loop order",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              },
              {
                "label": "Room schedule",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              },
              {
                "label": "SemVer regex",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "Money refactor",
                "value": 0.6667,
                "lo": 0.2077,
                "hi": 0.9385,
                "n": 3
              },
              {
                "label": "SQL report",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              }
            ]
          },
          {
            "name": "Claude Haiku 4.5 (agent loop) · Claude Code",
            "points": [
              {
                "label": "Interval merge fix",
                "value": 0.6667,
                "lo": 0.2077,
                "hi": 0.9385,
                "n": 3
              },
              {
                "label": "DST day-length fix",
                "value": 0.6667,
                "lo": 0.2077,
                "hi": 0.9385,
                "n": 3
              },
              {
                "label": "CSV parser",
                "value": 0.6667,
                "lo": 0.2077,
                "hi": 0.9385,
                "n": 3
              },
              {
                "label": "Event-loop order",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "Room schedule",
                "value": 0.6667,
                "lo": 0.2077,
                "hi": 0.9385,
                "n": 3
              },
              {
                "label": "SemVer regex",
                "value": 0.6667,
                "lo": 0.2077,
                "hi": 0.9385,
                "n": 3
              },
              {
                "label": "Money refactor",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              },
              {
                "label": "SQL report",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              }
            ]
          },
          {
            "name": "Claude Sonnet 5.5 (single call) · Claude Code",
            "points": [
              {
                "label": "Interval merge fix",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "DST day-length fix",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "CSV parser",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "Event-loop order",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "Room schedule",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "SemVer regex",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "Money refactor",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "SQL report",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              }
            ]
          },
          {
            "name": "Claude Sonnet 5.5 (agent loop) · Claude Code",
            "points": [
              {
                "label": "Interval merge fix",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "DST day-length fix",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "CSV parser",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "Event-loop order",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "Room schedule",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "SemVer regex",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "Money refactor",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "SQL report",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              }
            ]
          },
          {
            "name": "GPT-6 Luna (single call) · Codex CLI",
            "points": [
              {
                "label": "Interval merge fix",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "DST day-length fix",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "CSV parser",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "Event-loop order",
                "value": 0,
                "lo": 0,
                "hi": 0.6576,
                "n": 2
              },
              {
                "label": "Room schedule",
                "value": 0.5,
                "lo": 0.0945,
                "hi": 0.9055,
                "n": 2
              },
              {
                "label": "SemVer regex",
                "value": 0.5,
                "lo": 0.0945,
                "hi": 0.9055,
                "n": 2
              },
              {
                "label": "Money refactor",
                "value": 0,
                "lo": 0,
                "hi": 0.6576,
                "n": 2
              },
              {
                "label": "SQL report",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              }
            ]
          },
          {
            "name": "GPT-6 Luna (agent loop) · Codex CLI",
            "points": [
              {
                "label": "Interval merge fix",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "CSV parser",
                "value": 0.5,
                "lo": 0.0945,
                "hi": 0.9055,
                "n": 2
              },
              {
                "label": "Event-loop order",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "Room schedule",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "SemVer regex",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              },
              {
                "label": "Money refactor",
                "value": 0.5,
                "lo": 0.0945,
                "hi": 0.9055,
                "n": 2
              },
              {
                "label": "SQL report",
                "value": 1,
                "lo": 0.3424,
                "hi": 1,
                "n": 2
              }
            ]
          }
        ],
        "note": "Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.",
        "whisker": "ci95",
        "sourceIds": [
          "agent-agent-loop",
          "agent-provider-h2h-hard"
        ]
      },
      {
        "id": "agent-loop-total-time",
        "title": "Total time per attempt: single call vs agent loop",
        "subtitle": "Median per configuration; whiskers = fastest and slowest attempt",
        "kind": "dot-range",
        "unit": "seconds",
        "yLabel": "Seconds",
        "series": [
          {
            "name": "Total time per attempt",
            "points": [
              {
                "label": "Claude Haiku 4.5 (single call) · Claude Code",
                "value": 39.01,
                "lo": 15.27,
                "hi": 75.13,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (agent loop) · Claude Code",
                "value": 56.77,
                "lo": 24.53,
                "hi": 223.7,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (single call) · Claude Code",
                "value": 7.75,
                "lo": 2.26,
                "hi": 34.79,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (agent loop) · Claude Code",
                "value": 7.41,
                "lo": 2.75,
                "hi": 24.19,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (single call) · Codex CLI",
                "value": 5.16,
                "lo": 3.59,
                "hi": 11.32,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (agent loop) · Codex CLI",
                "value": 9.32,
                "lo": 3.78,
                "hi": 15.89,
                "n": 14
              }
            ]
          }
        ],
        "note": "Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-agent-loop",
          "agent-provider-h2h-hard"
        ]
      },
      {
        "id": "agent-loop-tokens",
        "title": "Tokens per attempt: single call vs agent loop",
        "subtitle": "Median per configuration; whiskers = fewest and most",
        "kind": "grouped-bar",
        "unit": "tokens",
        "yLabel": "Tokens",
        "series": [
          {
            "name": "Input tokens (cache reads included)",
            "points": [
              {
                "label": "Claude Haiku 4.5 (single call) · Claude Code",
                "value": 3941,
                "lo": 3879,
                "hi": 4221,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (agent loop) · Claude Code",
                "value": 71691,
                "lo": 41732,
                "hi": 516306,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (single call) · Claude Code",
                "value": 2281,
                "lo": 2234,
                "hi": 2669,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (agent loop) · Claude Code",
                "value": 9550,
                "lo": 9398,
                "hi": 33040,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (single call) · Codex CLI",
                "value": 11582,
                "lo": 11526,
                "hi": 11818,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (agent loop) · Codex CLI",
                "value": 15530,
                "lo": 15391,
                "hi": 39009,
                "n": 14
              }
            ]
          },
          {
            "name": "Output tokens",
            "points": [
              {
                "label": "Claude Haiku 4.5 (single call) · Claude Code",
                "value": 5064,
                "lo": 1899,
                "hi": 9321,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (agent loop) · Claude Code",
                "value": 7912,
                "lo": 2541,
                "hi": 20654,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (single call) · Claude Code",
                "value": 1050,
                "lo": 176,
                "hi": 3895,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (agent loop) · Claude Code",
                "value": 876,
                "lo": 219,
                "hi": 3243,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (single call) · Codex CLI",
                "value": 345,
                "lo": 36,
                "hi": 634,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (agent loop) · Codex CLI",
                "value": 480,
                "lo": 143,
                "hi": 858,
                "n": 14
              }
            ]
          }
        ],
        "note": "Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-agent-loop",
          "agent-provider-h2h-hard"
        ]
      },
      {
        "id": "agent-loop-tool-calls",
        "title": "Tool calls per agent-loop attempt",
        "subtitle": "Median per configuration; whiskers = fewest and most. A single call makes none",
        "kind": "dot-range",
        "unit": "count",
        "yLabel": "Tool calls",
        "series": [
          {
            "name": "Tool calls per attempt",
            "points": [
              {
                "label": "Claude Haiku 4.5 (agent loop) · Claude Code",
                "value": 3,
                "lo": 2,
                "hi": 18,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (agent loop) · Claude Code",
                "value": 0,
                "lo": 0,
                "hi": 3,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (agent loop) · Codex CLI",
                "value": 0,
                "lo": 0,
                "hi": 1,
                "n": 14
              }
            ]
          }
        ],
        "note": "Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-agent-loop"
        ]
      },
      {
        "id": "agent-loop-cost-per-pass",
        "title": "List-price cost per strict pass: single call vs agent loop (calculation)",
        "subtitle": "All attempts in a configuration divided by its strict passes",
        "kind": "bar",
        "unit": "usd",
        "yLabel": "USD per strict pass",
        "series": [
          {
            "name": "Cost per strict pass",
            "points": [
              {
                "label": "Claude Haiku 4.5 (single call) · Claude Code",
                "value": 0.0672,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (agent loop) · Claude Code",
                "value": 0.14225,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (single call) · Claude Code",
                "value": 0.01435,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (agent loop) · Claude Code",
                "value": 0.02746,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (single call) · Codex CLI",
                "value": 0.00116,
                "n": 16
              },
              {
                "label": "GPT-6 Luna (agent loop) · Codex CLI",
                "value": 0.00099,
                "n": 14
              }
            ]
          }
        ],
        "note": "Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.",
        "sourceIds": [
          "agent-agent-loop",
          "agent-provider-h2h-hard",
          "calc-repricing",
          "price-anthropic",
          "price-openai"
        ]
      }
    ],
    "tables": [
      {
        "id": "agent-loop-cells",
        "title": "Every single-call and agent-loop cell",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "origin",
            "label": "Cell",
            "unit": "text"
          },
          {
            "key": "strict",
            "label": "Strict passes",
            "unit": "text"
          },
          {
            "key": "ci",
            "label": "95% interval",
            "unit": "text"
          },
          {
            "key": "lenient",
            "label": "Answer correct (lenient)",
            "unit": "text"
          },
          {
            "key": "formatMisses",
            "label": "Format misses",
            "unit": "count"
          },
          {
            "key": "wrong",
            "label": "Wrong answers",
            "unit": "count"
          },
          {
            "key": "noAnswer",
            "label": "No answer (error, time-out, turn limit)",
            "unit": "count"
          },
          {
            "key": "medianTotal",
            "label": "Median total (s)",
            "unit": "seconds"
          },
          {
            "key": "rangeTotal",
            "label": "Fastest to slowest (s)",
            "unit": "text"
          },
          {
            "key": "medianTools",
            "label": "Median tool calls",
            "unit": "count"
          },
          {
            "key": "usedTools",
            "label": "Attempts that ran a tool",
            "unit": "text"
          },
          {
            "key": "medianTokens",
            "label": "Median total tokens",
            "unit": "tokens"
          },
          {
            "key": "perPass",
            "label": "USD per strict pass (calculation)",
            "unit": "usd"
          },
          {
            "key": "contaminated",
            "label": "Contaminated (left out)",
            "unit": "count"
          }
        ],
        "rows": [
          {
            "config": "Claude Haiku 4.5 (single call) · Claude Code",
            "origin": "reference (hard head-to-head)",
            "strict": "11/24",
            "ci": "28% to 65%",
            "lenient": "16/24",
            "formatMisses": 5,
            "wrong": 8,
            "noAnswer": 0,
            "medianTotal": 39.01,
            "rangeTotal": "15.3 to 75.1",
            "medianTools": 0,
            "usedTools": "tools off",
            "medianTokens": 9038,
            "perPass": 0.0672,
            "contaminated": 0
          },
          {
            "config": "Claude Haiku 4.5 (agent loop) · Claude Code",
            "origin": "new run",
            "strict": "13/24",
            "ci": "35% to 72%",
            "lenient": "20/24",
            "formatMisses": 7,
            "wrong": 4,
            "noAnswer": 0,
            "medianTotal": 56.77,
            "rangeTotal": "24.5 to 223.7",
            "medianTools": 3,
            "usedTools": "24/24",
            "medianTokens": 78432,
            "perPass": 0.14225,
            "contaminated": 0
          },
          {
            "config": "Claude Sonnet 5.5 (single call) · Claude Code",
            "origin": "reference (hard head-to-head)",
            "strict": "24/24",
            "ci": "86% to 100%",
            "lenient": "24/24",
            "formatMisses": 0,
            "wrong": 0,
            "noAnswer": 0,
            "medianTotal": 7.75,
            "rangeTotal": "2.3 to 34.8",
            "medianTools": 0,
            "usedTools": "tools off",
            "medianTokens": 3380,
            "perPass": 0.01435,
            "contaminated": 0
          },
          {
            "config": "Claude Sonnet 5.5 (agent loop) · Claude Code",
            "origin": "new run",
            "strict": "16/16",
            "ci": "81% to 100%",
            "lenient": "16/16",
            "formatMisses": 0,
            "wrong": 0,
            "noAnswer": 0,
            "medianTotal": 7.41,
            "rangeTotal": "2.7 to 24.2",
            "medianTools": 0,
            "usedTools": "3/16",
            "medianTokens": 10483,
            "perPass": 0.02746,
            "contaminated": 0
          },
          {
            "config": "GPT-6 Luna (single call) · Codex CLI",
            "origin": "new run",
            "strict": "10/16",
            "ci": "39% to 82%",
            "lenient": "11/16",
            "formatMisses": 1,
            "wrong": 5,
            "noAnswer": 0,
            "medianTotal": 5.16,
            "rangeTotal": "3.6 to 11.3",
            "medianTools": 0,
            "usedTools": "tools off",
            "medianTokens": 11954,
            "perPass": 0.00116,
            "contaminated": 0
          },
          {
            "config": "GPT-6 Luna (agent loop) · Codex CLI",
            "origin": "new run",
            "strict": "12/14",
            "ci": "60% to 96%",
            "lenient": "12/14",
            "formatMisses": 0,
            "wrong": 2,
            "noAnswer": 0,
            "medianTotal": 9.32,
            "rangeTotal": "3.8 to 15.9",
            "medianTools": 0,
            "usedTools": "2/14",
            "medianTokens": 16058,
            "perPass": 0.00099,
            "contaminated": 2
          }
        ]
      },
      {
        "id": "agent-loop-audit",
        "title": "Sandbox audit of every agent-loop attempt",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "attempts",
            "label": "Attempts",
            "unit": "count"
          },
          {
            "key": "outsideEdits",
            "label": "Edits outside the work folder that ran",
            "unit": "count"
          },
          {
            "key": "refused",
            "label": "Outside edit or read attempts refused before they ran",
            "unit": "count"
          },
          {
            "key": "contaminated",
            "label": "Contaminated attempts (read outside the work folder)",
            "unit": "count"
          },
          {
            "key": "which",
            "label": "Which (task, repetition)",
            "unit": "text"
          },
          {
            "key": "timedOut",
            "label": "Time-outs (10 min)",
            "unit": "count"
          },
          {
            "key": "maxTurns",
            "label": "Stopped at 30 turns",
            "unit": "count"
          }
        ],
        "rows": [
          {
            "config": "Claude Haiku 4.5 (agent loop) · Claude Code",
            "attempts": 24,
            "outsideEdits": 0,
            "refused": 2,
            "contaminated": 0,
            "which": "none",
            "timedOut": 0,
            "maxTurns": 0
          },
          {
            "config": "Claude Sonnet 5.5 (agent loop) · Claude Code",
            "attempts": 16,
            "outsideEdits": 0,
            "refused": 0,
            "contaminated": 0,
            "which": "none",
            "timedOut": 0,
            "maxTurns": 0
          },
          {
            "config": "GPT-6 Luna (agent loop) · Codex CLI",
            "attempts": 16,
            "outsideEdits": 0,
            "refused": 0,
            "contaminated": 2,
            "which": "DST day-length fix r1, DST day-length fix r2",
            "timedOut": 0,
            "maxTurns": 0
          }
        ]
      }
    ],
    "related": [
      "hard-model-head-to-head",
      "cli-model-latency-tokens"
    ]
  },
  "sources": [
    {
      "id": "agent-provider-h2h-hard",
      "title": "Provider head-to-head, hard set: eight hard tasks with strict validators",
      "kind": "run",
      "date": "2026-10-06",
      "note": "Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.",
      "data": [
        "/benchmarks/raw/provider-h2h-hard/receipts.json"
      ]
    },
    {
      "id": "price-anthropic",
      "title": "Anthropic list prices (Claude models)",
      "kind": "price-list",
      "date": "2026-09-21",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "note": "Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price."
    },
    {
      "id": "price-openai",
      "title": "OpenAI list prices",
      "kind": "price-list",
      "date": "2026-10-03",
      "url": "https://developers.openai.com/api/docs/pricing",
      "note": "Token prices as listed by the vendor on 2026-10-03."
    },
    {
      "id": "calc-repricing",
      "title": "Repricing calculation",
      "kind": "calculation",
      "date": "2026-10-05",
      "note": "Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes."
    },
    {
      "id": "agent-agent-loop",
      "title": "Single call vs agent loop",
      "kind": "run",
      "date": "2026-10-06",
      "data": [
        "/benchmarks/raw/agent-loop/receipts.json"
      ],
      "note": "Receipts of the single call vs agent loop study (public-runs/single-call-vs-agent-loop). Every attempt is kept, failures and contaminated attempts included. Reference single-call cells (Claude Haiku 4.5 and Claude Sonnet 5.5, default effort) are read from raw/provider-h2h-hard."
    }
  ]
}
