{"$k":["slug","chart"],"$r":[["single-call-vs-agent-loop",{"id":"agent-loop-pass-rate","title":"Strict pass rate: single call vs agent loop on eight hard tasks","subtitle":"Same tasks and validators. Whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",0.4583,0.2789,0.6493,24],["Claude Haiku 4.5 (agent loop) · Claude Code",0.5417,0.3507,0.7211,24],["Claude Sonnet 5.5 (single call) · Claude Code",1,0.862,1,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",1,0.8064,1,16],["GPT-6 Luna (single call) · Codex CLI",0.625,0.3864,0.8152,16],["GPT-6 Luna (agent loop) · Codex CLI",0.8571,0.6006,0.9599,14]]}}],"note":"Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.","whisker":"ci95","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["single-call-vs-agent-loop",{"id":"agent-loop-by-task","title":"Strict passes per task: single call vs agent loop","subtitle":"Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Passed","series":{"$k":["name","points"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.4385,1,3],["DST day-length fix",0.3333,0.0615,0.7923,3],["CSV parser",0.6667,0.2077,0.9385,3],["Event-loop order",0,0,0.5615,3],["Room schedule",0,0,0.5615,3],["SemVer regex",1,0.4385,1,3],["Money refactor",0.6667,0.2077,0.9385,3],["SQL report",0,0,0.5615,3]]}],["Claude Haiku 4.5 (agent loop) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",0.6667,0.2077,0.9385,3],["DST day-length fix",0.6667,0.2077,0.9385,3],["CSV parser",0.6667,0.2077,0.9385,3],["Event-loop order",1,0.4385,1,3],["Room schedule",0.6667,0.2077,0.9385,3],["SemVer regex",0.6667,0.2077,0.9385,3],["Money refactor",0,0,0.5615,3],["SQL report",0,0,0.5615,3]]}],["Claude Sonnet 5.5 (single call) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.4385,1,3],["DST day-length fix",1,0.4385,1,3],["CSV parser",1,0.4385,1,3],["Event-loop order",1,0.4385,1,3],["Room schedule",1,0.4385,1,3],["SemVer regex",1,0.4385,1,3],["Money refactor",1,0.4385,1,3],["SQL report",1,0.4385,1,3]]}],["Claude Sonnet 5.5 (agent loop) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.3424,1,2],["DST day-length fix",1,0.3424,1,2],["CSV parser",1,0.3424,1,2],["Event-loop order",1,0.3424,1,2],["Room schedule",1,0.3424,1,2],["SemVer regex",1,0.3424,1,2],["Money refactor",1,0.3424,1,2],["SQL report",1,0.3424,1,2]]}],["GPT-6 Luna (single call) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.3424,1,2],["DST day-length fix",1,0.3424,1,2],["CSV parser",1,0.3424,1,2],["Event-loop order",0,0,0.6576,2],["Room schedule",0.5,0.0945,0.9055,2],["SemVer regex",0.5,0.0945,0.9055,2],["Money refactor",0,0,0.6576,2],["SQL report",1,0.3424,1,2]]}],["GPT-6 Luna (agent loop) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.3424,1,2],["CSV parser",0.5,0.0945,0.9055,2],["Event-loop order",1,0.3424,1,2],["Room schedule",1,0.3424,1,2],["SemVer regex",1,0.3424,1,2],["Money refactor",0.5,0.0945,0.9055,2],["SQL report",1,0.3424,1,2]]}]]},"note":"Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.","whisker":"ci95","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["single-call-vs-agent-loop",{"id":"agent-loop-total-time","title":"Total time per attempt: single call vs agent loop","subtitle":"Median per configuration; whiskers = fastest and slowest attempt","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",39.01,15.27,75.13,24],["Claude Haiku 4.5 (agent loop) · Claude Code",56.77,24.53,223.7,24],["Claude Sonnet 5.5 (single call) · Claude Code",7.75,2.26,34.79,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",7.41,2.75,24.19,16],["GPT-6 Luna (single call) · Codex CLI",5.16,3.59,11.32,16],["GPT-6 Luna (agent loop) · Codex CLI",9.32,3.78,15.89,14]]}}],"note":"Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.","whisker":"minmax","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["single-call-vs-agent-loop",{"id":"agent-loop-tokens","title":"Tokens per attempt: single call vs agent loop","subtitle":"Median per configuration; whiskers = fewest and most","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Input tokens (cache reads included)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",3941,3879,4221,24],["Claude Haiku 4.5 (agent loop) · Claude Code",71691,41732,516306,24],["Claude Sonnet 5.5 (single call) · Claude Code",2281,2234,2669,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",9550,9398,33040,16],["GPT-6 Luna (single call) · Codex CLI",11582,11526,11818,16],["GPT-6 Luna (agent loop) · Codex CLI",15530,15391,39009,14]]}},{"name":"Output tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",5064,1899,9321,24],["Claude Haiku 4.5 (agent loop) · Claude Code",7912,2541,20654,24],["Claude Sonnet 5.5 (single call) · Claude Code",1050,176,3895,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",876,219,3243,16],["GPT-6 Luna (single call) · Codex CLI",345,36,634,16],["GPT-6 Luna (agent loop) · Codex CLI",480,143,858,14]]}}],"note":"Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.","whisker":"minmax","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["single-call-vs-agent-loop",{"id":"agent-loop-tool-calls","title":"Tool calls per agent-loop attempt","subtitle":"Median per configuration; whiskers = fewest and most. A single call makes none","kind":"dot-range","unit":"count","yLabel":"Tool calls","series":[{"name":"Tool calls per attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (agent loop) · Claude Code",3,2,18,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",0,0,3,16],["GPT-6 Luna (agent loop) · Codex CLI",0,0,1,14]]}}],"note":"Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.","whisker":"minmax","sourceIds":["agent-agent-loop"]}],["single-call-vs-agent-loop",{"id":"agent-loop-cost-per-pass","title":"List-price cost per strict pass: single call vs agent loop (calculation)","subtitle":"All attempts in a configuration divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",0.0672,24],["Claude Haiku 4.5 (agent loop) · Claude Code",0.14225,24],["Claude Sonnet 5.5 (single call) · Claude Code",0.01435,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",0.02746,16],["GPT-6 Luna (single call) · Codex CLI",0.00116,16],["GPT-6 Luna (agent loop) · Codex CLI",0.00099,14]]}}],"note":"Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.","sourceIds":["agent-agent-loop","agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]}],["llm-speed-anatomy",{"id":"speed-anatomy-first-text","title":"Time to first text: a 250-line answer, six models","subtitle":"Median of 4 calls per model; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds to first text","series":[{"name":"Time to first text","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",4,2.84,6.38,4],["Claude Sonnet 5.5 · Claude Code",1.96,0.88,4.09,4],["Claude Opus 5.5 · Claude Code",1.97,1.7,2.35,4],["Claude Fable 5.1 · Claude Code",4.43,2.27,4.64,4],["GPT-6.1 Sol (low) · Codex CLI",3.52,2.75,4.42,4],["GPT-6 Luna (low) · Codex CLI",3.3,3.19,3.47,4]]}}],"note":"Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-output-speed","title":"Output speed after the first text: visible tokens per second (calculation)","subtitle":"Median of 4 calls per model; whiskers = slowest and fastest call","kind":"dot-range","unit":"tokens","yLabel":"Visible tokens per second","series":[{"name":"Visible tokens per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",153.2,152.6,216.1,4],["Claude Sonnet 5.5 · Claude Code",231.7,230.3,233,4],["Claude Opus 5.5 · Claude Code",155.5,154.6,156.4,4],["Claude Fable 5.1 · Claude Code",122.6,120.9,131.4,4],["GPT-6.1 Sol (low) · Codex CLI",79.6,71.6,80.5,4],["GPT-6 Luna (low) · Codex CLI",129.1,55.5,259.1,4]]}}],"note":"Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-chars-per-second","title":"Output speed in characters per second after the first text (calculation)","subtitle":"Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call","kind":"dot-range","unit":"count","yLabel":"Characters per second","series":[{"name":"Characters per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",547,546,548,3],["Claude Sonnet 5.5 · Claude Code",517,513,519,4],["Claude Opus 5.5 · Claude Code",347,345,349,4],["Claude Fable 5.1 · Claude Code",273,270,293,4],["GPT-6.1 Sol (low) · Codex CLI",323,291,327,4],["GPT-6 Luna (low) · Codex CLI",524,225,1052,4]]}}],"note":"Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}]]}