{"$k":["slug","chart"],"$r":[["model-head-to-head",{"id":"h2h-pass-rate","title":"Pass rate on five validated tasks","subtitle":"Every call counts; failures and timeouts are non-passes","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Pass rate","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Fable 5.1 · Claude Code",1,0.7961,1,15],["Claude Sonnet 5.5 · Claude Code",0.8,0.5481,0.9295,15],["Claude Opus 5.5 (high) · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 (low) · Claude Code",1,0.7961,1,15],["Claude Haiku 4.5 · Claude Code",1,0.7961,1,15],["GPT-6.1 Sol (high) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (low) · Codex CLI",1,0.7225,1,10]]}}],"note":"Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-total-latency","title":"Total time per call","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.94,1.41,9.83,15,true],["Claude Sonnet 5.5 · Claude Code",2.31,2.17,7.73,15,true],["Claude Opus 5.5 (high) · Claude Code",2.71,2.45,11.78,15,true],["Claude Opus 5.5 · Claude Code",2.75,2.47,8.91,15,true],["Claude Opus 5.5 (low) · Claude Code",2.83,2.35,6.62,15,true],["Claude Haiku 4.5 · Claude Code",4.43,3.16,23.57,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.6,4.05,19.52,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.65,4.1,25.46,15,false],["GPT-6.1 Sol (low) · Codex CLI",6.26,4.65,10.47,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-first-useful-latency","title":"Time to first useful output","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Time to first useful output","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.2,0.95,7.9,15,true],["Claude Sonnet 5.5 · Claude Code",1.56,0.99,6.39,15,true],["Claude Opus 5.5 (high) · Claude Code",2.04,1.4,9.94,15,true],["Claude Opus 5.5 · Claude Code",1.92,1.56,7.23,15,true],["Claude Opus 5.5 (low) · Claude Code",2.39,1.45,4.9,15,true],["Claude Haiku 4.5 · Claude Code",3.63,2.78,22.27,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.32,3.64,16.37,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.05,3.36,17.82,15,false],["GPT-6.1 Sol (low) · Codex CLI",5.14,4.02,8.5,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-input-tokens","title":"Input tokens per call: what the CLI sends","subtitle":"Mean per call, split into prompt-cache reads and other input","kind":"stacked-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Cache read","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",2760,15],["Claude Sonnet 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (high) · Claude Code",1463,15],["Claude Opus 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (low) · Claude Code",1463,15],["Claude Haiku 4.5 · Claude Code",0,15],["GPT-6.1 Sol (high) · Codex CLI",6716,15],["GPT-6.1 Sol (medium) · Codex CLI",5180,15],["GPT-6.1 Sol (low) · Codex CLI",8064,10]]}},{"name":"Other input","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",473,15],["Claude Sonnet 5.5 · Claude Code",685,15],["Claude Opus 5.5 (high) · Claude Code",619,15],["Claude Opus 5.5 · Claude Code",680,15],["Claude Opus 5.5 (low) · Claude Code",618,15],["Claude Haiku 4.5 · Claude Code",3790,15],["GPT-6.1 Sol (high) · Codex CLI",5406,15],["GPT-6.1 Sol (medium) · Codex CLI",6943,15],["GPT-6.1 Sol (low) · Codex CLI",4059,10]]}}],"note":"The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-output-tokens","title":"Output tokens per call","subtitle":"Median per configuration; reasoning tokens where the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",64,15],["Claude Sonnet 5.5 · Claude Code",107,15],["Claude Opus 5.5 (high) · Claude Code",78,15],["Claude Opus 5.5 · Claude Code",64,15],["Claude Opus 5.5 (low) · Claude Code",64,15],["Claude Haiku 4.5 · Claude Code",367,15],["GPT-6.1 Sol (high) · Codex CLI",42,15],["GPT-6.1 Sol (medium) · Codex CLI",42,15],["GPT-6.1 Sol (low) · Codex CLI",42,10]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0,15],["Claude Sonnet 5.5 · Claude Code",0,15],["Claude Opus 5.5 (high) · Claude Code",34,15],["Claude Opus 5.5 · Claude Code",0,15],["Claude Opus 5.5 (low) · Claude Code",0,15],["Claude Haiku 4.5 · Claude Code",297,15],["GPT-6.1 Sol (high) · Codex CLI",21,15],["GPT-6.1 Sol (medium) · Codex CLI",20,15],["GPT-6.1 Sol (low) · Codex CLI",20,10]]}}],"note":"Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-list-price-per-call","title":"List-price cost per call (calculation)","subtitle":"Reported tokens × list price; the calls ran on subscriptions","kind":"dot-range","unit":"usd","yLabel":"USD per call","series":[{"name":"Cost per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",0.00987,0.0049,0.05843,15,true],["Claude Sonnet 5.5 · Claude Code",0.0036,0.00342,0.01021,15,true],["Claude Opus 5.5 (high) · Claude Code",0.00694,0.00592,0.02708,15,true],["Claude Opus 5.5 · Claude Code",0.00688,0.00592,0.02226,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00688,0.00582,0.01793,15,true],["Claude Haiku 4.5 · Claude Code",0.00566,0.00513,0.01804,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01047,0.0066,0.02812,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01018,0.0054,0.02686,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00769,0.00742,0.02649,10,false]]}}],"note":"Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["model-head-to-head",{"id":"h2h-cost-per-pass","title":"List-price cost per passing answer (calculation)","subtitle":"All calls in a configuration, failures included, divided by its passes","kind":"bar","unit":"usd","yLabel":"USD per passing answer","series":[{"name":"Cost per pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.00624,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00829,15,true],["Claude Haiku 4.5 · Claude Code",0.00836,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00998,10,false],["Claude Opus 5.5 · Claude Code",0.01009,15,true],["Claude Opus 5.5 (high) · Claude Code",0.01049,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01322,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01564,15,false],["Claude Fable 5.1 · Claude Code",0.02054,15,true]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["hard-model-head-to-head",{"id":"hard-h2h-pass-rate","title":"Pass rate on eight hard tasks","subtitle":"Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.4583,0.2789,0.6493,24]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.6667,0.4671,0.8203,24]]}}],"note":"Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-total-latency","title":"Total time per call on hard tasks (separate batches)","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call on hard tasks (separate batches)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",7.75,2.26,34.79,24,true],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,24,true],["Claude Opus 5.5 (high) · Claude Code",11.03,3.63,63,24,true],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16,true],["Claude Fable 5.1 · Claude Code",16.13,4.46,90,24,true],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16,true],["Claude Haiku 4.5 · Claude Code",39.01,15.27,75.13,24,false]]}}],"note":"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-first-useful-latency","title":"Time to first useful output on hard tasks","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Time to first useful output on hard tasks","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",5.95,0.86,30.57,24,true],["Claude Opus 5.5 · Claude Code",6.78,2.39,21.77,24,true],["Claude Opus 5.5 (high) · Claude Code",7.13,2.15,56.23,24,true],["GPT-6.1 Sol (medium) · Codex CLI",10.23,6.09,40.41,16,true],["Claude Fable 5.1 · Claude Code",11.63,2,85.33,24,true],["GPT-6.1 Sol (high) · Codex CLI",12.69,8.93,75.91,16,true],["Claude Haiku 4.5 · Claude Code",35.54,12.88,70.31,24,false]]}}],"note":"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-output-tokens","title":"Output tokens per call on hard tasks","subtitle":"Median per configuration; reasoning tokens as the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1050,24],["Claude Opus 5.5 · Claude Code",945,24],["Claude Opus 5.5 (high) · Claude Code",1052,24],["GPT-6.1 Sol (medium) · Codex CLI",335,16],["Claude Fable 5.1 · Claude Code",1366,24],["GPT-6.1 Sol (high) · Codex CLI",436,16],["Claude Haiku 4.5 · Claude Code",5064,24]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",585,24],["Claude Opus 5.5 · Claude Code",529,24],["Claude Opus 5.5 (high) · Claude Code",614,24],["GPT-6.1 Sol (medium) · Codex CLI",150,16],["Claude Fable 5.1 · Claude Code",889,24],["GPT-6.1 Sol (high) · Codex CLI",225,16],["Claude Haiku 4.5 · Claude Code",4556,24]]}}],"note":"Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-cost-per-pass","title":"List-price cost per strict pass on hard tasks (calculation)","subtitle":"All calls in a configuration, failures and format misses included, divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.01435,24,true],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16,false],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16,false],["Claude Opus 5.5 · Claude Code",0.02824,24,false],["Claude Opus 5.5 (high) · Claude Code",0.03337,24,false],["Claude Haiku 4.5 · Claude Code",0.0672,24,false],["Claude Fable 5.1 · Claude Code",0.09331,24,false]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.","sourceIds":["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]}],["coding-agents-head-to-head",{"id":"coding-agents-pass-rate","title":"Coding sessions that passed every hidden check","subtitle":"A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","yLabel":"Passed","viz":"IntervalDotPlot","series":[{"name":"Passed every hidden check","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.7575,1,12],["Claude Opus 5.5 · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",1,0.7575,1,12]]}}],"note":"6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.","sourceIds":["agent-coding-agents"]}],["coding-agents-head-to-head",{"id":"coding-agents-wall-time","title":"Time per coding session","subtitle":"Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)","kind":"dot-range","unit":"seconds","whisker":"minmax","yLabel":"Seconds","viz":"LatencyLanes","series":[{"name":"Wall time per session","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",23.1,18.7,44.5,12,true],["Claude Opus 5.5 · Claude Code",56.9,29.8,185.8,12,"\u0001"],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",113.4,78.5,221.9,12,"\u0001"]]}}],"note":"CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.","sourceIds":["agent-coding-agents"]}],["coding-agents-head-to-head",{"id":"coding-agents-tool-calls","title":"Tool calls per coding session","subtitle":"Median; whiskers = fewest and most of 12 sessions (not an interval)","kind":"dot-range","unit":"calls","whisker":"minmax","yLabel":"Tool calls","viz":"IntervalDotPlot","series":[{"name":"Tool calls per session","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",7.5,3,14,12],["Claude Opus 5.5 · Claude Code",7.5,5,14,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",12.5,8,18,12]]}}],"note":"Claude Code counts its tool calls (Bash, Read, Edit, Write, Glob, Grep). Codex CLI counts shell commands and file changes; it has no separate read tool, so it reads files with shell commands. Turns are not compared: Codex reports one turn per run.","sourceIds":["agent-coding-agents"]}],["coding-agents-head-to-head",{"id":"coding-agents-cost-per-pass","title":"List-price cost per passing coding session (calculation)","subtitle":"Reported tokens of all 12 sessions × list price, divided by the passes","kind":"bar","unit":"usd","yLabel":"USD per pass","viz":"CostBars","series":[{"name":"List-price cost per pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.085,12],["Claude Opus 5.5 · Claude Code",0.2229,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",0.0978,12]]}}],"note":"Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.","sourceIds":["agent-coding-agents","calc-repricing","price-anthropic","price-openai"]}],["effort-ladder",{"id":"effort-ladder-pass-rate","title":"Strict pass rate by effort on eight hard tasks","subtitle":"Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 · Claude Code",1,0.8064,1,16],["GPT-6.1 Sol (low) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16]]}}],"note":"Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.","whisker":"ci95","sourceIds":["agent-effort-ladder"]}],["effort-ladder",{"id":"effort-ladder-total-latency","title":"Total time per call by effort on hard tasks","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call by effort on hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",5.82,2.78,19.96,16],["Claude Sonnet 5.5 (medium) · Claude Code",7.63,2.71,24.01,16],["Claude Sonnet 5.5 (high) · Claude Code",8.81,2.93,35.81,16],["Claude Sonnet 5.5 · Claude Code",7.97,2.26,21.61,16],["Claude Opus 5.5 (low) · Claude Code",7.5,3.34,15.82,16],["Claude Opus 5.5 (medium) · Claude Code",9.72,4.78,31.36,16],["Claude Opus 5.5 (high) · Claude Code",10.11,3.63,63,16],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,16],["GPT-6.1 Sol (low) · Codex CLI",13.62,7.94,44.29,16],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16]]}}],"note":"Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.","whisker":"minmax","sourceIds":["agent-effort-ladder"]}],["effort-ladder",{"id":"effort-ladder-output-tokens","title":"Output tokens per call by effort on hard tasks","subtitle":"Median per configuration; reasoning tokens as the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",667,16],["Claude Sonnet 5.5 (medium) · Claude Code",770,16],["Claude Sonnet 5.5 (high) · Claude Code",1192,16],["Claude Sonnet 5.5 · Claude Code",1054,16],["Claude Opus 5.5 (low) · Claude Code",594,16],["Claude Opus 5.5 (medium) · Claude Code",853,16],["Claude Opus 5.5 (high) · Claude Code",1052,16],["Claude Opus 5.5 · Claude Code",945,16],["GPT-6.1 Sol (low) · Codex CLI",284,16],["GPT-6.1 Sol (medium) · Codex CLI",335,16],["GPT-6.1 Sol (high) · Codex CLI",436,16]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",273,16],["Claude Sonnet 5.5 (medium) · Claude Code",422,16],["Claude Sonnet 5.5 (high) · Claude Code",745,16],["Claude Sonnet 5.5 · Claude Code",668,16],["Claude Opus 5.5 (low) · Claude Code",87,16],["Claude Opus 5.5 (medium) · Claude Code",518,16],["Claude Opus 5.5 (high) · Claude Code",614,16],["Claude Opus 5.5 · Claude Code",538,16],["GPT-6.1 Sol (low) · Codex CLI",63,16],["GPT-6.1 Sol (medium) · Codex CLI",150,16],["GPT-6.1 Sol (high) · Codex CLI",225,16]]}}],"note":"Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean \"not reported\". Their content is never captured. More tokens is not better or worse by itself.","sourceIds":["agent-effort-ladder"]}],["effort-ladder",{"id":"effort-ladder-cost-per-pass","title":"List-price cost per strict pass by effort (calculation)","subtitle":"All calls in a configuration divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.01219,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.01352,16],["Claude Sonnet 5.5 (high) · Claude Code",0.01671,16],["Claude Sonnet 5.5 · Claude Code",0.01398,16],["Claude Opus 5.5 (low) · Claude Code",0.02115,16],["Claude Opus 5.5 (medium) · Claude Code",0.02947,16],["Claude Opus 5.5 (high) · Claude Code",0.03368,16],["Claude Opus 5.5 · Claude Code",0.02893,16],["GPT-6.1 Sol (low) · Codex CLI",0.01284,16],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16]]}}],"note":"Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.","sourceIds":["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-share","title":"Reasoning share of output tokens per call on hard tasks (calculation)","subtitle":"Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)","kind":"bar","unit":"percent","yLabel":"Reasoning share of output tokens (%)","series":[{"name":"Median call","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16]]}}],"note":"Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.","whisker":"minmax","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-cost-per-call","title":"List-price cost per call: reasoning, remaining output and input (calculation)","subtitle":"Mean per call on the hard tasks; the three parts add up to the call","kind":"stacked-bar","unit":"usd","yLabel":"USD per call (list price)","series":{"$k":["name","points"],"$r":[["Reasoning (output tokens)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.053696,24],["Claude Opus 5.5 (high) · Claude Code",0.017969,24],["Claude Haiku 4.5 · Claude Code",0.024492,24],["Claude Opus 5.5 · Claude Code",0.012528,24],["GPT-6.1 Sol (medium) · Codex CLI",0.002273,16],["GPT-6.1 Sol (high) · Codex CLI",0.004114,16],["Claude Sonnet 5.5 · Claude Code",0.006665,24]]}],["Remaining output (visible-answer estimate)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.018702,24],["Claude Opus 5.5 (high) · Claude Code",0.00799,24],["Claude Haiku 4.5 · Claude Code",0.001795,24],["Claude Opus 5.5 · Claude Code",0.008003,24],["GPT-6.1 Sol (medium) · Codex CLI",0.003025,16],["GPT-6.1 Sol (high) · Codex CLI",0.002905,16],["Claude Sonnet 5.5 · Claude Code",0.003672,24]]}],["Input (prompt, cache priced)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.02091,24],["Claude Opus 5.5 (high) · Claude Code",0.007407,24],["Claude Haiku 4.5 · Claude Code",0.00451,24],["Claude Opus 5.5 · Claude Code",0.00771,24],["GPT-6.1 Sol (medium) · Codex CLI",0.020339,16],["GPT-6.1 Sol (high) · Codex CLI",0.008117,16],["Claude Sonnet 5.5 · Claude Code",0.004012,24]]}]]},"note":"Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-by-effort","title":"Reasoning cost per strict pass by effort, with the total (calculation)","subtitle":"List price ÷ strict passes; every cell is 8 tasks × 2 repetitions","kind":"grouped-bar","unit":"usd","yLabel":"USD per strict pass (list price)","series":[{"name":"Reasoning cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.004317,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.005946,16],["Claude Sonnet 5.5 (high) · Claude Code",0.009369,16],["Claude Sonnet 5.5 · Claude Code",0.006299,16],["Claude Opus 5.5 (low) · Claude Code",0.005031,16],["Claude Opus 5.5 (medium) · Claude Code",0.01344,16],["Claude Opus 5.5 (high) · Claude Code",0.018034,16],["Claude Opus 5.5 · Claude Code",0.013104,16],["GPT-6.1 Sol (low) · Codex CLI",0.001223,16],["GPT-6.1 Sol (medium) · Codex CLI",0.002273,16],["GPT-6.1 Sol (high) · Codex CLI",0.004114,16]]}},{"name":"Total cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.012191,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.01352,16],["Claude Sonnet 5.5 (high) · Claude Code",0.016705,16],["Claude Sonnet 5.5 · Claude Code",0.013978,16],["Claude Opus 5.5 (low) · Claude Code",0.021152,16],["Claude Opus 5.5 (medium) · Claude Code",0.029475,16],["Claude Opus 5.5 (high) · Claude Code",0.033677,16],["Claude Opus 5.5 · Claude Code",0.028925,16],["GPT-6.1 Sol (low) · Codex CLI",0.012837,16],["GPT-6.1 Sol (medium) · Codex CLI",0.025637,16],["GPT-6.1 Sol (high) · Codex CLI",0.015137,16]]}}],"note":"Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). \"Default\" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-short-vs-hard","title":"Reasoning share on short tasks vs hard tasks (calculation)","subtitle":"Median call per configuration; five short tasks and eight hard tasks","kind":"grouped-bar","unit":"percent","yLabel":"Reasoning share of output tokens (median call, %)","series":[{"name":"Eight hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16]]}},{"name":"Five short tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",90.19,73.1,97.59,15],["Claude Sonnet 5.5 · Claude Code",0,0,72.75,15],["Claude Opus 5.5 · Claude Code",0,0,93.33,15],["Claude Opus 5.5 (high) · Claude Code",43.59,0,93.33,15],["Claude Fable 5.1 · Claude Code",0,0,74.01,15],["GPT-6.1 Sol (medium) · Codex CLI",41.05,0,71.43,15],["GPT-6.1 Sol (high) · Codex CLI",58.06,0,75.76,15]]}}],"note":"Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.","whisker":"minmax","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["llm-speed-anatomy",{"id":"speed-anatomy-first-text","title":"Time to first text: a 250-line answer, six models","subtitle":"Median of 4 calls per model; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds to first text","series":[{"name":"Time to first text","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",4,2.84,6.38,4],["Claude Sonnet 5.5 · Claude Code",1.96,0.88,4.09,4],["Claude Opus 5.5 · Claude Code",1.97,1.7,2.35,4],["Claude Fable 5.1 · Claude Code",4.43,2.27,4.64,4],["GPT-6.1 Sol (low) · Codex CLI",3.52,2.75,4.42,4],["GPT-6 Luna (low) · Codex CLI",3.3,3.19,3.47,4]]}}],"note":"Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-output-speed","title":"Output speed after the first text: visible tokens per second (calculation)","subtitle":"Median of 4 calls per model; whiskers = slowest and fastest call","kind":"dot-range","unit":"tokens","yLabel":"Visible tokens per second","series":[{"name":"Visible tokens per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",153.2,152.6,216.1,4],["Claude Sonnet 5.5 · Claude Code",231.7,230.3,233,4],["Claude Opus 5.5 · Claude Code",155.5,154.6,156.4,4],["Claude Fable 5.1 · Claude Code",122.6,120.9,131.4,4],["GPT-6.1 Sol (low) · Codex CLI",79.6,71.6,80.5,4],["GPT-6 Luna (low) · Codex CLI",129.1,55.5,259.1,4]]}}],"note":"Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-chars-per-second","title":"Output speed in characters per second after the first text (calculation)","subtitle":"Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call","kind":"dot-range","unit":"count","yLabel":"Characters per second","series":[{"name":"Characters per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",547,546,548,3],["Claude Sonnet 5.5 · Claude Code",517,513,519,4],["Claude Opus 5.5 · Claude Code",347,345,349,4],["Claude Fable 5.1 · Claude Code",273,270,293,4],["GPT-6.1 Sol (low) · Codex CLI",323,291,327,4],["GPT-6 Luna (low) · Codex CLI",524,225,1052,4]]}}],"note":"Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-prompt-size","title":"Time to first text as the prompt grows","subtitle":"Median of 3 calls per size; whiskers = fastest and slowest call","kind":"line","unit":"seconds","xLabel":"Prompt-size target (approximate Haiku tokens; calibration calculation)","yLabel":"Seconds to first text","series":{"$k":["name","points"],"$r":[["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.93,1.85,2.04,3],["16k",2.27,2.22,2.47,3],["64k",2.78,2.45,2.89,3]]}],["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.45,1.23,1.72,3],["16k",1.78,1.64,2.11,3],["64k",3.07,1.38,3.61,3]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.51,1.46,2.01,3],["16k",1.74,1.7,2.97,3],["64k",1.79,1.72,3.72,3]]}],["GPT-6.1 Sol (low) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["1k",3.36,3.36,4.75,3],["16k",4.02,3.3,4.28,3],["64k",3.93,3.42,4.38,3]]}]]},"note":"Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-total-by-size","title":"Total time per call by prompt size","subtitle":"Median of 3 calls per bar; whiskers = fastest and slowest call","kind":"grouped-bar","unit":"seconds","yLabel":"Seconds, whole call","series":{"$k":["name","points"],"$r":[["1k prompt",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",2.34,2.22,2.46,3],["Claude Sonnet 5.5 · Claude Code",1.78,1.57,2.12,3],["Claude Opus 5.5 · Claude Code",1.83,1.82,2.41,3],["GPT-6.1 Sol (low) · Codex CLI",3.43,3.43,4.92,3]]}],["16k prompt",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",2.79,2.58,2.84,3],["Claude Sonnet 5.5 · Claude Code",2.1,1.98,2.48,3],["Claude Opus 5.5 · Claude Code",2.36,2.11,3.4,3],["GPT-6.1 Sol (low) · Codex CLI",4.14,3.96,4.68,3]]}],["64k prompt",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",3.13,2.84,3.28,3],["Claude Sonnet 5.5 · Claude Code",3.44,1.74,4.38,3],["Claude Opus 5.5 · Claude Code",2.35,2.26,4.29,3],["GPT-6.1 Sol (low) · Codex CLI",3.96,3.47,4.44,3]]}]]},"note":"Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-lookup-correct","title":"Exact lookup answers at the 1k, 16k and 64k prompt-size targets","subtitle":"All sizes together per model; whiskers = 95% Wilson intervals","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Lookups answered exactly","series":[{"name":"Exact answer","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,0.7009,1,9],["Claude Sonnet 5.5 · Claude Code",1,0.7009,1,9],["Claude Opus 5.5 · Claude Code",0.5556,0.2667,0.8112,9],["GPT-6.1 Sol (low) · Codex CLI",1,0.7009,1,9]]}}],"note":"Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.","whisker":"ci95","sourceIds":["agent-speed-anatomy"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-pass-rate","title":"Pass rate on 4 harder tasks","subtitle":"Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Passed","whisker":"ci95","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.5,0.2538,0.7462,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}}],"note":"Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row.\n\nCounted calls are new calls.","sourceIds":["agent-harder-tasks"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-tool-attempts","title":"Calls that tried a tool although tools were off","subtitle":"Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse","kind":"dot-range","unit":"rate","yLabel":"Calls with a tool attempt","polarity":"none","whisker":"ci95","series":[{"name":"Tool attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0,0,0.1936,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.3125,0.1416,0.556,16],["Claude Haiku 4.5 · Claude Code",0.0833,0.0149,0.3539,12]]}}],"note":"Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator.\n\nIt is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.","sourceIds":["agent-harder-tasks"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-pass-by-task","title":"Strict pass rate by task","subtitle":"One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Strict pass rate","whisker":"ci95","series":{"$k":["name","points"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",0.75,0.3006,0.9544,4],["Sudoku, 22 givens",0.25,0.0456,0.6994,4],["6x6 Skyscrapers",1,0.5101,1,4],["Seeded shuffle output",0.75,0.3006,0.9544,4]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",1,0.4385,1,3],["Sudoku, 22 givens",0,0,0.5615,3],["6x6 Skyscrapers",0.3333,0.0615,0.7923,3],["Seeded shuffle output",0.3333,0.0615,0.7923,3]]}],["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",1,0.5101,1,4],["Sudoku, 22 givens",0,0,0.4899,4],["6x6 Skyscrapers",0,0,0.4899,4],["Seeded shuffle output",0.5,0.15,0.85,4]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",0,0,0.5615,3],["Sudoku, 22 givens",0,0,0.5615,3],["6x6 Skyscrapers",0,0,0.5615,3],["Seeded shuffle output",0,0,0.5615,3]]}]]},"note":"Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.","sourceIds":["agent-harder-tasks"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-total-latency","title":"Total time per call on harder tasks","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","whisker":"minmax","series":[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",120.24,46.24,273.46,13,false],["Claude Opus 5.5 · Claude Code",80.34,3.82,279.5,9,false],["Claude Sonnet 5.5 · Claude Code",70.43,4.32,210.08,12,false],["Claude Haiku 4.5 · Claude Code",108.98,25.73,223.95,10,false]]}}],"note":"Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session.\n\nArena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.","sourceIds":["agent-harder-tasks"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-output-tokens","title":"Output tokens per call on harder tasks","subtitle":"Median per configuration; reasoning tokens as the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","polarity":"none","whisker":"minmax","series":[{"name":"Output tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",4994,2099,13413,13],["Claude Opus 5.5 · Claude Code",8420,279,40044,9],["Claude Sonnet 5.5 · Claude Code",9287,407,27921,12],["Claude Haiku 4.5 · Claude Code",12508,2965,26532,10]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",4971,2070,13372,13],["Claude Opus 5.5 · Claude Code",8352,21,9897,9],["Claude Sonnet 5.5 · Claude Code",6557,63,27902,12],["Claude Haiku 4.5 · Claude Code",12483,2924,26510,10]]}}],"note":"Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured.\n\nClaude Code used an output-token cap setting of 16,000. Some reported totals exceeded it.\n\nCodex CLI had no cap. More tokens is not better or worse by itself.","sourceIds":["agent-harder-tasks"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-cost-per-pass","title":"List-price cost per strict pass on harder tasks (calculation)","subtitle":"All calls in a configuration, failures and format misses included, divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n","highlight"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.08293,16,true],["Claude Sonnet 5.5 · Claude Code",0.23843,16,false],["Claude Opus 5.5 · Claude Code",0.59333,12,false]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. These cells show a lower bound.\n\nAssume each unpriced call cost its cell’s median priced call. This sensitivity calculation gives GPT-6.1 Sol (medium) $0.100, Opus 5.5 $0.633 and Sonnet 5.5 $0.303. Opus 5.5 figures are provisional: its cache-read price is under re-check.\n\nHighlights mark the observed frontier of these lower-bound costs. Unknown timeout costs can change it; this is not a cost ranking. Claude Haiku 4.5 · Claude Code had no strict pass, so it has no cost per pass.","sourceIds":["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"]}]]}