{"$k":["slug","chart"],"$r":[["model-head-to-head",{"id":"h2h-pass-rate","title":"Pass rate on five validated tasks","subtitle":"Every call counts; failures and timeouts are non-passes","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Pass rate","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Fable 5.1 · Claude Code",1,0.7961,1,15],["Claude Sonnet 5.5 · Claude Code",0.8,0.5481,0.9295,15],["Claude Opus 5.5 (high) · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 (low) · Claude Code",1,0.7961,1,15],["Claude Haiku 4.5 · Claude Code",1,0.7961,1,15],["GPT-6.1 Sol (high) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (low) · Codex CLI",1,0.7225,1,10]]}}],"note":"Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-total-latency","title":"Total time per call","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.94,1.41,9.83,15,true],["Claude Sonnet 5.5 · Claude Code",2.31,2.17,7.73,15,true],["Claude Opus 5.5 (high) · Claude Code",2.71,2.45,11.78,15,true],["Claude Opus 5.5 · Claude Code",2.75,2.47,8.91,15,true],["Claude Opus 5.5 (low) · Claude Code",2.83,2.35,6.62,15,true],["Claude Haiku 4.5 · Claude Code",4.43,3.16,23.57,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.6,4.05,19.52,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.65,4.1,25.46,15,false],["GPT-6.1 Sol (low) · Codex CLI",6.26,4.65,10.47,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-first-useful-latency","title":"Time to first useful output","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Time to first useful output","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.2,0.95,7.9,15,true],["Claude Sonnet 5.5 · Claude Code",1.56,0.99,6.39,15,true],["Claude Opus 5.5 (high) · Claude Code",2.04,1.4,9.94,15,true],["Claude Opus 5.5 · Claude Code",1.92,1.56,7.23,15,true],["Claude Opus 5.5 (low) · Claude Code",2.39,1.45,4.9,15,true],["Claude Haiku 4.5 · Claude Code",3.63,2.78,22.27,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.32,3.64,16.37,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.05,3.36,17.82,15,false],["GPT-6.1 Sol (low) · Codex CLI",5.14,4.02,8.5,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-input-tokens","title":"Input tokens per call: what the CLI sends","subtitle":"Mean per call, split into prompt-cache reads and other input","kind":"stacked-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Cache read","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",2760,15],["Claude Sonnet 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (high) · Claude Code",1463,15],["Claude Opus 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (low) · Claude Code",1463,15],["Claude Haiku 4.5 · Claude Code",0,15],["GPT-6.1 Sol (high) · Codex CLI",6716,15],["GPT-6.1 Sol (medium) · Codex CLI",5180,15],["GPT-6.1 Sol (low) · Codex CLI",8064,10]]}},{"name":"Other input","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",473,15],["Claude Sonnet 5.5 · Claude Code",685,15],["Claude Opus 5.5 (high) · Claude Code",619,15],["Claude Opus 5.5 · Claude Code",680,15],["Claude Opus 5.5 (low) · Claude Code",618,15],["Claude Haiku 4.5 · Claude Code",3790,15],["GPT-6.1 Sol (high) · Codex CLI",5406,15],["GPT-6.1 Sol (medium) · Codex CLI",6943,15],["GPT-6.1 Sol (low) · Codex CLI",4059,10]]}}],"note":"The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-output-tokens","title":"Output tokens per call","subtitle":"Median per configuration; reasoning tokens where the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",64,15],["Claude Sonnet 5.5 · Claude Code",107,15],["Claude Opus 5.5 (high) · Claude Code",78,15],["Claude Opus 5.5 · Claude Code",64,15],["Claude Opus 5.5 (low) · Claude Code",64,15],["Claude Haiku 4.5 · Claude Code",367,15],["GPT-6.1 Sol (high) · Codex CLI",42,15],["GPT-6.1 Sol (medium) · Codex CLI",42,15],["GPT-6.1 Sol (low) · Codex CLI",42,10]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0,15],["Claude Sonnet 5.5 · Claude Code",0,15],["Claude Opus 5.5 (high) · Claude Code",34,15],["Claude Opus 5.5 · Claude Code",0,15],["Claude Opus 5.5 (low) · Claude Code",0,15],["Claude Haiku 4.5 · Claude Code",297,15],["GPT-6.1 Sol (high) · Codex CLI",21,15],["GPT-6.1 Sol (medium) · Codex CLI",20,15],["GPT-6.1 Sol (low) · Codex CLI",20,10]]}}],"note":"Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-list-price-per-call","title":"List-price cost per call (calculation)","subtitle":"Reported tokens × list price; the calls ran on subscriptions","kind":"dot-range","unit":"usd","yLabel":"USD per call","series":[{"name":"Cost per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",0.00987,0.0049,0.05843,15,true],["Claude Sonnet 5.5 · Claude Code",0.0036,0.00342,0.01021,15,true],["Claude Opus 5.5 (high) · Claude Code",0.00694,0.00592,0.02708,15,true],["Claude Opus 5.5 · Claude Code",0.00688,0.00592,0.02226,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00688,0.00582,0.01793,15,true],["Claude Haiku 4.5 · Claude Code",0.00566,0.00513,0.01804,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01047,0.0066,0.02812,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01018,0.0054,0.02686,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00769,0.00742,0.02649,10,false]]}}],"note":"Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["model-head-to-head",{"id":"h2h-cost-per-pass","title":"List-price cost per passing answer (calculation)","subtitle":"All calls in a configuration, failures included, divided by its passes","kind":"bar","unit":"usd","yLabel":"USD per passing answer","series":[{"name":"Cost per pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.00624,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00829,15,true],["Claude Haiku 4.5 · Claude Code",0.00836,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00998,10,false],["Claude Opus 5.5 · Claude Code",0.01009,15,true],["Claude Opus 5.5 (high) · Claude Code",0.01049,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01322,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01564,15,false],["Claude Fable 5.1 · Claude Code",0.02054,15,true]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["effort-ladder",{"id":"effort-ladder-pass-rate","title":"Strict pass rate by effort on eight hard tasks","subtitle":"Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 · Claude Code",1,0.8064,1,16],["GPT-6.1 Sol (low) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16]]}}],"note":"Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.","whisker":"ci95","sourceIds":["agent-effort-ladder"]}],["effort-ladder",{"id":"effort-ladder-total-latency","title":"Total time per call by effort on hard tasks","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call by effort on hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",5.82,2.78,19.96,16],["Claude Sonnet 5.5 (medium) · Claude Code",7.63,2.71,24.01,16],["Claude Sonnet 5.5 (high) · Claude Code",8.81,2.93,35.81,16],["Claude Sonnet 5.5 · Claude Code",7.97,2.26,21.61,16],["Claude Opus 5.5 (low) · Claude Code",7.5,3.34,15.82,16],["Claude Opus 5.5 (medium) · Claude Code",9.72,4.78,31.36,16],["Claude Opus 5.5 (high) · Claude Code",10.11,3.63,63,16],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,16],["GPT-6.1 Sol (low) · Codex CLI",13.62,7.94,44.29,16],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16]]}}],"note":"Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.","whisker":"minmax","sourceIds":["agent-effort-ladder"]}],["effort-ladder",{"id":"effort-ladder-output-tokens","title":"Output tokens per call by effort on hard tasks","subtitle":"Median per configuration; reasoning tokens as the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",667,16],["Claude Sonnet 5.5 (medium) · Claude Code",770,16],["Claude Sonnet 5.5 (high) · Claude Code",1192,16],["Claude Sonnet 5.5 · Claude Code",1054,16],["Claude Opus 5.5 (low) · Claude Code",594,16],["Claude Opus 5.5 (medium) · Claude Code",853,16],["Claude Opus 5.5 (high) · Claude Code",1052,16],["Claude Opus 5.5 · Claude Code",945,16],["GPT-6.1 Sol (low) · Codex CLI",284,16],["GPT-6.1 Sol (medium) · Codex CLI",335,16],["GPT-6.1 Sol (high) · Codex CLI",436,16]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",273,16],["Claude Sonnet 5.5 (medium) · Claude Code",422,16],["Claude Sonnet 5.5 (high) · Claude Code",745,16],["Claude Sonnet 5.5 · Claude Code",668,16],["Claude Opus 5.5 (low) · Claude Code",87,16],["Claude Opus 5.5 (medium) · Claude Code",518,16],["Claude Opus 5.5 (high) · Claude Code",614,16],["Claude Opus 5.5 · Claude Code",538,16],["GPT-6.1 Sol (low) · Codex CLI",63,16],["GPT-6.1 Sol (medium) · Codex CLI",150,16],["GPT-6.1 Sol (high) · Codex CLI",225,16]]}}],"note":"Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean \"not reported\". Their content is never captured. More tokens is not better or worse by itself.","sourceIds":["agent-effort-ladder"]}],["effort-ladder",{"id":"effort-ladder-cost-per-pass","title":"List-price cost per strict pass by effort (calculation)","subtitle":"All calls in a configuration divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.01219,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.01352,16],["Claude Sonnet 5.5 (high) · Claude Code",0.01671,16],["Claude Sonnet 5.5 · Claude Code",0.01398,16],["Claude Opus 5.5 (low) · Claude Code",0.02115,16],["Claude Opus 5.5 (medium) · Claude Code",0.02947,16],["Claude Opus 5.5 (high) · Claude Code",0.03368,16],["Claude Opus 5.5 · Claude Code",0.02893,16],["GPT-6.1 Sol (low) · Codex CLI",0.01284,16],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16]]}}],"note":"Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.","sourceIds":["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-by-effort","title":"Reasoning cost per strict pass by effort, with the total (calculation)","subtitle":"List price ÷ strict passes; every cell is 8 tasks × 2 repetitions","kind":"grouped-bar","unit":"usd","yLabel":"USD per strict pass (list price)","series":[{"name":"Reasoning cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.004317,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.005946,16],["Claude Sonnet 5.5 (high) · Claude Code",0.009369,16],["Claude Sonnet 5.5 · Claude Code",0.006299,16],["Claude Opus 5.5 (low) · Claude Code",0.005031,16],["Claude Opus 5.5 (medium) · Claude Code",0.01344,16],["Claude Opus 5.5 (high) · Claude Code",0.018034,16],["Claude Opus 5.5 · Claude Code",0.013104,16],["GPT-6.1 Sol (low) · Codex CLI",0.001223,16],["GPT-6.1 Sol (medium) · Codex CLI",0.002273,16],["GPT-6.1 Sol (high) · Codex CLI",0.004114,16]]}},{"name":"Total cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.012191,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.01352,16],["Claude Sonnet 5.5 (high) · Claude Code",0.016705,16],["Claude Sonnet 5.5 · Claude Code",0.013978,16],["Claude Opus 5.5 (low) · Claude Code",0.021152,16],["Claude Opus 5.5 (medium) · Claude Code",0.029475,16],["Claude Opus 5.5 (high) · Claude Code",0.033677,16],["Claude Opus 5.5 · Claude Code",0.028925,16],["GPT-6.1 Sol (low) · Codex CLI",0.012837,16],["GPT-6.1 Sol (medium) · Codex CLI",0.025637,16],["GPT-6.1 Sol (high) · Codex CLI",0.015137,16]]}}],"note":"Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). \"Default\" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}]]}