{"$k":["slug","chart"],"$r":[["model-head-to-head",{"id":"h2h-pass-rate","title":"Pass rate on five validated tasks","subtitle":"Every call counts; failures and timeouts are non-passes","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Pass rate","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Fable 5.1 · Claude Code",1,0.7961,1,15],["Claude Sonnet 5.5 · Claude Code",0.8,0.5481,0.9295,15],["Claude Opus 5.5 (high) · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 (low) · Claude Code",1,0.7961,1,15],["Claude Haiku 4.5 · Claude Code",1,0.7961,1,15],["GPT-6.1 Sol (high) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (low) · Codex CLI",1,0.7225,1,10]]}}],"note":"Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-total-latency","title":"Total time per call","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.94,1.41,9.83,15,true],["Claude Sonnet 5.5 · Claude Code",2.31,2.17,7.73,15,true],["Claude Opus 5.5 (high) · Claude Code",2.71,2.45,11.78,15,true],["Claude Opus 5.5 · Claude Code",2.75,2.47,8.91,15,true],["Claude Opus 5.5 (low) · Claude Code",2.83,2.35,6.62,15,true],["Claude Haiku 4.5 · Claude Code",4.43,3.16,23.57,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.6,4.05,19.52,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.65,4.1,25.46,15,false],["GPT-6.1 Sol (low) · Codex CLI",6.26,4.65,10.47,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-first-useful-latency","title":"Time to first useful output","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Time to first useful output","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.2,0.95,7.9,15,true],["Claude Sonnet 5.5 · Claude Code",1.56,0.99,6.39,15,true],["Claude Opus 5.5 (high) · Claude Code",2.04,1.4,9.94,15,true],["Claude Opus 5.5 · Claude Code",1.92,1.56,7.23,15,true],["Claude Opus 5.5 (low) · Claude Code",2.39,1.45,4.9,15,true],["Claude Haiku 4.5 · Claude Code",3.63,2.78,22.27,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.32,3.64,16.37,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.05,3.36,17.82,15,false],["GPT-6.1 Sol (low) · Codex CLI",5.14,4.02,8.5,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-input-tokens","title":"Input tokens per call: what the CLI sends","subtitle":"Mean per call, split into prompt-cache reads and other input","kind":"stacked-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Cache read","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",2760,15],["Claude Sonnet 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (high) · Claude Code",1463,15],["Claude Opus 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (low) · Claude Code",1463,15],["Claude Haiku 4.5 · Claude Code",0,15],["GPT-6.1 Sol (high) · Codex CLI",6716,15],["GPT-6.1 Sol (medium) · Codex CLI",5180,15],["GPT-6.1 Sol (low) · Codex CLI",8064,10]]}},{"name":"Other input","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",473,15],["Claude Sonnet 5.5 · Claude Code",685,15],["Claude Opus 5.5 (high) · Claude Code",619,15],["Claude Opus 5.5 · Claude Code",680,15],["Claude Opus 5.5 (low) · Claude Code",618,15],["Claude Haiku 4.5 · Claude Code",3790,15],["GPT-6.1 Sol (high) · Codex CLI",5406,15],["GPT-6.1 Sol (medium) · Codex CLI",6943,15],["GPT-6.1 Sol (low) · Codex CLI",4059,10]]}}],"note":"The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-output-tokens","title":"Output tokens per call","subtitle":"Median per configuration; reasoning tokens where the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",64,15],["Claude Sonnet 5.5 · Claude Code",107,15],["Claude Opus 5.5 (high) · Claude Code",78,15],["Claude Opus 5.5 · Claude Code",64,15],["Claude Opus 5.5 (low) · Claude Code",64,15],["Claude Haiku 4.5 · Claude Code",367,15],["GPT-6.1 Sol (high) · Codex CLI",42,15],["GPT-6.1 Sol (medium) · Codex CLI",42,15],["GPT-6.1 Sol (low) · Codex CLI",42,10]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0,15],["Claude Sonnet 5.5 · Claude Code",0,15],["Claude Opus 5.5 (high) · Claude Code",34,15],["Claude Opus 5.5 · Claude Code",0,15],["Claude Opus 5.5 (low) · Claude Code",0,15],["Claude Haiku 4.5 · Claude Code",297,15],["GPT-6.1 Sol (high) · Codex CLI",21,15],["GPT-6.1 Sol (medium) · Codex CLI",20,15],["GPT-6.1 Sol (low) · Codex CLI",20,10]]}}],"note":"Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-list-price-per-call","title":"List-price cost per call (calculation)","subtitle":"Reported tokens × list price; the calls ran on subscriptions","kind":"dot-range","unit":"usd","yLabel":"USD per call","series":[{"name":"Cost per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",0.00987,0.0049,0.05843,15,true],["Claude Sonnet 5.5 · Claude Code",0.0036,0.00342,0.01021,15,true],["Claude Opus 5.5 (high) · Claude Code",0.00694,0.00592,0.02708,15,true],["Claude Opus 5.5 · Claude Code",0.00688,0.00592,0.02226,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00688,0.00582,0.01793,15,true],["Claude Haiku 4.5 · Claude Code",0.00566,0.00513,0.01804,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01047,0.0066,0.02812,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01018,0.0054,0.02686,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00769,0.00742,0.02649,10,false]]}}],"note":"Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["model-head-to-head",{"id":"h2h-cost-per-pass","title":"List-price cost per passing answer (calculation)","subtitle":"All calls in a configuration, failures included, divided by its passes","kind":"bar","unit":"usd","yLabel":"USD per passing answer","series":[{"name":"Cost per pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.00624,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00829,15,true],["Claude Haiku 4.5 · Claude Code",0.00836,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00998,10,false],["Claude Opus 5.5 · Claude Code",0.01009,15,true],["Claude Opus 5.5 (high) · Claude Code",0.01049,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01322,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01564,15,false],["Claude Fable 5.1 · Claude Code",0.02054,15,true]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["effort-ladder",{"id":"effort-ladder-pass-rate","title":"Strict pass rate by effort on eight hard tasks","subtitle":"Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 · Claude Code",1,0.8064,1,16],["GPT-6.1 Sol (low) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16]]}}],"note":"Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.","whisker":"ci95","sourceIds":["agent-effort-ladder"]}],["effort-ladder",{"id":"effort-ladder-total-latency","title":"Total time per call by effort on hard tasks","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call by effort on hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",5.82,2.78,19.96,16],["Claude Sonnet 5.5 (medium) · Claude Code",7.63,2.71,24.01,16],["Claude Sonnet 5.5 (high) · Claude Code",8.81,2.93,35.81,16],["Claude Sonnet 5.5 · Claude Code",7.97,2.26,21.61,16],["Claude Opus 5.5 (low) · Claude Code",7.5,3.34,15.82,16],["Claude Opus 5.5 (medium) · Claude Code",9.72,4.78,31.36,16],["Claude Opus 5.5 (high) · Claude Code",10.11,3.63,63,16],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,16],["GPT-6.1 Sol (low) · Codex CLI",13.62,7.94,44.29,16],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16]]}}],"note":"Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.","whisker":"minmax","sourceIds":["agent-effort-ladder"]}],["effort-ladder",{"id":"effort-ladder-output-tokens","title":"Output tokens per call by effort on hard tasks","subtitle":"Median per configuration; reasoning tokens as the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",667,16],["Claude Sonnet 5.5 (medium) · Claude Code",770,16],["Claude Sonnet 5.5 (high) · Claude Code",1192,16],["Claude Sonnet 5.5 · Claude Code",1054,16],["Claude Opus 5.5 (low) · Claude Code",594,16],["Claude Opus 5.5 (medium) · Claude Code",853,16],["Claude Opus 5.5 (high) · Claude Code",1052,16],["Claude Opus 5.5 · Claude Code",945,16],["GPT-6.1 Sol (low) · Codex CLI",284,16],["GPT-6.1 Sol (medium) · Codex CLI",335,16],["GPT-6.1 Sol (high) · Codex CLI",436,16]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",273,16],["Claude Sonnet 5.5 (medium) · Claude Code",422,16],["Claude Sonnet 5.5 (high) · Claude Code",745,16],["Claude Sonnet 5.5 · Claude Code",668,16],["Claude Opus 5.5 (low) · Claude Code",87,16],["Claude Opus 5.5 (medium) · Claude Code",518,16],["Claude Opus 5.5 (high) · Claude Code",614,16],["Claude Opus 5.5 · Claude Code",538,16],["GPT-6.1 Sol (low) · Codex CLI",63,16],["GPT-6.1 Sol (medium) · Codex CLI",150,16],["GPT-6.1 Sol (high) · Codex CLI",225,16]]}}],"note":"Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean \"not reported\". Their content is never captured. More tokens is not better or worse by itself.","sourceIds":["agent-effort-ladder"]}],["effort-ladder",{"id":"effort-ladder-cost-per-pass","title":"List-price cost per strict pass by effort (calculation)","subtitle":"All calls in a configuration divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.01219,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.01352,16],["Claude Sonnet 5.5 (high) · Claude Code",0.01671,16],["Claude Sonnet 5.5 · Claude Code",0.01398,16],["Claude Opus 5.5 (low) · Claude Code",0.02115,16],["Claude Opus 5.5 (medium) · Claude Code",0.02947,16],["Claude Opus 5.5 (high) · Claude Code",0.03368,16],["Claude Opus 5.5 · Claude Code",0.02893,16],["GPT-6.1 Sol (low) · Codex CLI",0.01284,16],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16]]}}],"note":"Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.","sourceIds":["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"]}],["cli-model-latency-tokens",{"id":"cli-vs-api-exact-reply-latency","title":"CLI vs API: time for a one-line answer","subtitle":"Matched cohort, fixed exact reply, 5 runs per configuration","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.97,0.65,1.5,5],["OpenAI API · GPT-6.1 Sol · low",1.02,0.96,1.87,5],["OpenAI API · GPT-6.1 Sol · high",1.52,1.35,2.23,5],["Codex CLI · GPT-6 Luna · none",3.19,2.88,3.83,5],["Codex CLI · GPT-6.1 Sol · low",4.18,3.86,4.53,5],["Codex CLI · GPT-6.1 Sol · high",4.19,3.81,4.69,5]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.82,0.51,1.37,5],["OpenAI API · GPT-6.1 Sol · low",0.87,0.84,1.74,5],["OpenAI API · GPT-6.1 Sol · high",1.34,1.26,2.12,5],["Codex CLI · GPT-6 Luna · none",2.79,2.46,3.42,5],["Codex CLI · GPT-6.1 Sol · low",3.75,3.44,4.1,5],["Codex CLI · GPT-6.1 Sol · high",3.79,3.37,4.3,5]]}}],"note":"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.","sourceIds":["agent-provider-explorer"]}],["cli-model-latency-tokens",{"id":"cli-vs-api-small-coding-latency","title":"CLI vs API: time for a small coding task","subtitle":"Matched cohort, small coding task, 3 runs per configuration","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",4.01,3.83,4.35,3],["OpenAI API · GPT-6.1 Sol · low",6,5.44,6.2,3],["Codex CLI · GPT-6 Luna · none",9.23,8.99,11.68,3],["OpenAI API · GPT-6.1 Sol · high",9.56,9.44,10.94,3],["Codex CLI · GPT-6.1 Sol · low",14.15,13.02,14.41,3],["Codex CLI · GPT-6.1 Sol · high",17.85,17.68,22.42,3]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.67,0.62,0.81,3],["OpenAI API · GPT-6.1 Sol · low",1.05,0.97,1.4,3],["Codex CLI · GPT-6 Luna · none",8.68,8.27,11.01,3],["OpenAI API · GPT-6.1 Sol · high",5.31,4.99,6.42,3],["Codex CLI · GPT-6.1 Sol · low",13.6,12.52,13.83,3],["Codex CLI · GPT-6.1 Sol · high",17.27,17.13,21.86,3]]}}],"note":"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.","sourceIds":["agent-provider-explorer"]}],["cli-model-latency-tokens",{"id":"cli-vs-api-prompt-overhead","title":"Hidden prompt: input tokens for the same one-line request","subtitle":"Reported input tokens, matched cohort","kind":"bar","unit":"tokens","yLabel":"Input tokens per call","series":[{"name":"Input tokens","points":{"$k":["label","value","n"],"$r":[["OpenAI API · GPT-6 Luna · none",17,5],["OpenAI API · GPT-6.1 Sol · low",17,5],["OpenAI API · GPT-6.1 Sol · high",17,5],["Codex CLI · GPT-6 Luna · none",18859,5],["Codex CLI · GPT-6.1 Sol · low",19551,5],["Codex CLI · GPT-6.1 Sol · high",19555,5]]}}],"note":"The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.","sourceIds":["agent-provider-explorer"]}],["thinking-token-bill",{"id":"thinking-bill-by-effort","title":"Reasoning cost per strict pass by effort, with the total (calculation)","subtitle":"List price ÷ strict passes; every cell is 8 tasks × 2 repetitions","kind":"grouped-bar","unit":"usd","yLabel":"USD per strict pass (list price)","series":[{"name":"Reasoning cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.004317,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.005946,16],["Claude Sonnet 5.5 (high) · Claude Code",0.009369,16],["Claude Sonnet 5.5 · Claude Code",0.006299,16],["Claude Opus 5.5 (low) · Claude Code",0.005031,16],["Claude Opus 5.5 (medium) · Claude Code",0.01344,16],["Claude Opus 5.5 (high) · Claude Code",0.018034,16],["Claude Opus 5.5 · Claude Code",0.013104,16],["GPT-6.1 Sol (low) · Codex CLI",0.001223,16],["GPT-6.1 Sol (medium) · Codex CLI",0.002273,16],["GPT-6.1 Sol (high) · Codex CLI",0.004114,16]]}},{"name":"Total cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.012191,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.01352,16],["Claude Sonnet 5.5 (high) · Claude Code",0.016705,16],["Claude Sonnet 5.5 · Claude Code",0.013978,16],["Claude Opus 5.5 (low) · Claude Code",0.021152,16],["Claude Opus 5.5 (medium) · Claude Code",0.029475,16],["Claude Opus 5.5 (high) · Claude Code",0.033677,16],["Claude Opus 5.5 · Claude Code",0.028925,16],["GPT-6.1 Sol (low) · Codex CLI",0.012837,16],["GPT-6.1 Sol (medium) · Codex CLI",0.025637,16],["GPT-6.1 Sol (high) · Codex CLI",0.015137,16]]}}],"note":"Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). \"Default\" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}]]}