{"$k":["slug","chart"],"$r":[["model-head-to-head",{"id":"h2h-pass-rate","title":"Pass rate on five validated tasks","subtitle":"Every call counts; failures and timeouts are non-passes","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Pass rate","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Fable 5.1 · Claude Code",1,0.7961,1,15],["Claude Sonnet 5.5 · Claude Code",0.8,0.5481,0.9295,15],["Claude Opus 5.5 (high) · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 (low) · Claude Code",1,0.7961,1,15],["Claude Haiku 4.5 · Claude Code",1,0.7961,1,15],["GPT-6.1 Sol (high) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (low) · Codex CLI",1,0.7225,1,10]]}}],"note":"Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-total-latency","title":"Total time per call","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.94,1.41,9.83,15,true],["Claude Sonnet 5.5 · Claude Code",2.31,2.17,7.73,15,true],["Claude Opus 5.5 (high) · Claude Code",2.71,2.45,11.78,15,true],["Claude Opus 5.5 · Claude Code",2.75,2.47,8.91,15,true],["Claude Opus 5.5 (low) · Claude Code",2.83,2.35,6.62,15,true],["Claude Haiku 4.5 · Claude Code",4.43,3.16,23.57,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.6,4.05,19.52,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.65,4.1,25.46,15,false],["GPT-6.1 Sol (low) · Codex CLI",6.26,4.65,10.47,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-first-useful-latency","title":"Time to first useful output","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Time to first useful output","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.2,0.95,7.9,15,true],["Claude Sonnet 5.5 · Claude Code",1.56,0.99,6.39,15,true],["Claude Opus 5.5 (high) · Claude Code",2.04,1.4,9.94,15,true],["Claude Opus 5.5 · Claude Code",1.92,1.56,7.23,15,true],["Claude Opus 5.5 (low) · Claude Code",2.39,1.45,4.9,15,true],["Claude Haiku 4.5 · Claude Code",3.63,2.78,22.27,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.32,3.64,16.37,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.05,3.36,17.82,15,false],["GPT-6.1 Sol (low) · Codex CLI",5.14,4.02,8.5,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-input-tokens","title":"Input tokens per call: what the CLI sends","subtitle":"Mean per call, split into prompt-cache reads and other input","kind":"stacked-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Cache read","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",2760,15],["Claude Sonnet 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (high) · Claude Code",1463,15],["Claude Opus 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (low) · Claude Code",1463,15],["Claude Haiku 4.5 · Claude Code",0,15],["GPT-6.1 Sol (high) · Codex CLI",6716,15],["GPT-6.1 Sol (medium) · Codex CLI",5180,15],["GPT-6.1 Sol (low) · Codex CLI",8064,10]]}},{"name":"Other input","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",473,15],["Claude Sonnet 5.5 · Claude Code",685,15],["Claude Opus 5.5 (high) · Claude Code",619,15],["Claude Opus 5.5 · Claude Code",680,15],["Claude Opus 5.5 (low) · Claude Code",618,15],["Claude Haiku 4.5 · Claude Code",3790,15],["GPT-6.1 Sol (high) · Codex CLI",5406,15],["GPT-6.1 Sol (medium) · Codex CLI",6943,15],["GPT-6.1 Sol (low) · Codex CLI",4059,10]]}}],"note":"The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-output-tokens","title":"Output tokens per call","subtitle":"Median per configuration; reasoning tokens where the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",64,15],["Claude Sonnet 5.5 · Claude Code",107,15],["Claude Opus 5.5 (high) · Claude Code",78,15],["Claude Opus 5.5 · Claude Code",64,15],["Claude Opus 5.5 (low) · Claude Code",64,15],["Claude Haiku 4.5 · Claude Code",367,15],["GPT-6.1 Sol (high) · Codex CLI",42,15],["GPT-6.1 Sol (medium) · Codex CLI",42,15],["GPT-6.1 Sol (low) · Codex CLI",42,10]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0,15],["Claude Sonnet 5.5 · Claude Code",0,15],["Claude Opus 5.5 (high) · Claude Code",34,15],["Claude Opus 5.5 · Claude Code",0,15],["Claude Opus 5.5 (low) · Claude Code",0,15],["Claude Haiku 4.5 · Claude Code",297,15],["GPT-6.1 Sol (high) · Codex CLI",21,15],["GPT-6.1 Sol (medium) · Codex CLI",20,15],["GPT-6.1 Sol (low) · Codex CLI",20,10]]}}],"note":"Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-list-price-per-call","title":"List-price cost per call (calculation)","subtitle":"Reported tokens × list price; the calls ran on subscriptions","kind":"dot-range","unit":"usd","yLabel":"USD per call","series":[{"name":"Cost per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",0.00987,0.0049,0.05843,15,true],["Claude Sonnet 5.5 · Claude Code",0.0036,0.00342,0.01021,15,true],["Claude Opus 5.5 (high) · Claude Code",0.00694,0.00592,0.02708,15,true],["Claude Opus 5.5 · Claude Code",0.00688,0.00592,0.02226,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00688,0.00582,0.01793,15,true],["Claude Haiku 4.5 · Claude Code",0.00566,0.00513,0.01804,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01047,0.0066,0.02812,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01018,0.0054,0.02686,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00769,0.00742,0.02649,10,false]]}}],"note":"Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["model-head-to-head",{"id":"h2h-cost-per-pass","title":"List-price cost per passing answer (calculation)","subtitle":"All calls in a configuration, failures included, divided by its passes","kind":"bar","unit":"usd","yLabel":"USD per passing answer","series":[{"name":"Cost per pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.00624,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00829,15,true],["Claude Haiku 4.5 · Claude Code",0.00836,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00998,10,false],["Claude Opus 5.5 · Claude Code",0.01009,15,true],["Claude Opus 5.5 (high) · Claude Code",0.01049,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01322,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01564,15,false],["Claude Fable 5.1 · Claude Code",0.02054,15,true]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["hard-model-head-to-head",{"id":"hard-h2h-pass-rate","title":"Pass rate on eight hard tasks","subtitle":"Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.4583,0.2789,0.6493,24]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.6667,0.4671,0.8203,24]]}}],"note":"Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-total-latency","title":"Total time per call on hard tasks (separate batches)","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call on hard tasks (separate batches)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",7.75,2.26,34.79,24,true],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,24,true],["Claude Opus 5.5 (high) · Claude Code",11.03,3.63,63,24,true],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16,true],["Claude Fable 5.1 · Claude Code",16.13,4.46,90,24,true],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16,true],["Claude Haiku 4.5 · Claude Code",39.01,15.27,75.13,24,false]]}}],"note":"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-first-useful-latency","title":"Time to first useful output on hard tasks","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Time to first useful output on hard tasks","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",5.95,0.86,30.57,24,true],["Claude Opus 5.5 · Claude Code",6.78,2.39,21.77,24,true],["Claude Opus 5.5 (high) · Claude Code",7.13,2.15,56.23,24,true],["GPT-6.1 Sol (medium) · Codex CLI",10.23,6.09,40.41,16,true],["Claude Fable 5.1 · Claude Code",11.63,2,85.33,24,true],["GPT-6.1 Sol (high) · Codex CLI",12.69,8.93,75.91,16,true],["Claude Haiku 4.5 · Claude Code",35.54,12.88,70.31,24,false]]}}],"note":"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-output-tokens","title":"Output tokens per call on hard tasks","subtitle":"Median per configuration; reasoning tokens as the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1050,24],["Claude Opus 5.5 · Claude Code",945,24],["Claude Opus 5.5 (high) · Claude Code",1052,24],["GPT-6.1 Sol (medium) · Codex CLI",335,16],["Claude Fable 5.1 · Claude Code",1366,24],["GPT-6.1 Sol (high) · Codex CLI",436,16],["Claude Haiku 4.5 · Claude Code",5064,24]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",585,24],["Claude Opus 5.5 · Claude Code",529,24],["Claude Opus 5.5 (high) · Claude Code",614,24],["GPT-6.1 Sol (medium) · Codex CLI",150,16],["Claude Fable 5.1 · Claude Code",889,24],["GPT-6.1 Sol (high) · Codex CLI",225,16],["Claude Haiku 4.5 · Claude Code",4556,24]]}}],"note":"Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-cost-per-pass","title":"List-price cost per strict pass on hard tasks (calculation)","subtitle":"All calls in a configuration, failures and format misses included, divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.01435,24,true],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16,false],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16,false],["Claude Opus 5.5 · Claude Code",0.02824,24,false],["Claude Opus 5.5 (high) · Claude Code",0.03337,24,false],["Claude Haiku 4.5 · Claude Code",0.0672,24,false],["Claude Fable 5.1 · Claude Code",0.09331,24,false]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.","sourceIds":["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]}],["cost-thought-experiments",{"id":"repriced-cost-per-resolved","title":"Thought experiment: the same tokens at other list prices","subtitle":"Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price","kind":"bar","unit":"usd","yLabel":"USD per resolved instance","series":[{"name":"Repriced cost per resolved instance","points":{"$k":["label","value","highlight"],"$r":[["Claude Fable 5.1",12.852,false],["Claude Opus 5",8.723,false],["Claude Opus 5.5",5.753,false],["Claude Sonnet 5.5",3.489,true],["GPT-6.1 Sol",2.097,false],["Claude Haiku 4.5",1.745,false],["Gemini 3.x Flash",1.016,false],["Jev 1.13 (router)",0.042,false]]}}],"note":"Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.","sourceIds":["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic","price-google","price-openai","price-jev"]}],["cost-thought-experiments",{"id":"prompt-cache-savings","title":"Thought experiment: what prompt caching saved","subtitle":"The same recorded tokens with and without cache pricing","kind":"grouped-bar","unit":"usd","yLabel":"USD for all attempts","series":[{"name":"With caching (as recorded)","points":{"$k":["label","value","highlight"],"$r":[["Claude Haiku 4.5",43.61,false],["Claude Sonnet 5.5",87.23,true],["Claude Opus 5.5",143.83,false],["Claude Fable 5.1",321.31,false]]}},{"name":"Without caching","points":{"$k":["label","value"],"$r":[["Claude Haiku 4.5",171.66],["Claude Sonnet 5.5",343.33],["Claude Opus 5.5",686.66],["Claude Fable 5.1",1716.65]]}}],"note":"94.0% of recorded input tokens were cache reads. Calculation, not a run.","sourceIds":["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic"]}],["prompt-cache-break-even",{"id":"cache-break-even-reads","title":"Reuses before a cached prefix costs less, by model and write type (calculation)","subtitle":"The break-even point: reuses at which the cached and the uncached cost are equal, at list prices","kind":"grouped-bar","unit":"score","yLabel":"Reuses at break-even","series":{"$k":["name","points"],"$r":[["1-hour write (2× input), whole prefix new",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",1.11],["Claude Sonnet 5.5",1.11],["Claude Opus 5.5 (cache read $0.2 per M)",1.05],["Claude Opus 5.5 (cache read $0.4 per M)",1.11],["Claude Fable 5.1",1.03]]}],["1-hour write, pooled n = 6 session share, 19% already cached (as recorded)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0.72],["Claude Sonnet 5.5",0.72],["Claude Opus 5.5 (cache read $0.2 per M)",0.67],["Claude Opus 5.5 (cache read $0.4 per M)",0.72],["Claude Fable 5.1",0.65]]}],["5-minute write (1.25× input, an assumption)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0.28],["Claude Sonnet 5.5",0.28],["Claude Opus 5.5 (cache read $0.2 per M)",0.26],["Claude Opus 5.5 (cache read $0.4 per M)",0.28],["Claude Fable 5.1",0.26]]}]]},"note":"Calculation, not a run: break-even reuses = (write price − input price) ÷ (input price − cache-read price). A cached prefix costs less once the reuses pass that point. On a new prefix, a 1-hour write needs 2 reuses (the 3rd request). With 19% already cached it needs 1 reuse, and a 5-minute write needs 1 reuse. The 1.25 of the 5-minute write is an assumption. The caching study protocol states it as Anthropic’s published figure. No 5-minute write occurred in the recorded sessions. We recorded the 19% on Sonnet 5.5 and Opus 5.5 sessions. For Haiku 4.5 and Fable 5.1 it is a what-if. GPT-6.1 Sol and GPT-6 Luna list no write surcharge, so their break-even is 0 reuses and the chart leaves them out. The price list gives Opus 5.5 a cache read of $0.2 per million. Another table of the product lists $0.4. We did not check the vendor price, so the chart shows both.","sourceIds":["calc-cache-pricing","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-share","title":"Reasoning share of output tokens per call on hard tasks (calculation)","subtitle":"Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)","kind":"bar","unit":"percent","yLabel":"Reasoning share of output tokens (%)","series":[{"name":"Median call","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16]]}}],"note":"Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.","whisker":"minmax","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-cost-per-call","title":"List-price cost per call: reasoning, remaining output and input (calculation)","subtitle":"Mean per call on the hard tasks; the three parts add up to the call","kind":"stacked-bar","unit":"usd","yLabel":"USD per call (list price)","series":{"$k":["name","points"],"$r":[["Reasoning (output tokens)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.053696,24],["Claude Opus 5.5 (high) · Claude Code",0.017969,24],["Claude Haiku 4.5 · Claude Code",0.024492,24],["Claude Opus 5.5 · Claude Code",0.012528,24],["GPT-6.1 Sol (medium) · Codex CLI",0.002273,16],["GPT-6.1 Sol (high) · Codex CLI",0.004114,16],["Claude Sonnet 5.5 · Claude Code",0.006665,24]]}],["Remaining output (visible-answer estimate)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.018702,24],["Claude Opus 5.5 (high) · Claude Code",0.00799,24],["Claude Haiku 4.5 · Claude Code",0.001795,24],["Claude Opus 5.5 · Claude Code",0.008003,24],["GPT-6.1 Sol (medium) · Codex CLI",0.003025,16],["GPT-6.1 Sol (high) · Codex CLI",0.002905,16],["Claude Sonnet 5.5 · Claude Code",0.003672,24]]}],["Input (prompt, cache priced)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.02091,24],["Claude Opus 5.5 (high) · Claude Code",0.007407,24],["Claude Haiku 4.5 · Claude Code",0.00451,24],["Claude Opus 5.5 · Claude Code",0.00771,24],["GPT-6.1 Sol (medium) · Codex CLI",0.020339,16],["GPT-6.1 Sol (high) · Codex CLI",0.008117,16],["Claude Sonnet 5.5 · Claude Code",0.004012,24]]}]]},"note":"Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-short-vs-hard","title":"Reasoning share on short tasks vs hard tasks (calculation)","subtitle":"Median call per configuration; five short tasks and eight hard tasks","kind":"grouped-bar","unit":"percent","yLabel":"Reasoning share of output tokens (median call, %)","series":[{"name":"Eight hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16]]}},{"name":"Five short tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",90.19,73.1,97.59,15],["Claude Sonnet 5.5 · Claude Code",0,0,72.75,15],["Claude Opus 5.5 · Claude Code",0,0,93.33,15],["Claude Opus 5.5 (high) · Claude Code",43.59,0,93.33,15],["Claude Fable 5.1 · Claude Code",0,0,74.01,15],["GPT-6.1 Sol (medium) · Codex CLI",41.05,0,71.43,15],["GPT-6.1 Sol (high) · Codex CLI",58.06,0,75.76,15]]}}],"note":"Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.","whisker":"minmax","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["llm-speed-anatomy",{"id":"speed-anatomy-first-text","title":"Time to first text: a 250-line answer, six models","subtitle":"Median of 4 calls per model; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds to first text","series":[{"name":"Time to first text","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",4,2.84,6.38,4],["Claude Sonnet 5.5 · Claude Code",1.96,0.88,4.09,4],["Claude Opus 5.5 · Claude Code",1.97,1.7,2.35,4],["Claude Fable 5.1 · Claude Code",4.43,2.27,4.64,4],["GPT-6.1 Sol (low) · Codex CLI",3.52,2.75,4.42,4],["GPT-6 Luna (low) · Codex CLI",3.3,3.19,3.47,4]]}}],"note":"Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-output-speed","title":"Output speed after the first text: visible tokens per second (calculation)","subtitle":"Median of 4 calls per model; whiskers = slowest and fastest call","kind":"dot-range","unit":"tokens","yLabel":"Visible tokens per second","series":[{"name":"Visible tokens per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",153.2,152.6,216.1,4],["Claude Sonnet 5.5 · Claude Code",231.7,230.3,233,4],["Claude Opus 5.5 · Claude Code",155.5,154.6,156.4,4],["Claude Fable 5.1 · Claude Code",122.6,120.9,131.4,4],["GPT-6.1 Sol (low) · Codex CLI",79.6,71.6,80.5,4],["GPT-6 Luna (low) · Codex CLI",129.1,55.5,259.1,4]]}}],"note":"Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-chars-per-second","title":"Output speed in characters per second after the first text (calculation)","subtitle":"Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call","kind":"dot-range","unit":"count","yLabel":"Characters per second","series":[{"name":"Characters per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",547,546,548,3],["Claude Sonnet 5.5 · Claude Code",517,513,519,4],["Claude Opus 5.5 · Claude Code",347,345,349,4],["Claude Fable 5.1 · Claude Code",273,270,293,4],["GPT-6.1 Sol (low) · Codex CLI",323,291,327,4],["GPT-6 Luna (low) · Codex CLI",524,225,1052,4]]}}],"note":"Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}]]}