{"$k":["slug","chart"],"$r":[["model-head-to-head",{"id":"h2h-pass-rate","title":"Pass rate on five validated tasks","subtitle":"Every call counts; failures and timeouts are non-passes","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Pass rate","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Fable 5.1 · Claude Code",1,0.7961,1,15],["Claude Sonnet 5.5 · Claude Code",0.8,0.5481,0.9295,15],["Claude Opus 5.5 (high) · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 (low) · Claude Code",1,0.7961,1,15],["Claude Haiku 4.5 · Claude Code",1,0.7961,1,15],["GPT-6.1 Sol (high) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (low) · Codex CLI",1,0.7225,1,10]]}}],"note":"Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-total-latency","title":"Total time per call","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.94,1.41,9.83,15,true],["Claude Sonnet 5.5 · Claude Code",2.31,2.17,7.73,15,true],["Claude Opus 5.5 (high) · Claude Code",2.71,2.45,11.78,15,true],["Claude Opus 5.5 · Claude Code",2.75,2.47,8.91,15,true],["Claude Opus 5.5 (low) · Claude Code",2.83,2.35,6.62,15,true],["Claude Haiku 4.5 · Claude Code",4.43,3.16,23.57,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.6,4.05,19.52,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.65,4.1,25.46,15,false],["GPT-6.1 Sol (low) · Codex CLI",6.26,4.65,10.47,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-first-useful-latency","title":"Time to first useful output","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Time to first useful output","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.2,0.95,7.9,15,true],["Claude Sonnet 5.5 · Claude Code",1.56,0.99,6.39,15,true],["Claude Opus 5.5 (high) · Claude Code",2.04,1.4,9.94,15,true],["Claude Opus 5.5 · Claude Code",1.92,1.56,7.23,15,true],["Claude Opus 5.5 (low) · Claude Code",2.39,1.45,4.9,15,true],["Claude Haiku 4.5 · Claude Code",3.63,2.78,22.27,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.32,3.64,16.37,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.05,3.36,17.82,15,false],["GPT-6.1 Sol (low) · Codex CLI",5.14,4.02,8.5,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-input-tokens","title":"Input tokens per call: what the CLI sends","subtitle":"Mean per call, split into prompt-cache reads and other input","kind":"stacked-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Cache read","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",2760,15],["Claude Sonnet 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (high) · Claude Code",1463,15],["Claude Opus 5.5 · Claude Code",1401,15],["Claude Opus 5.5 (low) · Claude Code",1463,15],["Claude Haiku 4.5 · Claude Code",0,15],["GPT-6.1 Sol (high) · Codex CLI",6716,15],["GPT-6.1 Sol (medium) · Codex CLI",5180,15],["GPT-6.1 Sol (low) · Codex CLI",8064,10]]}},{"name":"Other input","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",473,15],["Claude Sonnet 5.5 · Claude Code",685,15],["Claude Opus 5.5 (high) · Claude Code",619,15],["Claude Opus 5.5 · Claude Code",680,15],["Claude Opus 5.5 (low) · Claude Code",618,15],["Claude Haiku 4.5 · Claude Code",3790,15],["GPT-6.1 Sol (high) · Codex CLI",5406,15],["GPT-6.1 Sol (medium) · Codex CLI",6943,15],["GPT-6.1 Sol (low) · Codex CLI",4059,10]]}}],"note":"The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-output-tokens","title":"Output tokens per call","subtitle":"Median per configuration; reasoning tokens where the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",64,15],["Claude Sonnet 5.5 · Claude Code",107,15],["Claude Opus 5.5 (high) · Claude Code",78,15],["Claude Opus 5.5 · Claude Code",64,15],["Claude Opus 5.5 (low) · Claude Code",64,15],["Claude Haiku 4.5 · Claude Code",367,15],["GPT-6.1 Sol (high) · Codex CLI",42,15],["GPT-6.1 Sol (medium) · Codex CLI",42,15],["GPT-6.1 Sol (low) · Codex CLI",42,10]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0,15],["Claude Sonnet 5.5 · Claude Code",0,15],["Claude Opus 5.5 (high) · Claude Code",34,15],["Claude Opus 5.5 · Claude Code",0,15],["Claude Opus 5.5 (low) · Claude Code",0,15],["Claude Haiku 4.5 · Claude Code",297,15],["GPT-6.1 Sol (high) · Codex CLI",21,15],["GPT-6.1 Sol (medium) · Codex CLI",20,15],["GPT-6.1 Sol (low) · Codex CLI",20,10]]}}],"note":"Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-list-price-per-call","title":"List-price cost per call (calculation)","subtitle":"Reported tokens × list price; the calls ran on subscriptions","kind":"dot-range","unit":"usd","yLabel":"USD per call","series":[{"name":"Cost per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",0.00987,0.0049,0.05843,15,true],["Claude Sonnet 5.5 · Claude Code",0.0036,0.00342,0.01021,15,true],["Claude Opus 5.5 (high) · Claude Code",0.00694,0.00592,0.02708,15,true],["Claude Opus 5.5 · Claude Code",0.00688,0.00592,0.02226,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00688,0.00582,0.01793,15,true],["Claude Haiku 4.5 · Claude Code",0.00566,0.00513,0.01804,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01047,0.0066,0.02812,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01018,0.0054,0.02686,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00769,0.00742,0.02649,10,false]]}}],"note":"Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["model-head-to-head",{"id":"h2h-cost-per-pass","title":"List-price cost per passing answer (calculation)","subtitle":"All calls in a configuration, failures included, divided by its passes","kind":"bar","unit":"usd","yLabel":"USD per passing answer","series":[{"name":"Cost per pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.00624,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00829,15,true],["Claude Haiku 4.5 · Claude Code",0.00836,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00998,10,false],["Claude Opus 5.5 · Claude Code",0.01009,15,true],["Claude Opus 5.5 (high) · Claude Code",0.01049,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01322,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01564,15,false],["Claude Fable 5.1 · Claude Code",0.02054,15,true]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["hard-model-head-to-head",{"id":"hard-h2h-pass-rate","title":"Pass rate on eight hard tasks","subtitle":"Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.4583,0.2789,0.6493,24]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.6667,0.4671,0.8203,24]]}}],"note":"Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-total-latency","title":"Total time per call on hard tasks (separate batches)","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call on hard tasks (separate batches)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",7.75,2.26,34.79,24,true],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,24,true],["Claude Opus 5.5 (high) · Claude Code",11.03,3.63,63,24,true],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16,true],["Claude Fable 5.1 · Claude Code",16.13,4.46,90,24,true],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16,true],["Claude Haiku 4.5 · Claude Code",39.01,15.27,75.13,24,false]]}}],"note":"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-first-useful-latency","title":"Time to first useful output on hard tasks","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Time to first useful output on hard tasks","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",5.95,0.86,30.57,24,true],["Claude Opus 5.5 · Claude Code",6.78,2.39,21.77,24,true],["Claude Opus 5.5 (high) · Claude Code",7.13,2.15,56.23,24,true],["GPT-6.1 Sol (medium) · Codex CLI",10.23,6.09,40.41,16,true],["Claude Fable 5.1 · Claude Code",11.63,2,85.33,24,true],["GPT-6.1 Sol (high) · Codex CLI",12.69,8.93,75.91,16,true],["Claude Haiku 4.5 · Claude Code",35.54,12.88,70.31,24,false]]}}],"note":"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-output-tokens","title":"Output tokens per call on hard tasks","subtitle":"Median per configuration; reasoning tokens as the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1050,24],["Claude Opus 5.5 · Claude Code",945,24],["Claude Opus 5.5 (high) · Claude Code",1052,24],["GPT-6.1 Sol (medium) · Codex CLI",335,16],["Claude Fable 5.1 · Claude Code",1366,24],["GPT-6.1 Sol (high) · Codex CLI",436,16],["Claude Haiku 4.5 · Claude Code",5064,24]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",585,24],["Claude Opus 5.5 · Claude Code",529,24],["Claude Opus 5.5 (high) · Claude Code",614,24],["GPT-6.1 Sol (medium) · Codex CLI",150,16],["Claude Fable 5.1 · Claude Code",889,24],["GPT-6.1 Sol (high) · Codex CLI",225,16],["Claude Haiku 4.5 · Claude Code",4556,24]]}}],"note":"Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-cost-per-pass","title":"List-price cost per strict pass on hard tasks (calculation)","subtitle":"All calls in a configuration, failures and format misses included, divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.01435,24,true],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16,false],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16,false],["Claude Opus 5.5 · Claude Code",0.02824,24,false],["Claude Opus 5.5 (high) · Claude Code",0.03337,24,false],["Claude Haiku 4.5 · Claude Code",0.0672,24,false],["Claude Fable 5.1 · Claude Code",0.09331,24,false]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.","sourceIds":["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]}],["caching-consistency",{"id":"consistency-pass-rate","title":"Same prompt, 10 times: strict pass rate","subtitle":"One series per prompt; whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","yLabel":"Passed","series":{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0,0,0.2775,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["JSON object",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0.1,0.0179,0.4042,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["Code fix",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,0.7225,1,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}]]},"note":"Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.","whisker":"ci95","sourceIds":["agent-caching-consistency"]}],["caching-consistency",{"id":"consistency-distinct-answers","title":"Same prompt, 10 times: how many different answers","subtitle":"Distinct normalized answers over 10 repetitions (1 = the same answer every time)","kind":"grouped-bar","unit":"count","yLabel":"Distinct answers","series":{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,10],["Claude Sonnet 5.5 · Claude Code",1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,10]]}],["JSON object",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,10],["Claude Sonnet 5.5 · Claude Code",1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,10]]}],["Code fix",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 · Claude Code",6,10],["Claude Sonnet 5.5 · Claude Code",3,10],["GPT-6.1 Sol (medium) · Codex CLI",6,10]]}]]},"note":"Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.","sourceIds":["agent-caching-consistency"]}],["caching-consistency",{"id":"consistency-latency-spread","title":"Same prompt, 10 times: time per call","subtitle":"Median; whiskers = fastest and slowest of 10 calls","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",5.06,4.42,6.2,10],["Claude Sonnet 5.5 · Claude Code",6.89,5.81,7.81,10],["GPT-6.1 Sol (medium) · Codex CLI",13.38,12.29,17.97,10]]}],["JSON object",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",7.03,5.28,12.27,10],["Claude Sonnet 5.5 · Claude Code",2.89,2.68,5.3,10],["GPT-6.1 Sol (medium) · Codex CLI",6.42,5.25,8.26,10]]}],["Code fix",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",5.95,4.89,7.33,10],["Claude Sonnet 5.5 · Claude Code",2.67,2.32,4.34,10],["GPT-6.1 Sol (medium) · Codex CLI",11.29,9.08,14.85,10]]}]]},"note":"Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.","whisker":"minmax","sourceIds":["agent-caching-consistency"]}],["agent-memory",{"id":"memory-full-pass","title":"Full pass rate by kind of memory","subtitle":"Hidden tests pass and every convention check passes · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["No memory",0.6,0.3575,0.8018,15,"\u0001"],["/init CLAUDE.md",0.6,0.3575,0.8018,15,"\u0001"],["Curated, 11 lines",1,0.7961,1,15,true],["Raw notes, 60 lines",0.9333,0.7018,0.9881,15,"\u0001"],["Dreamed notes",1,0.7961,1,15,"\u0001"],["Handbook, 210 lines",1,0.7961,1,15,"\u0001"],["Stop hook only",0.8,0.5481,0.9295,15,"\u0001"],["Curated + hook",1,0.7961,1,15,"\u0001"]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.2,0.0567,0.5098,10],["/init CLAUDE.md",0.2,0.0567,0.5098,10],["Curated, 11 lines",0.7,0.3968,0.8922,10],["Raw notes, 60 lines",0.6,0.3127,0.8318,10],["Dreamed notes",0.7,0.3968,0.8922,10],["Handbook, 210 lines",0.3,0.1078,0.6032,10],["Stop hook only",0.8,0.4902,0.9433,10],["Curated + hook",0.9,0.5958,0.9821,10]]}}],"note":"Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.","sourceIds":["agent-memory-study"]}],["agent-memory",{"id":"memory-team-knowledge-by-model","title":"Team knowledge followed, Sonnet vs Haiku","subtitle":"Changelog rule and late-fee rate, pooled · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.4,0.1982,0.6425,15],["/init CLAUDE.md",0.6667,0.4171,0.8482,15],["Curated, 11 lines",1,0.7961,1,15],["Raw notes, 60 lines",1,0.7961,1,15],["Dreamed notes",1,0.7961,1,15],["Handbook, 210 lines",1,0.7961,1,15],["Stop hook only",0.6667,0.4171,0.8482,15],["Curated + hook",1,0.7961,1,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0,0,0.2775,10],["/init CLAUDE.md",0.1,0.0179,0.4042,10],["Curated, 11 lines",0.8,0.4902,0.9433,10],["Raw notes, 60 lines",0.6,0.3127,0.8318,10],["Dreamed notes",0.8,0.4902,0.9433,10],["Handbook, 210 lines",0.3,0.1078,0.6032,10],["Stop hook only",0.8,0.4902,0.9433,10],["Curated + hook",1,0.7225,1,10]]}}],"note":"Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.","sourceIds":["agent-memory-study"]}],["agent-memory",{"id":"memory-broken-test-command","title":"A stale README command: who still ran it?","polarity":"lower","subtitle":"Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.8,0.5481,0.9295,15],["/init CLAUDE.md",0.8667,0.6212,0.9626,15],["Curated, 11 lines",0,0,0.2039,15],["Raw notes, 60 lines",0,0,0.2039,15],["Dreamed notes",0,0,0.2039,15],["Handbook, 210 lines",0,0,0.2039,15],["Stop hook only",0.6,0.3575,0.8018,15],["Curated + hook",0,0,0.2039,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",1,0.7225,1,10],["/init CLAUDE.md",1,0.7225,1,10],["Curated, 11 lines",0,0,0.2775,10],["Raw notes, 60 lines",1,0.7225,1,10],["Dreamed notes",0.1,0.0179,0.4042,10],["Handbook, 210 lines",0,0,0.2775,10],["Stop hook only",1,0.7225,1,10],["Curated + hook",0,0,0.2775,10]]}}],"note":"The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old \"run npm test\" note and a later correction. The /init file repeats the README.","sourceIds":["agent-memory-study"]}],["agent-memory",{"id":"memory-cost-per-full-pass","title":"List-price cost per fully correct result (calculation)","subtitle":"Sum of the CLI's cost estimates for a condition, divided by its full passes","kind":"grouped-bar","unit":"usd","series":[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","n"],"$r":[["No memory",0.1386,9],["/init CLAUDE.md",0.1359,9],["Curated, 11 lines",0.0818,15],["Raw notes, 60 lines",0.1009,14],["Dreamed notes",0.0896,15],["Handbook, 210 lines",0.1006,15],["Stop hook only",0.1278,12],["Curated + hook",0.0843,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","n"],"$r":[["No memory",0.3786,2],["/init CLAUDE.md",0.4317,2],["Curated, 11 lines",0.1095,7],["Raw notes, 60 lines",0.1255,6],["Dreamed notes",0.1153,7],["Handbook, 210 lines",0.2615,3],["Stop hook only",0.1419,8],["Curated + hook",0.096,9]]}}],"note":"Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.","sourceIds":["agent-memory-study"]}],["agent-memory",{"id":"memory-wall-time","title":"Time per session","subtitle":"Median wall time in seconds","kind":"grouped-bar","unit":"seconds","series":[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","n"],"$r":[["No memory",18,15],["/init CLAUDE.md",19,15],["Curated, 11 lines",21.9,15],["Raw notes, 60 lines",26.8,15],["Dreamed notes",27.2,15],["Handbook, 210 lines",23.6,15],["Stop hook only",27.3,15],["Curated + hook",22,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","n"],"$r":[["No memory",54.2,10],["/init CLAUDE.md",52.9,10],["Curated, 11 lines",51.7,10],["Raw notes, 60 lines",51,10],["Dreamed notes",51.8,10],["Handbook, 210 lines",49.9,10],["Stop hook only",68.5,10],["Curated + hook",52.8,10]]}}],"note":"Up to four sessions ran at a time on one machine. Sessions without memory were often shorter because they stopped to ask or skipped the changelog.","sourceIds":["agent-memory-study"]}],["routing-jev-vs-llm",{"id":"routing-exact-decisions","title":"Typed routing decisions answered exactly right","subtitle":"Share of asked cases where every scored question was acceptable","kind":"dot-range","unit":"rate","yLabel":"Exact","whisker":"ci95","series":[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8984,0.8191,0.9497,82,true],["Claude Haiku 4.5",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5",0.939,0.8651,0.9737,82,false]]}}],"note":"Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-key-accuracy","title":"Per-question accuracy","subtitle":"Each open question the router was asked; an unanswered question counts as wrong","kind":"dot-range","unit":"rate","yLabel":"Correct answers","whisker":"ci95","series":[{"name":"Key accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.9485,0.9077,0.9718,194,true],["Claude Haiku 4.5",0.9433,0.9013,0.968,194,false],["Claude Sonnet 5.5",0.9742,0.9411,0.9889,194,false]]}}],"note":"Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-exact-by-decision","title":"Exact rate by decision type","kind":"grouped-bar","unit":"rate","yLabel":"Exact","series":{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.8241,1,18,true],["Message intent",1,0.8389,1,20,true],["Is it a rule?",1,0.7575,1,12,true],["Context shape",0.7396,0.5789,0.8675,32,true]]}],["Claude Haiku 4.5",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9444,0.7424,0.9901,18,false],["Message intent",1,0.8389,1,20,false],["Is it a rule?",1,0.7575,1,12,false],["Context shape",0.75,0.5789,0.8675,32,false]]}],["Claude Sonnet 5.5",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.8241,1,18,false],["Message intent",1,0.8389,1,20,false],["Is it a rule?",1,0.7575,1,12,false],["Context shape",0.8438,0.6825,0.9314,32,false]]}]]},"note":"A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-cost-per-1000","title":"Cost per 1,000 routing decisions","subtitle":"List price × reported tokens per decision","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost","points":{"$k":["label","value","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.0337,246,true],["Claude Haiku 4.5",8.924,82,false],["Claude Sonnet 5.5",4.996,82,false]]}}],"note":"List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.","sourceIds":["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-decision-latency","title":"Time per routing decision","subtitle":"Median wall time, whisker to the 95th percentile","kind":"dot-range","unit":"ms","yLabel":"Time per decision","series":{"$k":["name","points"],"$r":[["Wall time (CLI)",[{"label":"Claude Haiku 4.5","value":12674,"lo":12674,"hi":34413,"n":82},{"label":"Claude Sonnet 5.5","value":2598,"lo":2598,"hi":4298,"n":82}]],["Wall time (direct API call)",[{"label":"Jev 1.13 (TypeSafe)","value":136.5,"lo":136.5,"hi":195.7,"n":246}]],["Model time (API)",[{"label":"Claude Haiku 4.5","value":10734,"lo":10734,"hi":32072,"n":82},{"label":"Claude Sonnet 5.5","value":1599,"lo":1599,"hi":2574,"n":82}]]]},"note":"Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.","whisker":"p50-p95","sourceIds":["agent-routing","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-decision-latency","title":"Time to make one routing decision","subtitle":"Median; whiskers = median to 95th percentile","kind":"dot-range","unit":"ms","yLabel":"Time per decision","whisker":"p50-p95","series":[{"name":"Decision time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",0.00142,0.00142,0.00233,20000,true],["Jev 1.13 (TypeSafe)",136.5,136.5,195.7,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",2597,2597,4298,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",12543,12543,34481,82,"\u0001"]]}}],"note":"The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.","sourceIds":["agent-routing-overhead","agent-routing","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-cli-vs-model-time","title":"Where an LLM router’s time goes: model vs CLI","subtitle":"Median per call; whiskers = median to 95th percentile","kind":"dot-range","unit":"ms","yLabel":"Time per call","whisker":"p50-p95","series":[{"name":"Model API time","points":[{"label":"Claude Sonnet 5.5 (effort low, via Claude Code)","value":1596,"lo":1596,"hi":2583,"n":82},{"label":"Claude Haiku 4.5 (thinking on, via Claude Code)","value":10508,"lo":10508,"hi":32132,"n":82}]},{"name":"CLI and harness time","points":[{"label":"Claude Sonnet 5.5 (effort low, via Claude Code)","value":973,"lo":973,"hi":1277,"n":82},{"label":"Claude Haiku 4.5 (thinking on, via Claude Code)","value":1698,"lo":1698,"hi":2677,"n":82}]}],"note":"Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.","sourceIds":["agent-routing-overhead","agent-routing"]}],["routing-overhead",{"id":"router-overhead-completed","title":"Routing calls that returned a decision","subtitle":"Completed calls ÷ calls; whiskers = 95% Wilson interval","kind":"dot-range","unit":"rate","yLabel":"Completed","whisker":"ci95","series":[{"name":"Completed","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",1,0.9998,1,20000,true],["Jev 1.13 (TypeSafe)",1,0.9846,1,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",1,0.9552,1,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",1,0.9552,1,82,"\u0001"]]}}],"note":"A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.","sourceIds":["agent-routing-overhead","agent-routing","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-cost-per-1000-tasks","title":"Added routing cost per 1,000 tasks (calculation)","subtitle":"Decisions per task from recorded runs × cost per decision","kind":"grouped-bar","unit":"usd","yLabel":"USD per 1,000 tasks","series":[{"name":"Every model call routed (49.5 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0],["Jev 1.13 (TypeSafe)",1.67],["Claude Sonnet 5.5 (effort low, via Claude Code)",247.3],["Claude Haiku 4.5 (thinking on, via Claude Code)",441.74]]}},{"name":"Only System One decisions (7 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0],["Jev 1.13 (TypeSafe)",0.24],["Claude Sonnet 5.5 (effort low, via Claude Code)",34.97],["Claude Haiku 4.5 (thinking on, via Claude Code)",62.47]]}}],"note":"A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).","sourceIds":["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-delay-per-task","title":"Added routing delay per task (calculation)","subtitle":"Decisions per task × median decision time, if every decision waits in line","kind":"grouped-bar","unit":"seconds","yLabel":"Seconds per task","series":[{"name":"Every model call routed (49.5 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0.0000703],["Jev 1.13 (TypeSafe)",6.7568],["Claude Sonnet 5.5 (effort low, via Claude Code)",128.5515],["Claude Haiku 4.5 (thinking on, via Claude Code)",620.8785]]}},{"name":"Only System One decisions (7 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0.0000099],["Jev 1.13 (TypeSafe)",0.9555],["Claude Sonnet 5.5 (effort low, via Claude Code)",18.179],["Claude Haiku 4.5 (thinking on, via Claude Code)",87.801]]}}],"note":"A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.","sourceIds":["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]}],["single-call-vs-agent-loop",{"id":"agent-loop-pass-rate","title":"Strict pass rate: single call vs agent loop on eight hard tasks","subtitle":"Same tasks and validators. Whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",0.4583,0.2789,0.6493,24],["Claude Haiku 4.5 (agent loop) · Claude Code",0.5417,0.3507,0.7211,24],["Claude Sonnet 5.5 (single call) · Claude Code",1,0.862,1,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",1,0.8064,1,16],["GPT-6 Luna (single call) · Codex CLI",0.625,0.3864,0.8152,16],["GPT-6 Luna (agent loop) · Codex CLI",0.8571,0.6006,0.9599,14]]}}],"note":"Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.","whisker":"ci95","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["single-call-vs-agent-loop",{"id":"agent-loop-by-task","title":"Strict passes per task: single call vs agent loop","subtitle":"Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Passed","series":{"$k":["name","points"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.4385,1,3],["DST day-length fix",0.3333,0.0615,0.7923,3],["CSV parser",0.6667,0.2077,0.9385,3],["Event-loop order",0,0,0.5615,3],["Room schedule",0,0,0.5615,3],["SemVer regex",1,0.4385,1,3],["Money refactor",0.6667,0.2077,0.9385,3],["SQL report",0,0,0.5615,3]]}],["Claude Haiku 4.5 (agent loop) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",0.6667,0.2077,0.9385,3],["DST day-length fix",0.6667,0.2077,0.9385,3],["CSV parser",0.6667,0.2077,0.9385,3],["Event-loop order",1,0.4385,1,3],["Room schedule",0.6667,0.2077,0.9385,3],["SemVer regex",0.6667,0.2077,0.9385,3],["Money refactor",0,0,0.5615,3],["SQL report",0,0,0.5615,3]]}],["Claude Sonnet 5.5 (single call) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.4385,1,3],["DST day-length fix",1,0.4385,1,3],["CSV parser",1,0.4385,1,3],["Event-loop order",1,0.4385,1,3],["Room schedule",1,0.4385,1,3],["SemVer regex",1,0.4385,1,3],["Money refactor",1,0.4385,1,3],["SQL report",1,0.4385,1,3]]}],["Claude Sonnet 5.5 (agent loop) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.3424,1,2],["DST day-length fix",1,0.3424,1,2],["CSV parser",1,0.3424,1,2],["Event-loop order",1,0.3424,1,2],["Room schedule",1,0.3424,1,2],["SemVer regex",1,0.3424,1,2],["Money refactor",1,0.3424,1,2],["SQL report",1,0.3424,1,2]]}],["GPT-6 Luna (single call) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.3424,1,2],["DST day-length fix",1,0.3424,1,2],["CSV parser",1,0.3424,1,2],["Event-loop order",0,0,0.6576,2],["Room schedule",0.5,0.0945,0.9055,2],["SemVer regex",0.5,0.0945,0.9055,2],["Money refactor",0,0,0.6576,2],["SQL report",1,0.3424,1,2]]}],["GPT-6 Luna (agent loop) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.3424,1,2],["CSV parser",0.5,0.0945,0.9055,2],["Event-loop order",1,0.3424,1,2],["Room schedule",1,0.3424,1,2],["SemVer regex",1,0.3424,1,2],["Money refactor",0.5,0.0945,0.9055,2],["SQL report",1,0.3424,1,2]]}]]},"note":"Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.","whisker":"ci95","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["single-call-vs-agent-loop",{"id":"agent-loop-total-time","title":"Total time per attempt: single call vs agent loop","subtitle":"Median per configuration; whiskers = fastest and slowest attempt","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",39.01,15.27,75.13,24],["Claude Haiku 4.5 (agent loop) · Claude Code",56.77,24.53,223.7,24],["Claude Sonnet 5.5 (single call) · Claude Code",7.75,2.26,34.79,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",7.41,2.75,24.19,16],["GPT-6 Luna (single call) · Codex CLI",5.16,3.59,11.32,16],["GPT-6 Luna (agent loop) · Codex CLI",9.32,3.78,15.89,14]]}}],"note":"Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.","whisker":"minmax","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["single-call-vs-agent-loop",{"id":"agent-loop-tokens","title":"Tokens per attempt: single call vs agent loop","subtitle":"Median per configuration; whiskers = fewest and most","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Input tokens (cache reads included)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",3941,3879,4221,24],["Claude Haiku 4.5 (agent loop) · Claude Code",71691,41732,516306,24],["Claude Sonnet 5.5 (single call) · Claude Code",2281,2234,2669,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",9550,9398,33040,16],["GPT-6 Luna (single call) · Codex CLI",11582,11526,11818,16],["GPT-6 Luna (agent loop) · Codex CLI",15530,15391,39009,14]]}},{"name":"Output tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",5064,1899,9321,24],["Claude Haiku 4.5 (agent loop) · Claude Code",7912,2541,20654,24],["Claude Sonnet 5.5 (single call) · Claude Code",1050,176,3895,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",876,219,3243,16],["GPT-6 Luna (single call) · Codex CLI",345,36,634,16],["GPT-6 Luna (agent loop) · Codex CLI",480,143,858,14]]}}],"note":"Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.","whisker":"minmax","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["single-call-vs-agent-loop",{"id":"agent-loop-tool-calls","title":"Tool calls per agent-loop attempt","subtitle":"Median per configuration; whiskers = fewest and most. A single call makes none","kind":"dot-range","unit":"count","yLabel":"Tool calls","series":[{"name":"Tool calls per attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (agent loop) · Claude Code",3,2,18,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",0,0,3,16],["GPT-6 Luna (agent loop) · Codex CLI",0,0,1,14]]}}],"note":"Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.","whisker":"minmax","sourceIds":["agent-agent-loop"]}],["single-call-vs-agent-loop",{"id":"agent-loop-cost-per-pass","title":"List-price cost per strict pass: single call vs agent loop (calculation)","subtitle":"All attempts in a configuration divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",0.0672,24],["Claude Haiku 4.5 (agent loop) · Claude Code",0.14225,24],["Claude Sonnet 5.5 (single call) · Claude Code",0.01435,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",0.02746,16],["GPT-6 Luna (single call) · Codex CLI",0.00116,16],["GPT-6 Luna (agent loop) · Codex CLI",0.00099,14]]}}],"note":"Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.","sourceIds":["agent-agent-loop","agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]}],["haiku-thinking-on-off",{"id":"haiku-thinking-router-exact","title":"Haiku thinking study: typed routing decisions answered exactly right","subtitle":"Claude Haiku 4.5 and Claude Sonnet 5.5 (low effort); 82 decisions, the same cases for every arm","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Correct","series":[{"name":"Exact decisions (every scored question right)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0.8659,0.7755,0.9234,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5 (low) · Claude Code",0.939,0.8651,0.9737,82,false]]}},{"name":"Per-question accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0.9124,0.8642,0.9446,194,true],["Claude Haiku 4.5 (thinking on) · Claude Code",0.9433,0.9013,0.968,194,false],["Claude Sonnet 5.5 (low) · Claude Code",0.9742,0.9411,0.9889,194,false]]}}],"note":"Whiskers are 95% Wilson intervals. Questions within a decision are related; per-question intervals are descriptive, not an independent-question test. The thinking-on and Sonnet arms are the recorded 2026-10-05 routing run, reused, not rerun; the thinking-off arm ran later on another account. An unanswered question counts as wrong.","whisker":"ci95","factContext":"typed routing decisions, thinking on vs off","sourceIds":["agent-haiku-thinking","agent-routing"]}],["haiku-thinking-on-off",{"id":"haiku-thinking-router-latency","title":"Haiku thinking study: time per routing decision","subtitle":"Median wall time and model (API) time; whisker to the 95th percentile","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Wall time (CLI)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",4.66,4.66,8.18,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",12.54,12.54,34.48,82,false],["Claude Sonnet 5.5 (low) · Claude Code",2.6,2.6,4.3,82,false]]}},{"name":"Model time (API)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",3.79,3.79,7.43,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",10.51,10.51,32.13,82,false],["Claude Sonnet 5.5 (low) · Claude Code",1.6,1.6,2.58,82,false]]}}],"note":"Whiskers run from p50 to p95, not a confidence interval. Percentiles use the nearest-rank rule of the routing-overhead study, so the thinking-on and Sonnet medians match that study (the Jev-vs-LLM page interpolates between ranks and shows slightly different values). One call at a time through the Claude Code CLI; the thinking-on and Sonnet arms ran on another day.","whisker":"p50-p95","factContext":"typed routing decisions, thinking on vs off","sourceIds":["agent-haiku-thinking","agent-routing"]}],["haiku-thinking-on-off",{"id":"haiku-thinking-router-tokens","title":"Haiku thinking study: thinking and visible output tokens per routing decision","subtitle":"Mean per decision, as the Claude Code CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens per decision","series":[{"name":"Thinking tokens","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0,82],["Claude Haiku 4.5 (thinking on) · Claude Code",1101,82],["Claude Sonnet 5.5 (low) · Claude Code",2,82]]}},{"name":"Visible output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",366,82],["Claude Haiku 4.5 (thinking on) · Claude Code",318,82],["Claude Sonnet 5.5 (low) · Claude Code",105,82]]}}],"note":"Visible output = output tokens minus thinking tokens. It includes the structured answer the CLI asks for. Thinking tokens are counted by the CLI; their content is never captured. More tokens is not better or worse by itself.","factContext":"typed routing decisions, thinking on vs off","sourceIds":["agent-haiku-thinking","agent-routing"]}],["haiku-thinking-on-off",{"id":"haiku-thinking-router-cost","title":"Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)","subtitle":"Reported tokens × list price, per 1,000 decisions","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost per 1,000 decisions","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",3.364,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",8.924,82,false],["Claude Sonnet 5.5 (low) · Claude Code",7.324,82,false]]}}],"note":"Calculation, not a bill: the calls ran on a flat subscription. Reported input, cache and output tokens (thinking tokens are part of output) × list price, with the price table of the Jev-vs-LLM study. Sonnet cache writes use the one-hour rate ($4 per million tokens), as recorded in its receipts. The calculation matches the CLI-reported cost.","factContext":"typed routing decisions, thinking on vs off","sourceIds":["agent-haiku-thinking","agent-routing","calc-repricing","price-anthropic"]}],["json-schema-vs-instructions",{"id":"structured-output-pass-rate","title":"Does a JSON schema raise the pass rate? Instructions vs schema mode","subtitle":"Three extraction prompts pooled; whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Passed","series":[{"name":"Strict pass: the whole reply is the right JSON","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0,0,0.138,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}},{"name":"Right answer in any format (strict pass or format miss)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0.7083,0.5083,0.8509,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}}],"note":"Whiskers are 95% Wilson intervals (a calculation) over 24 calls and 12 calls per configuration; every error counts as a fail. Strict: the whole reply parses as JSON and matches the expected answer exactly. A format miss is a right answer inside a code fence or prose, so it is never a strict pass.","whisker":"ci95","sourceIds":["agent-structured-output"]}],["json-schema-vs-instructions",{"id":"structured-output-outcomes","title":"What each call produced: strict pass, format miss, wrong values or error","subtitle":"Counts of calls per configuration; the three prompts pooled","kind":"stacked-bar","unit":"count","yLabel":"Calls","series":{"$k":["name","points"],"$r":[["Strict pass",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",18,24],["Claude Sonnet 5.5 (instructions) · Claude Code",12,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",12,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",12,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",12,12]]}],["Format miss",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",17,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0,24],["Claude Sonnet 5.5 (instructions) · Claude Code",0,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",0,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",0,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",0,12]]}],["Wrong values",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",7,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",6,24],["Claude Sonnet 5.5 (instructions) · Claude Code",0,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",0,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",0,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",0,12]]}],["Error",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0,24],["Claude Sonnet 5.5 (instructions) · Claude Code",0,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",0,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",0,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",0,12]]}]]},"note":"Counts of calls, not rates; the pass-rate chart carries the same results with 95% intervals. Format miss: the right answer inside a code fence or prose. Wrong values: any other completed reply, with a wrong value, key or type (a reply that sits in a code fence and also has a wrong value is counted here). Error: the call did not complete.","sourceIds":["agent-structured-output"]}],["json-schema-vs-instructions",{"id":"structured-output-time","title":"Time per call, instructions vs schema mode","subtitle":"Median; whiskers = fastest and slowest completed call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Median time per call (the three prompts pooled)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",9.52,5.67,17,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",8.46,5.9,12.23,24],["Claude Sonnet 5.5 (instructions) · Claude Code",3.52,2.67,4.12,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",4.2,2.95,6.14,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",6.21,4.2,12.27,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",5.96,4.62,20.97,12]]}}],"note":"Whiskers are a range (fastest and slowest call), not a confidence interval. Wall time from process start to exit, so it includes CLI start-up; Codex CLI timings include its larger system prompt. The three prompts differ in length, which widens every range.","whisker":"minmax","sourceIds":["agent-structured-output"]}],["json-schema-vs-instructions",{"id":"structured-output-tokens","title":"Output and reasoning tokens per call, instructions vs schema mode","subtitle":"Median per call; ranges and sample sizes are in the note","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Median output tokens per call","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",1128,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",1036,24],["Claude Sonnet 5.5 (instructions) · Claude Code",368,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",424,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",117,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",123,12]]}},{"name":"Median reasoning tokens per call (thinking)","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",934,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",727,24],["Claude Sonnet 5.5 (instructions) · Claude Code",182,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",109,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",21,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",23,12]]}}],"note":"Output tokens include reasoning tokens. Claude schema-mode output includes the CLI’s structured-output tool call. Input counts include CLI context and are not compared. Ranges (not intervals): Claude Haiku 4.5 (instructions) · Claude Code: output 698 to 2059 (n = 24); reasoning 559 to 1830 (n = 24); Claude Haiku 4.5 (JSON schema) · Claude Code: output 716 to 1462 (n = 24); reasoning 522 to 1187 (n = 24); Claude Sonnet 5.5 (instructions) · Claude Code: output 154 to 456 (n = 12); reasoning 54 to 278 (n = 12); Claude Sonnet 5.5 (JSON schema) · Claude Code: output 273 to 541 (n = 12); reasoning 0 to 260 (n = 12); GPT-6.1 Sol (low, instructions) · Codex CLI: output 69 to 259 (n = 12); reasoning 0 to 60 (n = 12); GPT-6.1 Sol (low, JSON schema) · Codex CLI: output 98 to 181 (n = 12); reasoning 0 to 49 (n = 12).","sourceIds":["agent-structured-output"]}],["routing-holdout",{"id":"routing-holdout-exact","title":"Unseen routing decisions answered exactly right","subtitle":"Share of the 56 holdout cases where every scored question was acceptable","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Exact","series":[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"note":"Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-key-accuracy","title":"Per-question accuracy on unseen decisions","subtitle":"Each open question a router was asked; an unanswered question counts as wrong","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Correct answers","series":[{"name":"Key accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.904,0.8397,0.9442,125,true],["Claude Haiku 4.5 · Claude Code",0.816,0.739,0.8741,125,false],["Claude Sonnet 5.5 (low) · Claude Code",0.92,0.859,0.956,125,false]]}}],"note":"Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-by-purpose","title":"Exact rate on unseen decisions, by decision type","subtitle":"14 cases per decision type","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Exact","series":{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9286,0.6853,0.9873,14,true],["Message intent",0.8571,0.6006,0.9599,14,true],["Is it a rule?",0.9286,0.6853,0.9873,14,true],["Context shape",0.5714,0.3259,0.7862,14,true]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9286,0.6853,0.9873,14,false],["Message intent",0.9286,0.6853,0.9873,14,false],["Is it a rule?",0.9286,0.6853,0.9873,14,false],["Context shape",0.3571,0.1634,0.6124,14,false]]}],["Claude Sonnet 5.5 (low) · Claude Code",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.7847,1,14,false],["Message intent",1,0.7847,1,14,false],["Is it a rule?",0.9286,0.6853,0.9873,14,false],["Context shape",0.5714,0.3259,0.7862,14,false]]}]]},"note":"Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-tuned-vs-unseen","title":"Tuned case set vs unseen holdout: exact rate per router","subtitle":"Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Exact","series":[{"name":"Tuned set (routing-jev-vs-llm)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.9024,0.8191,0.9497,82,true],["Claude Haiku 4.5 · Claude Code",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5 (low) · Claude Code",0.939,0.8651,0.9737,82,false]]}},{"name":"Unseen holdout","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"note":"Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.","whisker":"ci95","sourceIds":["agent-routing-holdout","agent-routing"]}],["routing-holdout",{"id":"routing-holdout-latency","title":"Time per routing decision, by route","subtitle":"Median, whisker to the 95th percentile","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Wall time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.139,0.139,0.192,168,true],["Claude Haiku 4.5 · Claude Code",9.444,9.444,25.413,56,false],["Claude Sonnet 5.5 (low) · Claude Code",2.359,2.359,3.657,56,false]]}},{"name":"Model time (API, CLI-reported)","points":[{"label":"Claude Haiku 4.5 · Claude Code","value":7.522,"lo":7.522,"hi":23.913,"n":56},{"label":"Claude Sonnet 5.5 (low) · Claude Code","value":1.485,"lo":1.485,"hi":2.377,"n":56}]}],"note":"The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.","whisker":"p50-p95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-cost-per-1000","title":"Cost per 1,000 unseen routing decisions","subtitle":"Reported tokens per decision × list price","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost","points":{"$k":["label","value","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.03065,168,true],["Claude Haiku 4.5 · Claude Code",7.129,56,false],["Claude Sonnet 5.5 (low) · Claude Code",7.244,56,false]]}}],"note":"Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.","sourceIds":["agent-routing-holdout","calc-repricing","price-jev","price-anthropic"]}],["thinking-token-bill",{"id":"thinking-bill-share","title":"Reasoning share of output tokens per call on hard tasks (calculation)","subtitle":"Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)","kind":"bar","unit":"percent","yLabel":"Reasoning share of output tokens (%)","series":[{"name":"Median call","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16]]}}],"note":"Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.","whisker":"minmax","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-cost-per-call","title":"List-price cost per call: reasoning, remaining output and input (calculation)","subtitle":"Mean per call on the hard tasks; the three parts add up to the call","kind":"stacked-bar","unit":"usd","yLabel":"USD per call (list price)","series":{"$k":["name","points"],"$r":[["Reasoning (output tokens)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.053696,24],["Claude Opus 5.5 (high) · Claude Code",0.017969,24],["Claude Haiku 4.5 · Claude Code",0.024492,24],["Claude Opus 5.5 · Claude Code",0.012528,24],["GPT-6.1 Sol (medium) · Codex CLI",0.002273,16],["GPT-6.1 Sol (high) · Codex CLI",0.004114,16],["Claude Sonnet 5.5 · Claude Code",0.006665,24]]}],["Remaining output (visible-answer estimate)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.018702,24],["Claude Opus 5.5 (high) · Claude Code",0.00799,24],["Claude Haiku 4.5 · Claude Code",0.001795,24],["Claude Opus 5.5 · Claude Code",0.008003,24],["GPT-6.1 Sol (medium) · Codex CLI",0.003025,16],["GPT-6.1 Sol (high) · Codex CLI",0.002905,16],["Claude Sonnet 5.5 · Claude Code",0.003672,24]]}],["Input (prompt, cache priced)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.02091,24],["Claude Opus 5.5 (high) · Claude Code",0.007407,24],["Claude Haiku 4.5 · Claude Code",0.00451,24],["Claude Opus 5.5 · Claude Code",0.00771,24],["GPT-6.1 Sol (medium) · Codex CLI",0.020339,16],["GPT-6.1 Sol (high) · Codex CLI",0.008117,16],["Claude Sonnet 5.5 · Claude Code",0.004012,24]]}]]},"note":"Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["thinking-token-bill",{"id":"thinking-bill-short-vs-hard","title":"Reasoning share on short tasks vs hard tasks (calculation)","subtitle":"Median call per configuration; five short tasks and eight hard tasks","kind":"grouped-bar","unit":"percent","yLabel":"Reasoning share of output tokens (median call, %)","series":[{"name":"Eight hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16]]}},{"name":"Five short tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",90.19,73.1,97.59,15],["Claude Sonnet 5.5 · Claude Code",0,0,72.75,15],["Claude Opus 5.5 · Claude Code",0,0,93.33,15],["Claude Opus 5.5 (high) · Claude Code",43.59,0,93.33,15],["Claude Fable 5.1 · Claude Code",0,0,74.01,15],["GPT-6.1 Sol (medium) · Codex CLI",41.05,0,71.43,15],["GPT-6.1 Sol (high) · Codex CLI",58.06,0,75.76,15]]}}],"note":"Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.","whisker":"minmax","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["llm-speed-anatomy",{"id":"speed-anatomy-first-text","title":"Time to first text: a 250-line answer, six models","subtitle":"Median of 4 calls per model; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds to first text","series":[{"name":"Time to first text","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",4,2.84,6.38,4],["Claude Sonnet 5.5 · Claude Code",1.96,0.88,4.09,4],["Claude Opus 5.5 · Claude Code",1.97,1.7,2.35,4],["Claude Fable 5.1 · Claude Code",4.43,2.27,4.64,4],["GPT-6.1 Sol (low) · Codex CLI",3.52,2.75,4.42,4],["GPT-6 Luna (low) · Codex CLI",3.3,3.19,3.47,4]]}}],"note":"Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-output-speed","title":"Output speed after the first text: visible tokens per second (calculation)","subtitle":"Median of 4 calls per model; whiskers = slowest and fastest call","kind":"dot-range","unit":"tokens","yLabel":"Visible tokens per second","series":[{"name":"Visible tokens per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",153.2,152.6,216.1,4],["Claude Sonnet 5.5 · Claude Code",231.7,230.3,233,4],["Claude Opus 5.5 · Claude Code",155.5,154.6,156.4,4],["Claude Fable 5.1 · Claude Code",122.6,120.9,131.4,4],["GPT-6.1 Sol (low) · Codex CLI",79.6,71.6,80.5,4],["GPT-6 Luna (low) · Codex CLI",129.1,55.5,259.1,4]]}}],"note":"Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-chars-per-second","title":"Output speed in characters per second after the first text (calculation)","subtitle":"Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call","kind":"dot-range","unit":"count","yLabel":"Characters per second","series":[{"name":"Characters per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",547,546,548,3],["Claude Sonnet 5.5 · Claude Code",517,513,519,4],["Claude Opus 5.5 · Claude Code",347,345,349,4],["Claude Fable 5.1 · Claude Code",273,270,293,4],["GPT-6.1 Sol (low) · Codex CLI",323,291,327,4],["GPT-6 Luna (low) · Codex CLI",524,225,1052,4]]}}],"note":"Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-prompt-size","title":"Time to first text as the prompt grows","subtitle":"Median of 3 calls per size; whiskers = fastest and slowest call","kind":"line","unit":"seconds","xLabel":"Prompt-size target (approximate Haiku tokens; calibration calculation)","yLabel":"Seconds to first text","series":{"$k":["name","points"],"$r":[["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.93,1.85,2.04,3],["16k",2.27,2.22,2.47,3],["64k",2.78,2.45,2.89,3]]}],["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.45,1.23,1.72,3],["16k",1.78,1.64,2.11,3],["64k",3.07,1.38,3.61,3]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.51,1.46,2.01,3],["16k",1.74,1.7,2.97,3],["64k",1.79,1.72,3.72,3]]}],["GPT-6.1 Sol (low) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["1k",3.36,3.36,4.75,3],["16k",4.02,3.3,4.28,3],["64k",3.93,3.42,4.38,3]]}]]},"note":"Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-total-by-size","title":"Total time per call by prompt size","subtitle":"Median of 3 calls per bar; whiskers = fastest and slowest call","kind":"grouped-bar","unit":"seconds","yLabel":"Seconds, whole call","series":{"$k":["name","points"],"$r":[["1k prompt",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",2.34,2.22,2.46,3],["Claude Sonnet 5.5 · Claude Code",1.78,1.57,2.12,3],["Claude Opus 5.5 · Claude Code",1.83,1.82,2.41,3],["GPT-6.1 Sol (low) · Codex CLI",3.43,3.43,4.92,3]]}],["16k prompt",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",2.79,2.58,2.84,3],["Claude Sonnet 5.5 · Claude Code",2.1,1.98,2.48,3],["Claude Opus 5.5 · Claude Code",2.36,2.11,3.4,3],["GPT-6.1 Sol (low) · Codex CLI",4.14,3.96,4.68,3]]}],["64k prompt",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",3.13,2.84,3.28,3],["Claude Sonnet 5.5 · Claude Code",3.44,1.74,4.38,3],["Claude Opus 5.5 · Claude Code",2.35,2.26,4.29,3],["GPT-6.1 Sol (low) · Codex CLI",3.96,3.47,4.44,3]]}]]},"note":"Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.","whisker":"minmax","sourceIds":["agent-speed-anatomy"]}],["llm-speed-anatomy",{"id":"speed-anatomy-lookup-correct","title":"Exact lookup answers at the 1k, 16k and 64k prompt-size targets","subtitle":"All sizes together per model; whiskers = 95% Wilson intervals","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Lookups answered exactly","series":[{"name":"Exact answer","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,0.7009,1,9],["Claude Sonnet 5.5 · Claude Code",1,0.7009,1,9],["Claude Opus 5.5 · Claude Code",0.5556,0.2667,0.8112,9],["GPT-6.1 Sol (low) · Codex CLI",1,0.7009,1,9]]}}],"note":"Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.","whisker":"ci95","sourceIds":["agent-speed-anatomy"]}],["haiku-retry-or-escalate",{"id":"retry-escalate-call-cost-by-task","title":"List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation)","subtitle":"Median of each task’s calls, default effort, Claude Code","kind":"grouped-bar","unit":"usd","yLabel":"USD per call","series":[{"name":"Claude Haiku 4.5 · Claude Code","points":{"$k":["label","value","n"],"$r":[["Interval merge fix",0.01789,3],["DST day length",0.0293,3],["CSV parser",0.02919,3],["Event-loop order",0.03708,3],["Room schedule",0.03578,3],["SemVer regex",0.04217,3],["Money refactor",0.02039,3],["SQLite report query",0.03648,3]]}},{"name":"Claude Sonnet 5.5 · Claude Code","points":{"$k":["label","value","n"],"$r":[["Interval merge fix",0.00557,3],["DST day length",0.02532,3],["CSV parser",0.01464,3],["Event-loop order",0.01588,3],["Room schedule",0.01243,3],["SemVer regex",0.00514,3],["Money refactor",0.00961,3],["SQLite report query",0.01719,3]]}}],"note":"Calculation: reported tokens × list price for each call, then the median per task; the calls ran on a flat subscription. Haiku’s list price per token is lower. Its median calls cost more and contained more output tokens. This does not isolate the effect of effort or thinking. 3 calls per task and configuration; the table shows each call-cost range, not a confidence interval.","sourceIds":["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-pass-rate","title":"Pass rate on 4 harder tasks","subtitle":"Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Passed","whisker":"ci95","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.5,0.2538,0.7462,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}}],"note":"Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row.\n\nCounted calls are new calls.","sourceIds":["agent-harder-tasks"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-tool-attempts","title":"Calls that tried a tool although tools were off","subtitle":"Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse","kind":"dot-range","unit":"rate","yLabel":"Calls with a tool attempt","polarity":"none","whisker":"ci95","series":[{"name":"Tool attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0,0,0.1936,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.3125,0.1416,0.556,16],["Claude Haiku 4.5 · Claude Code",0.0833,0.0149,0.3539,12]]}}],"note":"Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator.\n\nIt is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.","sourceIds":["agent-harder-tasks"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-pass-by-task","title":"Strict pass rate by task","subtitle":"One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Strict pass rate","whisker":"ci95","series":{"$k":["name","points"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",0.75,0.3006,0.9544,4],["Sudoku, 22 givens",0.25,0.0456,0.6994,4],["6x6 Skyscrapers",1,0.5101,1,4],["Seeded shuffle output",0.75,0.3006,0.9544,4]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",1,0.4385,1,3],["Sudoku, 22 givens",0,0,0.5615,3],["6x6 Skyscrapers",0.3333,0.0615,0.7923,3],["Seeded shuffle output",0.3333,0.0615,0.7923,3]]}],["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",1,0.5101,1,4],["Sudoku, 22 givens",0,0,0.4899,4],["6x6 Skyscrapers",0,0,0.4899,4],["Seeded shuffle output",0.5,0.15,0.85,4]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",0,0,0.5615,3],["Sudoku, 22 givens",0,0,0.5615,3],["6x6 Skyscrapers",0,0,0.5615,3],["Seeded shuffle output",0,0,0.5615,3]]}]]},"note":"Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.","sourceIds":["agent-harder-tasks"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-total-latency","title":"Total time per call on harder tasks","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","whisker":"minmax","series":[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",120.24,46.24,273.46,13,false],["Claude Opus 5.5 · Claude Code",80.34,3.82,279.5,9,false],["Claude Sonnet 5.5 · Claude Code",70.43,4.32,210.08,12,false],["Claude Haiku 4.5 · Claude Code",108.98,25.73,223.95,10,false]]}}],"note":"Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session.\n\nArena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.","sourceIds":["agent-harder-tasks"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-output-tokens","title":"Output tokens per call on harder tasks","subtitle":"Median per configuration; reasoning tokens as the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","polarity":"none","whisker":"minmax","series":[{"name":"Output tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",4994,2099,13413,13],["Claude Opus 5.5 · Claude Code",8420,279,40044,9],["Claude Sonnet 5.5 · Claude Code",9287,407,27921,12],["Claude Haiku 4.5 · Claude Code",12508,2965,26532,10]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",4971,2070,13372,13],["Claude Opus 5.5 · Claude Code",8352,21,9897,9],["Claude Sonnet 5.5 · Claude Code",6557,63,27902,12],["Claude Haiku 4.5 · Claude Code",12483,2924,26510,10]]}}],"note":"Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured.\n\nClaude Code used an output-token cap setting of 16,000. Some reported totals exceeded it.\n\nCodex CLI had no cap. More tokens is not better or worse by itself.","sourceIds":["agent-harder-tasks"]}]]}