{"$k":["slug","chart"],"$r":[["swe-bench-verified",{"id":"swebench-same-instance-leaderboard","title":"Resolved rate on the same 33 SWE-bench Verified instances","subtitle":"Agent vs 11 public mini-SWE-agent v2 runs, one attempt each","kind":"dot-range","unit":"rate","yLabel":"Resolved","series":[{"name":"Resolved rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["GPT 5.2 (high)",0.8485,0.6908,0.9335,33,"\u0001"],["Gemini 3 Flash (high)",0.8182,0.6561,0.9139,33,"\u0001"],["GLM 5 (high)",0.7879,0.6225,0.8932,33,"\u0001"],["Agent (Sonnet 5.5, full pipeline)",0.7576,0.5898,0.8717,33,true],["Claude 4.5 Sonnet (high)",0.7576,0.5898,0.8717,33,"\u0001"],["Claude 4.5 Haiku (high)",0.7576,0.5898,0.8717,33,"\u0001"],["Claude 4.5 Opus (high)",0.7273,0.5578,0.8493,33,"\u0001"],["DeepSeek V3.2 (high)",0.7273,0.5578,0.8493,33,"\u0001"],["MiniMax M2.5 (high)",0.697,0.5266,0.8262,33,"\u0001"],["Claude 4.6 Opus",0.697,0.5266,0.8262,33,"\u0001"],["Kimi K2.5 (high)",0.697,0.5266,0.8262,33,"\u0001"],["GPT 5 mini",0.6364,0.4662,0.7781,33,"\u0001"]]}}],"note":"Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.","sourceIds":["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"]}],["swe-bench-verified",{"id":"swebench-cost-vs-resolved","title":"Cost per instance vs resolved rate","subtitle":"Same 33 instances. Panel = API list price; Agent = notional subscription estimate","kind":"scatter","unit":"rate","xLabel":"Mean model cost per instance (USD)","yLabel":"Resolved rate","series":[{"name":"Public panel (mini-SWE-agent v2)","points":{"$k":["label","x","value","n"],"$r":[["Claude 4.5 Opus (high)",0.861,0.7273,33],["Gemini 3 Flash (high)",0.357,0.8182,33],["MiniMax M2.5 (high)",0.075,0.697,33],["Claude 4.6 Opus",0.61,0.697,33],["GLM 5 (high)",0.525,0.7879,33],["GPT 5.2 (high)",0.533,0.8485,33],["Claude 4.5 Sonnet (high)",0.692,0.7576,33],["Kimi K2.5 (high)",0.179,0.697,33],["DeepSeek V3.2 (high)",0.463,0.7273,33],["Claude 4.5 Haiku (high)",0.363,0.7576,33],["GPT 5 mini",0.051,0.6364,33]]}},{"name":"Agent","points":[{"label":"Agent","x":2.807,"value":0.7576,"n":33,"highlight":true}]}],"note":"Agent's cost includes repository onboarding, planning, verification and review; it is a list-price estimate for subscription calls, not an invoice. Panel costs are published API costs.","sourceIds":["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","price-anthropic"]}],["swe-bench-verified",{"id":"swebench-cost-by-stage","title":"Where Agent's model spend goes","subtitle":"Share of notional model cost by stage, all 33 attempts","kind":"bar","unit":"usd","yLabel":"USD (notional)","series":[{"name":"Cost","points":{"$k":["label","value"],"$r":[["Act (edit and run)",44.05],["Research",20.56],["Verify",8.72],["Other",6.24],["Review",5.58],["Context compaction",5.41],["Onboarding notes",2.09]]}}],"note":"Total $92.64 over 33 attempts.","sourceIds":["agent-swebench-c1","agent-swebench-c2"]}],["swe-bench-verified",{"id":"swebench-views","title":"Every way to slice the run, with intervals","subtitle":"Agent resolved rate and 95% Wilson interval per declared view","kind":"dot-range","unit":"rate","yLabel":"Resolved","series":[{"name":"Agent","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Campaign 1: 25-instance sample",0.72,0.5242,0.8572,25,"\u0001"],["Campaign 2: 8 compiled-extension instances",0.875,0.5291,0.9776,8,"\u0001"],["Original seed draw of 25",0.76,0.5657,0.885,25,"\u0001"],["All 33 attempted",0.7576,0.5898,0.8717,33,true]]}},{"name":"Public panel mean, same instances","points":{"$k":["label","value","n"],"$r":[["Campaign 1: 25-instance sample",0.7091,25],["Campaign 2: 8 compiled-extension instances",0.8409,8],["Original seed draw of 25",0.7091,25],["All 33 attempted",0.741,33]]}}],"note":"The original draw and \"all 33\" mix two platform builds.","sourceIds":["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol"]}],["swe-bench-opus-vs-sonnet",{"id":"swebench-opus-sonnet-resolved","title":"Resolved on the same 3 SWE-bench Verified instances (interim)","subtitle":"One attempt per arm per instance, official grader · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","yLabel":"Resolved","viz":"IntervalDotPlot","series":[{"name":"Resolved","points":[{"label":"Claude Opus 5.5 (Agent, new build)","value":0.6667,"lo":0.2077,"hi":0.9385,"n":3},{"label":"Claude Sonnet 5.5 (Agent, older builds)","value":0.3333,"lo":0.0615,"hi":0.7923,"n":3}]}],"note":"Interim: 3 of 8 declared pairs are graded; 5 were never started; no resumed attempts are included here. With n = 3 the intervals span most of the axis, so this chart supports no ranking. The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.","sourceIds":["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2"]}],["blind-review-head-to-head",{"id":"blind-review-first-vs-latest","title":"First attempt vs latest attempt","subtitle":"Share of tasks where the blind panel preferred the AI change","kind":"dot-range","unit":"rate","yLabel":"Tasks preferred","series":[{"name":"AI preferred","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["First scored attempt",0.5,0.2538,0.7462,12,"\u0001"],["Latest attempt",0.75,0.4677,0.9111,12,true],["Public OSS tasks, latest",0.8,0.3755,0.9638,5,"\u0001"],["Private tasks, latest",0.7143,0.3589,0.9178,7,"\u0001"]]}}],"note":"Later attempts had the earlier attempts’ lessons, review replays and, on some tasks, operator answers. They are not independent first tries.","sourceIds":["agent-blind-review"]}],["model-head-to-head",{"id":"h2h-pass-rate","title":"Pass rate on five validated tasks","subtitle":"Every call counts; failures and timeouts are non-passes","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Pass rate","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Fable 5.1 · Claude Code",1,0.7961,1,15],["Claude Sonnet 5.5 · Claude Code",0.8,0.5481,0.9295,15],["Claude Opus 5.5 (high) · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 (low) · Claude Code",1,0.7961,1,15],["Claude Haiku 4.5 · Claude Code",1,0.7961,1,15],["GPT-6.1 Sol (high) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (low) · Codex CLI",1,0.7225,1,10]]}}],"note":"Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-total-latency","title":"Total time per call","subtitle":"Median per configuration; whiskers = fastest and slowest call","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.94,1.41,9.83,15,true],["Claude Sonnet 5.5 · Claude Code",2.31,2.17,7.73,15,true],["Claude Opus 5.5 (high) · Claude Code",2.71,2.45,11.78,15,true],["Claude Opus 5.5 · Claude Code",2.75,2.47,8.91,15,true],["Claude Opus 5.5 (low) · Claude Code",2.83,2.35,6.62,15,true],["Claude Haiku 4.5 · Claude Code",4.43,3.16,23.57,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.6,4.05,19.52,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.65,4.1,25.46,15,false],["GPT-6.1 Sol (low) · Codex CLI",6.26,4.65,10.47,10,false]]}}],"note":"One host, one network, one day. Whiskers are a range, not a confidence interval.","sourceIds":["agent-provider-h2h"]}],["model-head-to-head",{"id":"h2h-list-price-per-call","title":"List-price cost per call (calculation)","subtitle":"Reported tokens × list price; the calls ran on subscriptions","kind":"dot-range","unit":"usd","yLabel":"USD per call","series":[{"name":"Cost per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",0.00987,0.0049,0.05843,15,true],["Claude Sonnet 5.5 · Claude Code",0.0036,0.00342,0.01021,15,true],["Claude Opus 5.5 (high) · Claude Code",0.00694,0.00592,0.02708,15,true],["Claude Opus 5.5 · Claude Code",0.00688,0.00592,0.02226,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00688,0.00582,0.01793,15,true],["Claude Haiku 4.5 · Claude Code",0.00566,0.00513,0.01804,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01047,0.0066,0.02812,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01018,0.0054,0.02686,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00769,0.00742,0.02649,10,false]]}}],"note":"Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.","sourceIds":["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]}],["hard-model-head-to-head",{"id":"hard-h2h-pass-rate","title":"Pass rate on eight hard tasks","subtitle":"Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.4583,0.2789,0.6493,24]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.6667,0.4671,0.8203,24]]}}],"note":"Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-outcomes","title":"What happened on every call","subtitle":"Counts per configuration: strict passes, format misses and wrong answers","kind":"stacked-bar","unit":"count","yLabel":"Calls","series":{"$k":["name","points"],"$r":[["Strict pass",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",24,24],["Claude Opus 5.5 · Claude Code",24,24],["Claude Opus 5.5 (high) · Claude Code",24,24],["GPT-6.1 Sol (medium) · Codex CLI",16,16],["Claude Fable 5.1 · Claude Code",24,24],["GPT-6.1 Sol (high) · Codex CLI",16,16],["Claude Haiku 4.5 · Claude Code",11,24]]}],["Format miss (correct answer, wrong format)",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0,24],["Claude Opus 5.5 · Claude Code",0,24],["Claude Opus 5.5 (high) · Claude Code",0,24],["GPT-6.1 Sol (medium) · Codex CLI",0,16],["Claude Fable 5.1 · Claude Code",0,24],["GPT-6.1 Sol (high) · Codex CLI",0,16],["Claude Haiku 4.5 · Claude Code",5,24]]}],["Wrong answer",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0,24],["Claude Opus 5.5 · Claude Code",0,24],["Claude Opus 5.5 (high) · Claude Code",0,24],["GPT-6.1 Sol (medium) · Codex CLI",0,16],["Claude Fable 5.1 · Claude Code",0,24],["GPT-6.1 Sol (high) · Codex CLI",0,16],["Claude Haiku 4.5 · Claude Code",8,24]]}]]},"note":"A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-output-tokens","title":"Output tokens per call on hard tasks","subtitle":"Median per configuration; reasoning tokens as the CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1050,24],["Claude Opus 5.5 · Claude Code",945,24],["Claude Opus 5.5 (high) · Claude Code",1052,24],["GPT-6.1 Sol (medium) · Codex CLI",335,16],["Claude Fable 5.1 · Claude Code",1366,24],["GPT-6.1 Sol (high) · Codex CLI",436,16],["Claude Haiku 4.5 · Claude Code",5064,24]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",585,24],["Claude Opus 5.5 · Claude Code",529,24],["Claude Opus 5.5 (high) · Claude Code",614,24],["GPT-6.1 Sol (medium) · Codex CLI",150,16],["Claude Fable 5.1 · Claude Code",889,24],["GPT-6.1 Sol (high) · Codex CLI",225,16],["Claude Haiku 4.5 · Claude Code",4556,24]]}}],"note":"Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.","sourceIds":["agent-provider-h2h-hard"]}],["hard-model-head-to-head",{"id":"hard-h2h-cost-per-pass","title":"List-price cost per strict pass on hard tasks (calculation)","subtitle":"All calls in a configuration, failures and format misses included, divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.01435,24,true],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16,false],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16,false],["Claude Opus 5.5 · Claude Code",0.02824,24,false],["Claude Opus 5.5 (high) · Claude Code",0.03337,24,false],["Claude Haiku 4.5 · Claude Code",0.0672,24,false],["Claude Fable 5.1 · Claude Code",0.09331,24,false]]}}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.","sourceIds":["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]}],["hard-model-head-to-head",{"id":"hard-h2h-frontier","title":"Quality vs cost frontier on hard tasks","subtitle":"Strict pass rate against list-price cost per strict pass","kind":"scatter","unit":"rate","xLabel":"USD per strict pass (list-price calculation)","yLabel":"Strict pass rate","series":[{"name":"Claude Code","points":{"$k":["label","x","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.01435,1,24,true],["Claude Opus 5.5 · Claude Code",0.02824,1,24,false],["Claude Opus 5.5 (high) · Claude Code",0.03337,1,24,false],["Claude Fable 5.1 · Claude Code",0.09331,1,24,false],["Claude Haiku 4.5 · Claude Code",0.0672,0.4583,24,false]]}},{"name":"Codex CLI","points":[{"label":"GPT-6.1 Sol (medium) · Codex CLI","x":0.02564,"value":1,"n":16,"highlight":false},{"label":"GPT-6.1 Sol (high) · Codex CLI","x":0.01514,"value":1,"n":16,"highlight":false}]}],"note":"Upper-left is better. Highlighted points are on the frontier: no other configuration passes at least as often for at most the same cost per pass. Frontier: Claude Sonnet 5.5 · Claude Code. Costs are calculations from tokens. Pass rates with their 95% intervals are in the pass-rate chart.","sourceIds":["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]}],["coding-agents-head-to-head",{"id":"coding-agents-wall-time","title":"Time per coding session","subtitle":"Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)","kind":"dot-range","unit":"seconds","whisker":"minmax","yLabel":"Seconds","viz":"LatencyLanes","series":[{"name":"Wall time per session","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",23.1,18.7,44.5,12,true],["Claude Opus 5.5 · Claude Code",56.9,29.8,185.8,12,"\u0001"],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",113.4,78.5,221.9,12,"\u0001"]]}}],"note":"CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.","sourceIds":["agent-coding-agents"]}],["effort-ladder",{"id":"effort-ladder-pass-rate","title":"Strict pass rate by effort on eight hard tasks","subtitle":"Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 · Claude Code",1,0.8064,1,16],["GPT-6.1 Sol (low) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16]]}}],"note":"Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.","whisker":"ci95","sourceIds":["agent-effort-ladder"]}],["caching-consistency",{"id":"caching-read-share-by-turn","title":"Share of input read from the cache, by turn in a session","subtitle":"Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each","kind":"line","unit":"rate","xLabel":"Turn in the session","yLabel":"Input tokens read from cache","series":{"$k":["name","points"],"$r":[["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["Turn 1",0.1868,3],["Turn 2",0.9924,3],["Turn 3",0.9896,3],["Turn 4",0.992,3],["Turn 5",0.9074,3]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["Turn 1",0.1869,3],["Turn 2",0.9924,3],["Turn 3",0.9866,3],["Turn 4",0.9893,3],["Turn 5",0.9031,3]]}],["GPT-6.1 Sol (medium) · Codex CLI",{"$k":["label","value","n"],"$r":[["Turn 1",0.5545,3],["Turn 2",0.9871,3],["Turn 3",0.9892,3],["Turn 4",0.9898,3],["Turn 5",0.9786,3]]}]]},"note":"Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.","sourceIds":["agent-caching-consistency"]}],["caching-consistency",{"id":"consistency-pass-rate","title":"Same prompt, 10 times: strict pass rate","subtitle":"One series per prompt; whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","yLabel":"Passed","series":{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0,0,0.2775,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["JSON object",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0.1,0.0179,0.4042,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["Code fix",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,0.7225,1,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}]]},"note":"Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.","whisker":"ci95","sourceIds":["agent-caching-consistency"]}],["caching-consistency",{"id":"consistency-distinct-answers","title":"Same prompt, 10 times: how many different answers","subtitle":"Distinct normalized answers over 10 repetitions (1 = the same answer every time)","kind":"grouped-bar","unit":"count","yLabel":"Distinct answers","series":{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,10],["Claude Sonnet 5.5 · Claude Code",1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,10]]}],["JSON object",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,10],["Claude Sonnet 5.5 · Claude Code",1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,10]]}],["Code fix",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 · Claude Code",6,10],["Claude Sonnet 5.5 · Claude Code",3,10],["GPT-6.1 Sol (medium) · Codex CLI",6,10]]}]]},"note":"Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.","sourceIds":["agent-caching-consistency"]}],["agent-memory",{"id":"memory-knowledge-class","title":"Where memory helps: what the repo shows vs what only the team knows","subtitle":"Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)","kind":"grouped-bar","unit":"rate","series":{"$k":["name","points"],"$r":[["Rules the code already shows",{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.95,0.863,0.9829,60],["/init CLAUDE.md",0.95,0.863,0.9829,60],["Curated, 11 lines",1,0.9398,1,60],["Raw notes, 60 lines",1,0.9398,1,60],["Dreamed notes",1,0.9398,1,60],["Handbook, 210 lines",1,0.9398,1,60],["Stop hook only",1,0.9398,1,60],["Curated + hook",1,0.9398,1,60]]}],["Rules the folders hint at",{"$k":["label","value","lo","hi","n"],"$r":[["No memory",1,0.8454,1,21],["/init CLAUDE.md",1,0.8454,1,21],["Curated, 11 lines",1,0.8454,1,21],["Raw notes, 60 lines",1,0.8454,1,21],["Dreamed notes",1,0.8454,1,21],["Handbook, 210 lines",1,0.8454,1,21],["Stop hook only",1,0.8454,1,21],["Curated + hook",1,0.8454,1,21]]}],["Team knowledge only",{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.4,0.1982,0.6425,15],["/init CLAUDE.md",0.6667,0.4171,0.8482,15],["Curated, 11 lines",1,0.7961,1,15],["Raw notes, 60 lines",1,0.7961,1,15],["Dreamed notes",1,0.7961,1,15],["Handbook, 210 lines",1,0.7961,1,15],["Stop hook only",0.6667,0.4171,0.8482,15],["Curated + hook",1,0.7961,1,15]]}]]},"note":"Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.","sourceIds":["agent-memory-study"]}],["agent-memory",{"id":"memory-team-knowledge-by-model","title":"Team knowledge followed, Sonnet vs Haiku","subtitle":"Changelog rule and late-fee rate, pooled · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.4,0.1982,0.6425,15],["/init CLAUDE.md",0.6667,0.4171,0.8482,15],["Curated, 11 lines",1,0.7961,1,15],["Raw notes, 60 lines",1,0.7961,1,15],["Dreamed notes",1,0.7961,1,15],["Handbook, 210 lines",1,0.7961,1,15],["Stop hook only",0.6667,0.4171,0.8482,15],["Curated + hook",1,0.7961,1,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0,0,0.2775,10],["/init CLAUDE.md",0.1,0.0179,0.4042,10],["Curated, 11 lines",0.8,0.4902,0.9433,10],["Raw notes, 60 lines",0.6,0.3127,0.8318,10],["Dreamed notes",0.8,0.4902,0.9433,10],["Handbook, 210 lines",0.3,0.1078,0.6032,10],["Stop hook only",0.8,0.4902,0.9433,10],["Curated + hook",1,0.7225,1,10]]}}],"note":"Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.","sourceIds":["agent-memory-study"]}],["agent-memory",{"id":"memory-late-fee","title":"\"Charge our standard late fee\": what Sonnet 5.5 did","subtitle":"The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%","kind":"stacked-bar","unit":"count","series":{"$k":["name","points"],"$r":[["Used the current rate (1.25%)",{"$k":["label","value","n","highlight"],"$r":[["No memory",0,3,"\u0001"],["/init CLAUDE.md",0,3,"\u0001"],["Curated, 11 lines",3,3,true],["Raw notes, 60 lines",3,3,true],["Dreamed notes",3,3,true],["Handbook, 210 lines",3,3,true],["Stop hook only",0,3,"\u0001"],["Curated + hook",3,3,true]]}],["Asked for the rate, wrote no code",{"$k":["label","value","n"],"$r":[["No memory",3,3],["/init CLAUDE.md",2,3],["Curated, 11 lines",0,3],["Raw notes, 60 lines",0,3],["Dreamed notes",0,3],["Handbook, 210 lines",0,3],["Stop hook only",2,3],["Curated + hook",0,3]]}],["Guessed another rate",{"$k":["label","value","n"],"$r":[["No memory",0,3],["/init CLAUDE.md",1,3],["Curated, 11 lines",0,3],["Raw notes, 60 lines",0,3],["Dreamed notes",0,3],["Handbook, 210 lines",0,3],["Stop hook only",1,3],["Curated + hook",0,3]]}]]},"note":"Late-fee task, 3 sessions per condition. \"Asked\" means the agent searched the repository, found no rate and stopped with a question instead of code.","sourceIds":["agent-memory-study"]}],["agent-memory",{"id":"memory-broken-test-command","title":"A stale README command: who still ran it?","polarity":"lower","subtitle":"Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.8,0.5481,0.9295,15],["/init CLAUDE.md",0.8667,0.6212,0.9626,15],["Curated, 11 lines",0,0,0.2039,15],["Raw notes, 60 lines",0,0,0.2039,15],["Dreamed notes",0,0,0.2039,15],["Handbook, 210 lines",0,0,0.2039,15],["Stop hook only",0.6,0.3575,0.8018,15],["Curated + hook",0,0,0.2039,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",1,0.7225,1,10],["/init CLAUDE.md",1,0.7225,1,10],["Curated, 11 lines",0,0,0.2775,10],["Raw notes, 60 lines",1,0.7225,1,10],["Dreamed notes",0.1,0.0179,0.4042,10],["Handbook, 210 lines",0,0,0.2775,10],["Stop hook only",1,0.7225,1,10],["Curated + hook",0,0,0.2775,10]]}}],"note":"The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old \"run npm test\" note and a later correction. The /init file repeats the README.","sourceIds":["agent-memory-study"]}],["agent-memory",{"id":"memory-cost-per-full-pass","title":"List-price cost per fully correct result (calculation)","subtitle":"Sum of the CLI's cost estimates for a condition, divided by its full passes","kind":"grouped-bar","unit":"usd","series":[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","n"],"$r":[["No memory",0.1386,9],["/init CLAUDE.md",0.1359,9],["Curated, 11 lines",0.0818,15],["Raw notes, 60 lines",0.1009,14],["Dreamed notes",0.0896,15],["Handbook, 210 lines",0.1006,15],["Stop hook only",0.1278,12],["Curated + hook",0.0843,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","n"],"$r":[["No memory",0.3786,2],["/init CLAUDE.md",0.4317,2],["Curated, 11 lines",0.1095,7],["Raw notes, 60 lines",0.1255,6],["Dreamed notes",0.1153,7],["Handbook, 210 lines",0.2615,3],["Stop hook only",0.1419,8],["Curated + hook",0.096,9]]}}],"note":"Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.","sourceIds":["agent-memory-study"]}],["system-one-arena",{"id":"arena-accuracy","title":"Who decides right? Accuracy on 1,000+ checkable decisions","subtitle":"Share of graded items answered correctly, first presentation · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13",0.7681,0.7416,0.7927,1048,true],["Clef 27B",0.6994,0.671,0.7264,1048,"\u0001"],["Clef-Flash 9B",0.6517,0.6224,0.68,1048,"\u0001"],["Kev 4B",0.6403,0.6107,0.6688,1048,"\u0001"],["lev 4B",0.5859,0.5558,0.6153,1048,"\u0001"],["Laya",0.2739,0.2477,0.3016,1048,"\u0001"],["Julia-1",0.2586,0.233,0.2859,1048,"\u0001"]]}}],"note":"1048 graded items in five suites. An error, a timeout or a label outside the option set counts as wrong. Two models differ clearly only where the paired McNemar test says so (table below).","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-elo","title":"Tournament rating across every game","subtitle":"Elo from every round-robin game (K 32, start 1000, averaged over 200 game orders) · 95% bootstrap intervals","kind":"dot-range","unit":"score","whisker":"ci95","series":[{"name":"Elo","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Perfect player",1554,1498,1609,336,"\u0001"],["Jev 1.13",1040,1002,1082,336,true],["lev 4B",941,894,986,336,"\u0001"],["Clef 27B",941,901,984,336,"\u0001"],["Kev 4B",936,892,974,336,"\u0001"],["Clef-Flash 9B",936,890,982,336,"\u0001"],["Random player",896,858,948,336,"\u0001"],["Julia-1",879,831,921,336,"\u0001"],["Laya",877,840,921,336,"\u0001"]]}}],"note":"Tic-tac-toe, Connect Four, Nim, Dots and Boxes and Pong; every game counts once. The random and perfect players anchor the scale.","sourceIds":["system-one-arena"]}],["routing-jev-vs-llm",{"id":"routing-exact-decisions","title":"Typed routing decisions answered exactly right","subtitle":"Share of asked cases where every scored question was acceptable","kind":"dot-range","unit":"rate","yLabel":"Exact","whisker":"ci95","series":[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8984,0.8191,0.9497,82,true],["Claude Haiku 4.5",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5",0.939,0.8651,0.9737,82,false]]}}],"note":"Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-economics-scenarios","title":"Thought experiment: recorded agent work under different model mixes","subtitle":"50 benchmark runs, 2,362 model calls, repriced","kind":"bar","unit":"usd","yLabel":"USD for all runs","series":[{"name":"Repriced cost","points":{"$k":["label","value","highlight"],"$r":[["all Fable 5.1",369.58,false],["all Opus 5.5",170.92,false],["policy (Opus strong, Haiku ancillary)",161.62,false],["all Sonnet 5.5",108.54,true],["split (Sonnet main line, Haiku ancillary)",105.53,false],["all Haiku 4.5",54.27,false]]}}],"note":"Calculation, not a run: every recorded call ran on Sonnet 5.5 with routing off. Same tokens on every model; a different model or mix would take a different path.","sourceIds":["agent-routing","calc-repricing","price-anthropic"]}],["routing-overhead",{"id":"router-overhead-decision-latency","title":"Time to make one routing decision","subtitle":"Median; whiskers = median to 95th percentile","kind":"dot-range","unit":"ms","yLabel":"Time per decision","whisker":"p50-p95","series":[{"name":"Decision time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",0.00142,0.00142,0.00233,20000,true],["Jev 1.13 (TypeSafe)",136.5,136.5,195.7,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",2597,2597,4298,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",12543,12543,34481,82,"\u0001"]]}}],"note":"The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.","sourceIds":["agent-routing-overhead","agent-routing","agent-jev-live"]}],["routing-overhead",{"id":"cli-startup-tax","title":"CLI start-up tax on a one-word answer","subtitle":"Median of 5 runs; whiskers = fastest and slowest run","kind":"dot-range","unit":"ms","yLabel":"Time to output","whisker":"minmax","series":{"$k":["name","points"],"$r":[["First output event",[{"label":"Claude Code · Claude Haiku 4.5","value":563,"lo":519,"hi":726,"n":5},{"label":"Codex CLI (default model)","value":489,"lo":354,"hi":1304,"n":5}]],["First model output",[{"label":"Claude Code · Claude Haiku 4.5","value":1461,"lo":1206,"hi":2308,"n":5},{"label":"Codex CLI (default model)","value":5059,"lo":4391,"hi":5478,"n":5}]],["Total wall time",[{"label":"Claude Code · Claude Haiku 4.5","value":2529,"lo":2273,"hi":3382,"n":5},{"label":"Codex CLI (default model)","value":5999,"lo":5367,"hi":6506,"n":5}]]]},"note":"Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.","sourceIds":["agent-routing-overhead"]}],["routing-overhead",{"id":"cli-startup-input-tokens","title":"Input tokens a CLI sends for a one-word answer","subtitle":"Per call, mostly the CLI’s own system prompt and tool definitions","kind":"bar","unit":"tokens","yLabel":"Input tokens per call","series":[{"name":"Input tokens per call","points":[{"label":"Claude Code · Claude Haiku 4.5","value":6761,"n":5},{"label":"Codex CLI (default model)","value":17051,"n":5}]}],"note":"Claude Code sums its disjoint input, cache-read and cache-write fields. Codex CLI reports 17,051 input tokens, 13,184 of them read from the cache. The prompt itself is a few tokens.","sourceIds":["agent-routing-overhead"]}],["inference-provider-index",{"id":"provider-index-spread","title":"How much more the priciest provider charges than the cheapest","subtitle":"Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider","kind":"bar","unit":"ratio","yLabel":"Most expensive ÷ cheapest","series":[{"name":"Price spread","points":{"$k":["label","value","n","highlight"],"$r":[["DeepSeek V4 Flash 0423",12.57,15,true],["DeepSeek V4 Pro 0423",11.21,15,true],["gpt-oss-120b",6.92,20,true],["Llama 3.3 70B Instruct",6.71,10,true],["GLM 5.3",4.69,32,true],["Kimi K3",1.78,19,false],["Llama 4 Maverick",1.69,3,false],["Claude Haiku 4.5",1,4,false],["Claude Sonnet 5",1,5,false],["Claude Sonnet 5.5",1,5,false],["Claude Opus 4.8",1,5,false],["Claude Opus 5",1,5,false],["Claude Opus 5.5",1,5,false],["Claude Fable 5.1",1,4,false],["GPT-6 Sol",1,2,false],["GPT-6 Luna",1,2,false],["GPT-6 Astra",1,2,false],["GPT-5.5",1,2,false],["Gemini 3.8 Flash",1,2,false],["Gemini 3.5 Flash",1,2,false],["Gemini 3.5 Flash Lite",1,2,false],["Gemini 3.1 Pro Preview",1,2,false]]}}],"note":"A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.","sourceIds":["openrouter-api-snapshot"]}],["inference-provider-index",{"id":"gateway-markup-vs-first-party","title":"OpenRouter markup over the first-party list price","subtitle":"Per-token markup, and the markup after the Standard credit-purchase fee (calculation)","kind":"grouped-bar","unit":"percent","yLabel":"Markup (%)","series":{"$k":["name","points"],"$r":[["Per-token markup (input)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0],["Claude Sonnet 5",0],["Claude Sonnet 5.5",0],["Claude Opus 4.8",0],["Claude Opus 5",0],["Claude Opus 5.5",0],["Claude Fable 5.1",0],["GPT-6 Luna",0],["Gemini 3.8 Flash",0],["Gemini 3.5 Flash",0]]}],["Per-token markup (output)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0],["Claude Sonnet 5",0],["Claude Sonnet 5.5",0],["Claude Opus 4.8",0],["Claude Opus 5",0],["Claude Opus 5.5",0],["Claude Fable 5.1",0],["GPT-6 Luna",0],["Gemini 3.8 Flash",0],["Gemini 3.5 Flash",0]]}],["With the 5.5% card credit fee (input)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",5.5],["Claude Sonnet 5",5.5],["Claude Sonnet 5.5",5.5],["Claude Opus 4.8",5.5],["Claude Opus 5",5.5],["Claude Opus 5.5",5.5],["Claude Fable 5.1",5.5],["GPT-6 Luna",5.5],["Gemini 3.8 Flash",5.5],["Gemini 3.5 Flash",5.5]]}]]},"note":"A calculation: OpenRouter list price ÷ first-party list price − 1, then × (1 + 5.5%) for credits bought by card on the Standard plan (purchases large enough that the minimum fee does not apply). Fees as published on 2026-10-06; first-party prices from the vendors’ list prices.","sourceIds":["openrouter-api-snapshot","openrouter-fees","price-anthropic","price-openai","price-google"]}],["inference-provider-index",{"id":"provider-prices-claude-sonnet-5-5","title":"Claude Sonnet 5.5: price per million tokens by provider","subtitle":"Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06","kind":"grouped-bar","unit":"usd","yLabel":"USD per million tokens","series":{"$k":["name","points"],"$r":[["Input",{"$k":["label","value"],"$r":[["Amazon Bedrock",2],["Anthropic",2],["Azure",2],["Claude Platform on AWS",2],["Google Vertex",2]]}],["Output",{"$k":["label","value"],"$r":[["Amazon Bedrock",10],["Anthropic",10],["Azure",10],["Claude Platform on AWS",10],["Google Vertex",10]]}],["Cache read",{"$k":["label","value"],"$r":[["Amazon Bedrock",0.2],["Anthropic",0.2],["Azure",0.2],["Claude Platform on AWS",0.2],["Google Vertex",0.2]]}]]},"note":"Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.","factContext":"Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","sourceIds":["openrouter-api-snapshot"]}],["inference-provider-index",{"id":"provider-prices-gpt-oss-120b","title":"gpt-oss-120b: price per million tokens by provider","subtitle":"Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06","kind":"grouped-bar","unit":"usd","yLabel":"USD per million tokens","series":{"$k":["name","points"],"$r":[["Input",{"$k":["label","value"],"$r":[["CoreWeave (fp4)",0.03],["DekaLLM (bf16)",0.03],["DeepInfra (bf16)",0.037],["AkashML (bf16)",0.037],["Mancer 2 (fp8)",0.045],["Crusoe (bf16)",0.05],["Novita (fp4)",0.05],["DigitalOcean",0.06],["Google Vertex",0.09],["BaseTen (fp4)",0.1],["Amazon Bedrock",0.15],["Groq",0.15],["Nebius (fp4)",0.15],["Phala",0.15],["SiliconFlow (fp8)",0.15],["Together",0.15],["Parasail (fp4)",0.1],["Mara",0.15],["SambaNova",0.14],["Cerebras (fp16)",0.35]]}],["Output",{"$k":["label","value"],"$r":[["CoreWeave (fp4)",0.17],["DekaLLM (bf16)",0.18],["DeepInfra (bf16)",0.17],["AkashML (bf16)",0.187],["Mancer 2 (fp8)",0.25],["Crusoe (bf16)",0.25],["Novita (fp4)",0.25],["DigitalOcean",0.42],["Google Vertex",0.36],["BaseTen (fp4)",0.5],["Amazon Bedrock",0.6],["Groq",0.6],["Nebius (fp4)",0.6],["Phala",0.6],["SiliconFlow (fp8)",0.6],["Together",0.6],["Parasail (fp4)",0.75],["Mara",0.75],["SambaNova",0.95],["Cerebras (fp16)",0.75]]}],["Cache read",{"$k":["label","value"],"$r":[["CoreWeave (fp4)",0.03],["DekaLLM (bf16)",0.03],["AkashML (bf16)",0.037],["Crusoe (bf16)",0.05],["DigitalOcean",0.012],["BaseTen (fp4)",0.1],["Groq",0.075],["SiliconFlow (fp8)",0.075],["Parasail (fp4)",0.055],["Cerebras (fp16)",0.35]]}]]},"note":"Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.","factContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","sourceIds":["openrouter-api-snapshot"]}],["cost-thought-experiments",{"id":"repriced-cost-per-resolved","title":"Thought experiment: the same tokens at other list prices","subtitle":"Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price","kind":"bar","unit":"usd","yLabel":"USD per resolved instance","series":[{"name":"Repriced cost per resolved instance","points":{"$k":["label","value","highlight"],"$r":[["Claude Fable 5.1",12.852,false],["Claude Opus 5",8.723,false],["Claude Opus 5.5",5.753,false],["Claude Sonnet 5.5",3.489,true],["GPT-6.1 Sol",2.097,false],["Claude Haiku 4.5",1.745,false],["Gemini 3.x Flash",1.016,false],["Jev 1.13 (router)",0.042,false]]}}],"note":"Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.","sourceIds":["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic","price-google","price-openai","price-jev"]}],["cost-thought-experiments",{"id":"cost-per-resolved-agent-vs-panel","title":"Recorded cost per resolved instance: Agent vs the public panel","subtitle":"Same 33 SWE-bench Verified instances; all attempts in the numerator","kind":"bar","unit":"usd","yLabel":"USD per resolved instance","series":[{"name":"Cost per resolved instance","points":{"$k":["label","value","n","highlight"],"$r":[["Agent (notional)",3.706,25,true],["Claude 4.5 Opus (high)",1.184,24,"\u0001"],["Claude 4.5 Sonnet (high)",0.913,25,"\u0001"],["Claude 4.6 Opus",0.875,23,"\u0001"],["GLM 5 (high)",0.667,26,"\u0001"],["DeepSeek V3.2 (high)",0.637,24,"\u0001"],["GPT 5.2 (high)",0.628,28,"\u0001"],["Claude 4.5 Haiku (high)",0.479,25,"\u0001"],["Gemini 3 Flash (high)",0.436,27,"\u0001"],["Kimi K2.5 (high)",0.256,23,"\u0001"],["MiniMax M2.5 (high)",0.107,23,"\u0001"],["GPT 5 mini",0.08,21,"\u0001"]]}}],"note":"Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.","sourceIds":["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"]}],["cost-thought-experiments",{"id":"token-cost-mix","title":"Where the token dollars go","subtitle":"Recorded tokens at Sonnet 5.5 list price, by token kind","kind":"bar","unit":"usd","yLabel":"USD for all attempts","series":[{"name":"Cost","points":{"$k":["label","value"],"$r":[["Cache writes (1 h)",38.99],["Cache reads",30.62],["Output",17.61],["Uncached input",0.01]]}}],"note":"List-price calculation on recorded tokens: 153.1M cache reads, 9.7M cache writes, 1.8M output, 3.2k uncached input. An agent loop re-reads its context on every call, so cache reads dominate the token count; by price the largest part is cache writes (1 h).","sourceIds":["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic"]}],["cli-model-latency-tokens",{"id":"cli-vs-api-exact-reply-latency","title":"CLI vs API: time for a one-line answer","subtitle":"Matched cohort, fixed exact reply, 5 runs per configuration","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.97,0.65,1.5,5],["OpenAI API · GPT-6.1 Sol · low",1.02,0.96,1.87,5],["OpenAI API · GPT-6.1 Sol · high",1.52,1.35,2.23,5],["Codex CLI · GPT-6 Luna · none",3.19,2.88,3.83,5],["Codex CLI · GPT-6.1 Sol · low",4.18,3.86,4.53,5],["Codex CLI · GPT-6.1 Sol · high",4.19,3.81,4.69,5]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.82,0.51,1.37,5],["OpenAI API · GPT-6.1 Sol · low",0.87,0.84,1.74,5],["OpenAI API · GPT-6.1 Sol · high",1.34,1.26,2.12,5],["Codex CLI · GPT-6 Luna · none",2.79,2.46,3.42,5],["Codex CLI · GPT-6.1 Sol · low",3.75,3.44,4.1,5],["Codex CLI · GPT-6.1 Sol · high",3.79,3.37,4.3,5]]}}],"note":"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.","sourceIds":["agent-provider-explorer"]}],["cli-model-latency-tokens",{"id":"cli-vs-api-prompt-overhead","title":"Hidden prompt: input tokens for the same one-line request","subtitle":"Reported input tokens, matched cohort","kind":"bar","unit":"tokens","yLabel":"Input tokens per call","series":[{"name":"Input tokens","points":{"$k":["label","value","n"],"$r":[["OpenAI API · GPT-6 Luna · none",17,5],["OpenAI API · GPT-6.1 Sol · low",17,5],["OpenAI API · GPT-6.1 Sol · high",17,5],["Codex CLI · GPT-6 Luna · none",18859,5],["Codex CLI · GPT-6.1 Sol · low",19551,5],["Codex CLI · GPT-6.1 Sol · high",19555,5]]}}],"note":"The CLI wraps every request in its own system prompt and tool context; the bare API sends only the request. Part of the CLI input is served from cache.","sourceIds":["agent-provider-explorer"]}],["cli-model-latency-tokens",{"id":"scheduler-repair-claude-vs-codex","title":"Repairing a scheduler: Claude Code vs Codex vs API","subtitle":"Same prompt, medium effort, 296 behavioral checks, 3 runs each","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Code CLI · Sonnet 5.5 · medium",15,13.89,15.89,3,true],["Codex CLI · GPT-6.1 Sol · medium",61.16,59.9,69.51,3,"\u0001"],["OpenAI API · GPT-6.1 Sol · medium",17.32,16.28,18.61,3,"\u0001"]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Code CLI · Sonnet 5.5 · medium",7.55,6.77,7.63,3],["Codex CLI · GPT-6.1 Sol · medium",15.56,13.65,23.04,3],["OpenAI API · GPT-6.1 Sol · medium",7.46,6.68,9.05,3]]}}],"note":"All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.","sourceIds":["agent-provider-explorer"]}],["coding-calibration",{"id":"calibration-outcomes-by-slice","title":"Three real tasks, four platform builds","subtitle":"Tasks per slice: functional pass, verified delivery, pull request opened","kind":"grouped-bar","unit":"count","yLabel":"Tasks (of 3)","series":{"$k":["name","points"],"$r":[["Functional pass (offline gates)",{"$k":["label","value","n"],"$r":[["Baseline (capped)",2,3],["Fix wave 1 (capped)",2,3],["Uncapped, build f0ac3a8a",1,3],["Uncapped, build 236c0d3f",1,3]]}],["Verified delivery",{"$k":["label","value","n","highlight"],"$r":[["Baseline (capped)",0,3,false],["Fix wave 1 (capped)",0,3,false],["Uncapped, build f0ac3a8a",0,3,false],["Uncapped, build 236c0d3f",1,3,true]]}],["Pull request opened",{"$k":["label","value","n"],"$r":[["Baseline (capped)",1,3],["Fix wave 1 (capped)",2,3],["Uncapped, build f0ac3a8a",2,3],["Uncapped, build 236c0d3f",2,3]]}]]},"note":"One attempt per task per slice. The two capped slices stopped at 20 minutes or $5; the uncapped slices had no ceiling. Each slice is a different platform build, so a change is not a matched improvement.","sourceIds":["agent-coding-calibration"]}],["coding-calibration",{"id":"calibration-minutes-by-task","title":"Wall time per task, by slice","kind":"grouped-bar","unit":"minutes","yLabel":"Minutes","series":{"$k":["name","points"],"$r":[["Baseline (capped)",{"$k":["label","value","highlight"],"$r":[["fastify/session",9,false],["h3js/h3",11.8,false],["Kludex/uvicorn",17.9,false]]}],["Fix wave 1 (capped)",{"$k":["label","value","highlight"],"$r":[["fastify/session",15.5,false],["h3js/h3",13.5,false],["Kludex/uvicorn",20.3,false]]}],["Uncapped, build f0ac3a8a",{"$k":["label","value","highlight"],"$r":[["fastify/session",7.2,false],["h3js/h3",12.5,false],["Kludex/uvicorn",22.7,false]]}],["Uncapped, build 236c0d3f",{"$k":["label","value","highlight"],"$r":[["fastify/session",11.2,true],["h3js/h3",9.2,true],["Kludex/uvicorn",19.8,true]]}]]},"note":"Onboarding included. uvicorn needed more than the old 20 minute cap once the caps were removed.","sourceIds":["agent-coding-calibration"]}],["single-call-vs-agent-loop",{"id":"agent-loop-pass-rate","title":"Strict pass rate: single call vs agent loop on eight hard tasks","subtitle":"Same tasks and validators. Whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",0.4583,0.2789,0.6493,24],["Claude Haiku 4.5 (agent loop) · Claude Code",0.5417,0.3507,0.7211,24],["Claude Sonnet 5.5 (single call) · Claude Code",1,0.862,1,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",1,0.8064,1,16],["GPT-6 Luna (single call) · Codex CLI",0.625,0.3864,0.8152,16],["GPT-6 Luna (agent loop) · Codex CLI",0.8571,0.6006,0.9599,14]]}}],"note":"Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.","whisker":"ci95","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["single-call-vs-agent-loop",{"id":"agent-loop-tokens","title":"Tokens per attempt: single call vs agent loop","subtitle":"Median per configuration; whiskers = fewest and most","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens","series":[{"name":"Input tokens (cache reads included)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",3941,3879,4221,24],["Claude Haiku 4.5 (agent loop) · Claude Code",71691,41732,516306,24],["Claude Sonnet 5.5 (single call) · Claude Code",2281,2234,2669,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",9550,9398,33040,16],["GPT-6 Luna (single call) · Codex CLI",11582,11526,11818,16],["GPT-6 Luna (agent loop) · Codex CLI",15530,15391,39009,14]]}},{"name":"Output tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",5064,1899,9321,24],["Claude Haiku 4.5 (agent loop) · Claude Code",7912,2541,20654,24],["Claude Sonnet 5.5 (single call) · Claude Code",1050,176,3895,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",876,219,3243,16],["GPT-6 Luna (single call) · Codex CLI",345,36,634,16],["GPT-6 Luna (agent loop) · Codex CLI",480,143,858,14]]}}],"note":"Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.","whisker":"minmax","sourceIds":["agent-agent-loop","agent-provider-h2h-hard"]}],["haiku-thinking-on-off",{"id":"haiku-thinking-router-exact","title":"Haiku thinking study: typed routing decisions answered exactly right","subtitle":"Claude Haiku 4.5 and Claude Sonnet 5.5 (low effort); 82 decisions, the same cases for every arm","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Correct","series":[{"name":"Exact decisions (every scored question right)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0.8659,0.7755,0.9234,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5 (low) · Claude Code",0.939,0.8651,0.9737,82,false]]}},{"name":"Per-question accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0.9124,0.8642,0.9446,194,true],["Claude Haiku 4.5 (thinking on) · Claude Code",0.9433,0.9013,0.968,194,false],["Claude Sonnet 5.5 (low) · Claude Code",0.9742,0.9411,0.9889,194,false]]}}],"note":"Whiskers are 95% Wilson intervals. Questions within a decision are related; per-question intervals are descriptive, not an independent-question test. The thinking-on and Sonnet arms are the recorded 2026-10-05 routing run, reused, not rerun; the thinking-off arm ran later on another account. An unanswered question counts as wrong.","whisker":"ci95","factContext":"typed routing decisions, thinking on vs off","sourceIds":["agent-haiku-thinking","agent-routing"]}],["haiku-thinking-on-off",{"id":"haiku-thinking-router-tokens","title":"Haiku thinking study: thinking and visible output tokens per routing decision","subtitle":"Mean per decision, as the Claude Code CLI reports them","kind":"grouped-bar","unit":"tokens","yLabel":"Tokens per decision","series":[{"name":"Thinking tokens","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0,82],["Claude Haiku 4.5 (thinking on) · Claude Code",1101,82],["Claude Sonnet 5.5 (low) · Claude Code",2,82]]}},{"name":"Visible output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",366,82],["Claude Haiku 4.5 (thinking on) · Claude Code",318,82],["Claude Sonnet 5.5 (low) · Claude Code",105,82]]}}],"note":"Visible output = output tokens minus thinking tokens. It includes the structured answer the CLI asks for. Thinking tokens are counted by the CLI; their content is never captured. More tokens is not better or worse by itself.","factContext":"typed routing decisions, thinking on vs off","sourceIds":["agent-haiku-thinking","agent-routing"]}],["prompt-cache-break-even",{"id":"cache-break-even-reads","title":"Reuses before a cached prefix costs less, by model and write type (calculation)","subtitle":"The break-even point: reuses at which the cached and the uncached cost are equal, at list prices","kind":"grouped-bar","unit":"score","yLabel":"Reuses at break-even","series":{"$k":["name","points"],"$r":[["1-hour write (2× input), whole prefix new",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",1.11],["Claude Sonnet 5.5",1.11],["Claude Opus 5.5 (cache read $0.2 per M)",1.05],["Claude Opus 5.5 (cache read $0.4 per M)",1.11],["Claude Fable 5.1",1.03]]}],["1-hour write, pooled n = 6 session share, 19% already cached (as recorded)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0.72],["Claude Sonnet 5.5",0.72],["Claude Opus 5.5 (cache read $0.2 per M)",0.67],["Claude Opus 5.5 (cache read $0.4 per M)",0.72],["Claude Fable 5.1",0.65]]}],["5-minute write (1.25× input, an assumption)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0.28],["Claude Sonnet 5.5",0.28],["Claude Opus 5.5 (cache read $0.2 per M)",0.26],["Claude Opus 5.5 (cache read $0.4 per M)",0.28],["Claude Fable 5.1",0.26]]}]]},"note":"Calculation, not a run: break-even reuses = (write price − input price) ÷ (input price − cache-read price). A cached prefix costs less once the reuses pass that point. On a new prefix, a 1-hour write needs 2 reuses (the 3rd request). With 19% already cached it needs 1 reuse, and a 5-minute write needs 1 reuse. The 1.25 of the 5-minute write is an assumption. The caching study protocol states it as Anthropic’s published figure. No 5-minute write occurred in the recorded sessions. We recorded the 19% on Sonnet 5.5 and Opus 5.5 sessions. For Haiku 4.5 and Fable 5.1 it is a what-if. GPT-6.1 Sol and GPT-6 Luna list no write surcharge, so their break-even is 0 reuses and the chart leaves them out. The price list gives Opus 5.5 a cache read of $0.2 per million. Another table of the product lists $0.4. We did not check the vendor price, so the chart shows both.","sourceIds":["calc-cache-pricing","price-anthropic","price-openai"]}],["routing-holdout",{"id":"routing-holdout-exact","title":"Unseen routing decisions answered exactly right","subtitle":"Share of the 56 holdout cases where every scored question was acceptable","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Exact","series":[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"note":"Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["thinking-token-bill",{"id":"thinking-bill-share","title":"Reasoning share of output tokens per call on hard tasks (calculation)","subtitle":"Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)","kind":"bar","unit":"percent","yLabel":"Reasoning share of output tokens (%)","series":[{"name":"Median call","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16]]}}],"note":"Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.","whisker":"minmax","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}],["haiku-retry-or-escalate",{"id":"retry-escalate-haiku-by-task","title":"Strict pass rate of Haiku 4.5 on each hard task","subtitle":"Claude Haiku 4.5 · Claude Code, default effort. 3 calls per task","kind":"bar","unit":"rate","polarity":"higher","yLabel":"Strict passes","series":[{"name":"Strict pass rate (n = 3 per task)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.4385,1,3],["DST day length",0.3333,0.0615,0.7923,3],["CSV parser",0.6667,0.2077,0.9385,3],["Event-loop order",0,0,0.5615,3],["Room schedule",0,0,0.5615,3],["SemVer regex",1,0.4385,1,3],["Money refactor",0.6667,0.2077,0.9385,3],["SQLite report query",0,0,0.5615,3]]}}],"note":"Measured. Whiskers are 95% Wilson intervals on only 3 calls per task, so they are wide: 0 of 3 allows a true rate up to 56%, and 3 of 3 allows one as low as 44%. A format miss (a right answer in a code fence or with prose) is not a pass. The retry policies take each task’s rate from this chart.","whisker":"ci95","sourceIds":["agent-provider-h2h-hard"]}],["harder-tasks-head-to-head",{"id":"harder-h2h-pass-rate","title":"Pass rate on 4 harder tasks","subtitle":"Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Passed","whisker":"ci95","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.5,0.2538,0.7462,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}}],"note":"Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row.\n\nCounted calls are new calls.","sourceIds":["agent-harder-tasks"]}],["voice-agent-latency-budget",{"id":"latency-budget-fit","title":"How many of the steps fit each assumed budget (calculation)","subtitle":"Count of 14 measured steps; n = steps","kind":"grouped-bar","unit":"count","yLabel":"Steps that fit (of 14)","series":[{"name":"Fits at the median","points":{"$k":["label","value","n"],"$r":[["300 ms budget",3,14],["800 ms budget",3,14],["1,500 ms budget",8,14]]}},{"name":"Fits at the slow end","points":{"$k":["label","value","n"],"$r":[["300 ms budget",3,14],["800 ms budget",3,14],["1,500 ms budget",4,14]]}}],"note":"A calculation on measured times, with budgets of 300 ms, 800 ms and 1,500 ms (assumption: a budget a voice team might set). A step fits when its median, or its slow end, is at or below the budget. The slow end is the 95th percentile for steps with 30 or more runs and the slowest run for the rest, which is stricter. The steps differ in route and in what they time, so read the count as a summary of the table below, not as a ranking.","sourceIds":["calc-latency-budget","agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]}],["routing-at-scale",{"id":"routing-at-scale-daily-cost","title":"Daily cost of routing at 10,000 to 10 million decisions a day (calculation)","subtitle":"Decisions a day × cost per decision, USD per day, every decision routed. The rule-based policy costs $0 and is not on the log axis","kind":"grouped-bar","unit":"usd","yLabel":"USD per day","series":{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","n"],"$r":[["10,000 a day",0.337,82],["100,000 a day",3.37,82],["1 million a day",33.7,82],["10 million a day",337,82]]}],["Claude Sonnet 5.5 (low) · Claude Code",{"$k":["label","value","n"],"$r":[["10,000 a day",73.24,82],["100,000 a day",732.4,82],["1 million a day",7324,82],["10 million a day",73240,82]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","n"],"$r":[["10,000 a day",89.24,82],["100,000 a day",892.4,82],["1 million a day",8924,82],["10 million a day",89240,82]]}]]},"note":"This is a calculation, not a run. Daily cost = decisions a day × cost per decision. Cost per 1,000 decisions: Jev $0.0337 (calculation from provider-reported run cost, n = 82), Sonnet 5.5 $7.32, Haiku 4.5 $8.92. The Claude figures use one-hour cache-write rates and list price × the tokens the CLI reported (Sonnet n = 82, Haiku n = 82). \n\nThe rule-based policy makes no model call, so it costs $0 at every volume and has no bar: a log axis cannot show zero. The chart prices model calls only, not servers, retries or the work itself.","sourceIds":["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev"]}]]}