{"$k":["slug","title","seoTitle","description","question","answer","date","updated","tags","method","sourceIds","stats","charts","tables","hero"],"$r":[["swe-bench-verified","Agent on SWE-bench Verified vs 11 public models","SWE-bench Verified: Agent vs GPT, Claude and Gemini","Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.","How does Agent, a full worker pipeline on one model, do on SWE-bench Verified next to public single-model runs on the very same instances?","Agent resolved 25 of 33 attempted instances (75.8%, 95% interval 59% to 87%).","2026-10-05","2026-10-05",["swe-bench","coding-agents","leaderboard","cost","claude-sonnet"],[],["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol","price-anthropic"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["agent-rate-33","Agent resolved, all 33 attempted instances",0.7576,"rate","76% (25/33)",33,[0.5898,0.8717],"Both campaigns, one attempt each, failures and empty patches included."],["agent-rate-c1","Agent resolved, campaign 1 sample of 25",0.72,"rate","72% (18/25)",25,[0.5242,0.8572],"\u0001"],["agent-rate-original25","Agent resolved, original seed draw of 25 (no replacements)",0.76,"rate","76% (19/25)",25,[0.5657,0.885],"Mixes two platform builds."],["panel-mean-33","Public panel mean on the same 33 instances",0.741,"rate","74.1%",33,"\u0001","Mean of 11 public mini-SWE-agent v2 runs."],["cost-per-attempt","Agent model cost per attempt (notional)",2.81,"usd","$2.81",33,"\u0001","Subscription calls priced at list price; onboarding and failed attempts included."],["cost-per-resolved","Agent model cost per resolved instance (notional)",3.71,"usd","$3.71",25,"\u0001","\u0001"],["median-minutes","Median worker time per attempt",9.6,"minutes","9.6 min",33,"\u0001","Range 1.6 to 54.1 min."],["calls-per-attempt","Model calls per attempt",49.5,"calls","49",33,"\u0001","\u0001"]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["swebench-same-instance-leaderboard","Resolved rate on the same 33 SWE-bench Verified instances","Agent vs 11 public mini-SWE-agent v2 runs, one attempt each","dot-range","rate","Resolved",[{"name":"Resolved rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["GPT 5.2 (high)",0.8485,0.6908,0.9335,33,"\u0001"],["Gemini 3 Flash (high)",0.8182,0.6561,0.9139,33,"\u0001"],["GLM 5 (high)",0.7879,0.6225,0.8932,33,"\u0001"],["Agent (Sonnet 5.5, full pipeline)",0.7576,0.5898,0.8717,33,true],["Claude 4.5 Sonnet (high)",0.7576,0.5898,0.8717,33,"\u0001"],["Claude 4.5 Haiku (high)",0.7576,0.5898,0.8717,33,"\u0001"],["Claude 4.5 Opus (high)",0.7273,0.5578,0.8493,33,"\u0001"],["DeepSeek V3.2 (high)",0.7273,0.5578,0.8493,33,"\u0001"],["MiniMax M2.5 (high)",0.697,0.5266,0.8262,33,"\u0001"],["Claude 4.6 Opus",0.697,0.5266,0.8262,33,"\u0001"],["Kimi K2.5 (high)",0.697,0.5266,0.8262,33,"\u0001"],["GPT 5 mini",0.6364,0.4662,0.7781,33,"\u0001"]]}}],"Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"]],["swebench-by-difficulty-band","Resolved rate by difficulty band","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol"]],["swebench-cost-vs-resolved","Cost per instance vs resolved rate","\u0001","scatter","\u0001","\u0001",[],"\u0001",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","price-anthropic"]],["swebench-model-calls","Model calls per instance","\u0001","bar","\u0001","\u0001",[],"\u0001",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"]],["swebench-cost-by-stage","Where Agent's model spend goes","\u0001","bar","\u0001","\u0001",[],"\u0001",["agent-swebench-c1","agent-swebench-c2"]],["swebench-views","Every way to slice the run, with intervals","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol"]]]},[{"id":"swebench-per-instance","columns":[],"rows":[]},{"id":"swebench-panel","columns":[],"rows":[]}],"\u0001"],["swe-bench-opus-vs-sonnet","Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)","Opus 5.5 vs Sonnet 5.5 on SWE-bench Verified (interim)","Interim: 3 of 8 paired SWE-bench Verified instances graded. Opus 5.5 resolved 2, Sonnet 5.5 1 (McNemar p = 1.0). Cost 2.6×, a list-price calculation.","Does Agent resolve more SWE-bench Verified instances with Claude Opus 5.5 as its brain than with Claude Sonnet 5.5, and at what cost and time?","Interim, not a final result: 3 of 8 declared pairs are graded.","2026-10-06","2026-10-06",["swe-bench","claude-opus","claude-sonnet","agent-harness","interim"],[],["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["swebench-opus-resolved","Claude Opus 5.5 in Agent: resolved, interim (3 of 8 pairs graded)",0.6667,"rate","67% (2/3)",3,[0.2077,0.9385],"\u0001"],["swebench-sonnet-resolved","Claude Sonnet 5.5 in Agent: resolved on the same 3 instances",0.3333,"rate","33% (1/3)",3,[0.0615,0.7923],"\u0001"],["swebench-opus-sonnet-mcnemar","Exact McNemar p on the graded pairs",1,"score","p = 1.0",3,"\u0001","1 pair where only Opus resolved, 0 pairs where only Sonnet did. The test reads only these; p = 1.0 is no evidence of a difference."],["swebench-opus-sonnet-graded","Declared pairs graded so far",3,"count","3 of 8",8,"\u0001","5 not started: the usage gate stopped the campaign before django__django-11885. The recorded usage-reset date was 2026-10-09; no resumed attempts are included here. Every started attempt counts."],["swebench-opus-sonnet-cost-ratio","Opus vs Sonnet list-price cost on the same instances (calculation)",2.63,"ratio","2.6×",3,"\u0001","$22.76 vs $8.64 over the 3 graded instances: notional cost from the platform price table, not an invoice."],["swebench-opus-sonnet-minutes-ratio","Opus vs Sonnet worker minutes on the same instances (ratio of totals)",1.9,"ratio","1.9×",3,"\u0001","55.3 vs 29.1 worker minutes."]]},{"$k":["id","title","subtitle","kind","unit","whisker","yLabel","viz","series","note","sourceIds"],"$r":[["swebench-opus-sonnet-resolved","Resolved on the same 3 SWE-bench Verified instances (interim)","One attempt per arm per instance, official grader · 95% Wilson intervals","dot-range","rate","ci95","Resolved","IntervalDotPlot",[{"name":"Resolved","points":[{"label":"Claude Opus 5.5 (Agent, new build)","value":0.6667,"lo":0.2077,"hi":0.9385,"n":3},{"label":"Claude Sonnet 5.5 (Agent, older builds)","value":0.3333,"lo":0.0615,"hi":0.7923,"n":3}]}],"Interim: 3 of 8 declared pairs are graded; 5 were never started; no resumed attempts are included here. With n = 3 the intervals span most of the axis, so this chart supports no ranking. The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2"]],["swebench-opus-sonnet-cost-per-attempt","List-price cost per attempt (calculation)","\u0001","bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"]],["swebench-opus-sonnet-cost-by-instance","List-price cost per instance (calculation)","\u0001","grouped-bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"]],["swebench-opus-sonnet-minutes","Worker time per attempt","\u0001","dot-range","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2"]],["swebench-opus-sonnet-stage-cost","Where the cost goes: pipeline stages (calculation)","\u0001","grouped-bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"]]]},{"$k":["id","columns","rows"],"$r":[["swebench-opus-sonnet-pairs",[],[]],["swebench-opus-sonnet-graded",[],[]],["swebench-opus-sonnet-instances",[],[]]]},{"statIds":["swebench-opus-resolved","swebench-sonnet-resolved"],"testStatId":"swebench-opus-sonnet-mcnemar"}],["blind-review-head-to-head","AI pull requests vs merged human pull requests, judged blind","AI vs human pull requests: a blind multi-model review","A blind panel of Claude and GPT critics preferred Agent's change over the merged human change on 9 of 12 real tasks. Votes, scores, caveats.","When critics cannot see which change came from a person, do they prefer the AI worker’s pull request or the one the maintainers merged?","On the latest attempt per task, the blind panel preferred the AI change on 9 of 12 tasks (75%, 95% interval 47% to 91%).","2026-09-28","2026-10-05",["code-review","ai-vs-human","llm-as-judge","pull-requests"],[],["agent-blind-review"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["ai-preferred-latest","Tasks where the panel preferred the AI change (latest attempt)",0.75,"rate","75% (9/12)",12,[0.4677,0.9111],"Latest attempts come after earlier attempts on the same task; see caveats."],["ai-preferred-first","Tasks where the panel preferred the AI change (first scored attempt)",0.5,"rate","50% (6/12)",12,[0.2538,0.7462],"\u0001"],["ai-preferred-all-pairs","All scored pairs where the panel preferred the AI change",0.6,"rate","60% (12/20)",20,[0.3866,0.7812],"\u0001"],["ai-preferred-public","Public OSS tasks, latest attempt",0.8,"rate","80% (4/5)",5,[0.3755,0.9638],"\u0001"],["verdicts-ai","Single critic verdicts that preferred the AI change",0.6894,"rate","69% (91/132)",132,[0.606,0.762],"\u0001"],["review-spend","Notional spend: worker runs + critic panel",1310.5,"usd","$979 + $332",20,"\u0001","List-price estimates of subscription calls over all scored pairs."],["position-bias-flags","Pairs flagged for position bias",4,"count","4 of 20",20,"\u0001","A critic model flipped its preference when the two changes swapped places."]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["blind-review-first-vs-latest","First attempt vs latest attempt","Share of tasks where the blind panel preferred the AI change","dot-range","rate","Tasks preferred",[{"name":"AI preferred","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["First scored attempt",0.5,0.2538,0.7462,12,"\u0001"],["Latest attempt",0.75,0.4677,0.9111,12,true],["Public OSS tasks, latest",0.8,0.3755,0.9638,5,"\u0001"],["Private tasks, latest",0.7143,0.3589,0.9178,7,"\u0001"]]}}],"Later attempts had the earlier attempts’ lessons, review replays and, on some tasks, operator answers. They are not independent first tries.",["agent-blind-review"]],["blind-review-votes-by-task","Blind panel votes per task (latest attempt)","\u0001","stacked-bar","\u0001","\u0001",[],"\u0001",["agent-blind-review"]],["blind-review-dimension-scores","What the critics scored higher","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-blind-review"]],["blind-review-critic-agreement","Does the judge’s model family matter?","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-blind-review"]]]},[{"id":"blind-review-pairs","columns":[],"rows":[]}],"\u0001"],["model-head-to-head","Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head","Haiku vs Sonnet vs Opus vs Fable vs Codex: speed and tokens","130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.","On short tasks with strict validators, how do the Claude Code models and efforts compare with Codex on pass rate, speed and tokens?","127 of 130 calls passed (98%, 95% interval 93% to 99%), so on these short tasks pass rate barely separates the configurations (non-passes: Claude Sonnet 5.5 · Claude Code on Multi-step shift arithmetic ×3 (\"Output does not exactly match the expected text\")).","2026-10-05","2026-10-05",["head-to-head","claude-haiku","claude-sonnet","claude-opus","claude-fable","codex","latency"],[],["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"],{"$k":["id","label","value","unit","display","n","ci"],"$r":[["h2h-pass-all","Calls that passed their validator",0.9769,"rate","98% (127/130)",130,[0.9343,0.9921]],["h2h-configs","Configurations compared",9,"count","9","\u0001","\u0001"],["h2h-fastest","Fastest configuration (median total time)",1.94,"seconds","Claude Fable 5.1 · Claude Code: 1.9 s",15,"\u0001"],["h2h-slowest","Slowest configuration (median total time)",6.26,"seconds","GPT-6.1 Sol (low) · Codex CLI: 6.3 s",10,"\u0001"],["h2h-cheapest-per-pass","Lowest list-price cost per passing answer (calculation)",0.00624,"usd","Claude Sonnet 5.5 · Claude Code: $0.0062",15,"\u0001"],["h2h-codex-input-tokens","Median input tokens per call, Codex CLI vs Claude Code",12124,"tokens","12,124 vs 2,130",130,"\u0001"]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["h2h-pass-rate","Pass rate on five validated tasks","Every call counts; failures and timeouts are non-passes","dot-range","rate","Passed",[{"name":"Pass rate","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Fable 5.1 · Claude Code",1,0.7961,1,15],["Claude Sonnet 5.5 · Claude Code",0.8,0.5481,0.9295,15],["Claude Opus 5.5 (high) · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 (low) · Claude Code",1,0.7961,1,15],["Claude Haiku 4.5 · Claude Code",1,0.7961,1,15],["GPT-6.1 Sol (high) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (low) · Codex CLI",1,0.7225,1,10]]}}],"Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.",["agent-provider-h2h"]],["h2h-total-latency","Total time per call","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-provider-h2h"]],["h2h-first-useful-latency","Time to first useful output","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-provider-h2h"]],["h2h-input-tokens","Input tokens per call: what the CLI sends","\u0001","stacked-bar","\u0001","\u0001",[],"\u0001",["agent-provider-h2h"]],["h2h-output-tokens","Output tokens per call","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-provider-h2h"]],["h2h-list-price-per-call","List-price cost per call (calculation)","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]],["h2h-speed-vs-cost","Speed vs list-price cost","\u0001","scatter","\u0001","\u0001",[],"\u0001",["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]],["h2h-cost-per-pass","List-price cost per passing answer (calculation)","\u0001","bar","\u0001","\u0001",[],"\u0001",["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]],["h2h-frontier","Speed, cost and quality frontier","\u0001","scatter","\u0001","\u0001",[],"\u0001",["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]]]},[{"id":"h2h-pass-matrix","columns":[],"rows":[]}],"\u0001"],["hard-model-head-to-head","Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks","Hard tasks: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol","152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.","On 8 hard tasks with deterministic validators, does pass rate separate the Claude Code models and GPT-6.1 Sol through the Codex CLI, and what do speed, tokens and cost per pass add?","139 of 152 calls that reached a model passed strictly (91%). 6 of 7 configurations passed every call: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Fable 5.1 · Claude Code (24/24 each, 95% interval 86% to 100%) and GPT-6.1 Sol (medium) · Codex CLI, GPT-6.1 Sol (high) · Codex CLI (16/16 each, 95% interval 81% to 100%), so the hard set still has a ceiling for these models and pass rate does not separate them.","2026-10-06","2026-10-06",["head-to-head","hard-tasks","claude-haiku","claude-sonnet","claude-opus","claude-fable","gpt-6-1-sol","codex-cli","format-misses","latency"],[],["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["hard-h2h-pass-all","Calls that passed strictly (hard set)",0.9145,"rate","91% (139/152)",152,[0.8592,0.9493],"\u0001"],["hard-h2h-correct-all","Calls with a correct answer, format misses included (lenient reading)",0.9474,"rate","95% (144/152)",152,[0.8996,0.9731],"\u0001"],["hard-h2h-format-misses","Non-passes that were format misses, not wrong answers",5,"count","5 of 13 non-passes (8 wrong answers)",13,"\u0001","\u0001"],["hard-h2h-perfect-configs","Configurations that passed every call",6,"count","6 of 7 (4 at 24/24, 2 at 16/16)",7,"\u0001","\u0001"],["hard-h2h-fastest-perfect","Lowest observed median time among configurations that passed every call (separate batches)",7.75,"seconds","Claude Sonnet 5.5 · Claude Code: 7.7 s",24,"\u0001","The counted Claude and Codex batches ran hours apart on one host and network. Host load was not controlled; this does not isolate model speed."],["hard-h2h-cheapest-per-pass","Lowest list-price cost per strict pass (calculation)",0.01435,"usd","Claude Sonnet 5.5 · Claude Code: $0.0143",24,"\u0001","\u0001"],["hard-h2h-blocked","Attempts blocked before any model call (not scored)",30,"count","30 (Codex CLI; 0 model calls)",182,"\u0001","\u0001"]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["hard-h2h-pass-rate","Pass rate on eight hard tasks","Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","dot-range","rate","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.4583,0.2789,0.6493,24]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.6667,0.4671,0.8203,24]]}}],"Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.",["agent-provider-h2h-hard"]],["hard-h2h-outcomes","What happened on every call","\u0001","stacked-bar","\u0001","\u0001",[],"\u0001",["agent-provider-h2h-hard"]],["hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-provider-h2h-hard"]],["hard-h2h-first-useful-latency","Time to first useful output on hard tasks","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-provider-h2h-hard"]],["hard-h2h-output-tokens","Output tokens per call on hard tasks","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-provider-h2h-hard"]],["hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)","\u0001","bar","\u0001","\u0001",[],"\u0001",["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]],["hard-h2h-frontier","Quality vs cost frontier on hard tasks","\u0001","scatter","\u0001","\u0001",[],"\u0001",["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]]]},[{"id":"hard-h2h-pass-matrix","columns":[],"rows":[]},{"id":"hard-h2h-controls","columns":[],"rows":[]}],"\u0001"],["coding-agents-head-to-head","Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks","Claude Code vs Codex CLI: 6 coding tasks, hidden tests","36 graded sessions: Claude Code with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol. All passed every hidden test; time, tool calls and diffs differ.","On small real repository tasks graded by hidden tests, how do coding-agent CLIs compare when they run with their normal file and shell tools?","All 3 agents passed every hidden check in every session (12/12, 12/12, 12/12; 95% Wilson 76–100% each), so this task set cannot separate them on quality.","2026-10-06","2026-10-06",["claude-code","codex-cli","coding-agents","hidden-tests","head-to-head"],[],["agent-coding-agents","calc-repricing","price-anthropic","price-openai"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["coding-agents-pass-all","Sessions that passed every hidden check, all three agents",1,"rate","100% (36/36)",36,[0.9036,1],"The ceiling: 36 of 36 means this task set cannot rank the agents on quality."],["coding-agents-sessions","Graded sessions (6 tasks × 2 repetitions × 3 agents)",36,"count","36",36,"\u0001","0 timeouts, 0 errors or usage-limit stops, nothing retried. Gemini CLI not run: it asked for a browser login."],["coding-agents-median-time-fastest","Median time per session, Sonnet 5.5 in Claude Code",23.1,"seconds","23.1 s",12,"\u0001","Fastest 18.7 s, slowest 44.5 s (a range of 12 sessions, not an interval)."],["coding-agents-time-ratio","Median time, GPT-6.1 Sol in Codex CLI vs Sonnet 5.5 in Claude Code (ratio of medians)",4.92,"ratio","4.9×",12,"\u0001","113.4 s vs 23.1 s. The run ranges do not overlap. Codex time includes the work its standing instructions asked for (see the caveats)."],["coding-agents-tester-notes","Codex sessions that read the tester’s global notes (standing instructions)",12,"count","12 of 12",12,"\u0001","10 of 12 also wrote a WORKLOG.md and 9 reported a commit attempt; no task asked for either. No Claude Code session did any of this."],["coding-agents-outside-edits","Edits that landed outside the task repository",0,"count","0",36,"\u0001","3 attempted writes to a temp folder (1 Sonnet 5.5, 2 Opus 5.5) were refused by the Claude Code permission check."],["coding-agents-cost-total","List-price estimate of all 36 sessions (calculation, not an invoice)",4.87,"usd","$4.87",36,"\u0001","\u0001"]]},{"$k":["id","title","subtitle","kind","unit","whisker","yLabel","viz","series","note","sourceIds"],"$r":[["coding-agents-pass-rate","Coding sessions that passed every hidden check","A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals","dot-range","rate","ci95","Passed","IntervalDotPlot",[{"name":"Passed every hidden check","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.7575,1,12],["Claude Opus 5.5 · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",1,0.7575,1,12]]}}],"6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.",["agent-coding-agents"]],["coding-agents-wall-time","Time per coding session","\u0001","dot-range","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-coding-agents"]],["coding-agents-time-by-task","Time per coding task","\u0001","grouped-bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-coding-agents"]],["coding-agents-tool-calls","Tool calls per coding session","\u0001","dot-range","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-coding-agents"]],["coding-agents-token-mix","Tokens per session, as each CLI reports them","\u0001","stacked-bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-coding-agents"]],["coding-agents-diff-lines","Lines changed per session, by kind of file","\u0001","stacked-bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-coding-agents"]],["coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)","\u0001","bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-coding-agents","calc-repricing","price-anthropic","price-openai"]]]},[{"id":"coding-agents-cells","columns":[],"rows":[]},{"id":"coding-agents-tasks","columns":[],"rows":[]}],{"statIds":["coding-agents-time-ratio"]}],["effort-ladder","Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks","Effort ladder: does more AI effort buy quality?","176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.","On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate for Sonnet, Opus and GPT-6.1 Sol, and what does it cost in time, tokens and list price per pass?","Not on this set.","2026-10-06","2026-10-06",["effort","reasoning-effort","hard-tasks","claude-sonnet","claude-opus","gpt-6-1-sol","claude-code","codex-cli","latency","tokens"],[],["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"],{"$k":["id","label","value","unit","display","n","ci"],"$r":[["effort-ladder-pass-new","New effort-ladder calls that passed strictly",1,"rate","100% (96/96)",96,[0.9615,1]],["effort-ladder-pass-all","All effort-ladder calls that passed strictly, reference cells included",1,"rate","100% (176/176)",176,[0.9786,1]],["effort-ladder-cells","Configurations on the ladder (new + reference)",11,"count","11 (6 new, 5 reference)",176,"\u0001"],["effort-ladder-format-misses","Format misses and wrong answers on the ladder",0,"count","0 format misses, 0 wrong answers",176,"\u0001"]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","whisker","sourceIds"],"$r":[["effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks","Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals","dot-range","rate","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 · Claude Code",1,0.8064,1,16],["GPT-6.1 Sol (low) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16]]}}],"Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.","ci95",["agent-effort-ladder"]],["effort-ladder-total-latency","Total time per call by effort on hard tasks","\u0001","dot-range","\u0001","\u0001",[],"\u0001","\u0001",["agent-effort-ladder"]],["effort-ladder-time-by-effort","Median total time per call, by effort","\u0001","line","\u0001","\u0001",[],"\u0001","\u0001",["agent-effort-ladder"]],["effort-ladder-output-tokens","Output tokens per call by effort on hard tasks","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001","\u0001",["agent-effort-ladder"]],["effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)","\u0001","bar","\u0001","\u0001",[],"\u0001","\u0001",["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"]]]},[{"id":"effort-ladder-cells","columns":[],"rows":[]}],"\u0001"],["caching-consistency","Prompt caching and run-to-run consistency in Claude Code and Codex CLI","Prompt caching savings and LLM consistency, measured","135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.","When a CLI session reuses a fixed context, how much input comes from the cache, what does that save at list price, and does it change latency? When the same prompt runs 10 times, how much do the pass rate, the answer and the time vary?","Caching: inside one Claude Code session, turns 2-5 read 97% of their input from the cache on average (turn 1: 19%, the CLI's own prefix).","2026-10-06","2026-10-06",["prompt-caching","consistency","variance","claude-code","codex-cli","claude-haiku","claude-sonnet","claude-opus","gpt-6-1-sol","calculation"],[],["agent-caching-consistency","calc-cache-pricing","price-anthropic"],{"$k":["id","label","value","unit","display","note","n"],"$r":[["caching-saving-sonnet","List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)",0.4996,"rate","50% ($0.1350 vs $0.2698)","Calculation from recorded tokens and list prices; not a bill.","\u0001"],["caching-saving-opus","List-price saving from the cache over 15 turns, Claude Opus 5.5 · Claude Code (calculation)",0.5313,"rate","53% ($0.2551 vs $0.5442)","Calculation from recorded tokens and list prices; not a bill.","\u0001"],["caching-cross-session-reuse","Later sessions whose first turn read the ledger from an earlier session’s cache",0,"count","0 of 4","\u0001",4],["consistency-perfect-cells","Model-and-prompt cells that passed all 10 repetitions",7,"count","7 of 9","\u0001",9],["caching-consistency-calls","Calls in this study (every one counted)",135,"calls","135 (45 cache turns, 90 repeated prompts)","\u0001","\u0001"]]},{"$k":["id","title","subtitle","kind","unit","xLabel","yLabel","series","note","sourceIds"],"$r":[["caching-read-share-by-turn","Share of input read from the cache, by turn in a session","Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each","line","rate","Turn in the session","Input tokens read from cache",{"$k":["name","points"],"$r":[["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["Turn 1",0.1868,3],["Turn 2",0.9924,3],["Turn 3",0.9896,3],["Turn 4",0.992,3],["Turn 5",0.9074,3]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["Turn 1",0.1869,3],["Turn 2",0.9924,3],["Turn 3",0.9866,3],["Turn 4",0.9893,3],["Turn 5",0.9031,3]]}],["GPT-6.1 Sol (medium) · Codex CLI",{"$k":["label","value","n"],"$r":[["Turn 1",0.5545,3],["Turn 2",0.9871,3],["Turn 3",0.9892,3],["Turn 4",0.9898,3],["Turn 5",0.9786,3]]}]]},"Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.",["agent-caching-consistency"]],["caching-cost-with-without","List-price cost of 5-question sessions with and without the cache (calculation)","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-caching-consistency","calc-cache-pricing","price-anthropic"]],["caching-latency-first-vs-later","Time per turn: first turn vs later turns in a cached session","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-caching-consistency"]],["consistency-pass-rate","Same prompt, 10 times: strict pass rate","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-caching-consistency"]],["consistency-distinct-answers","Same prompt, 10 times: how many different answers","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-caching-consistency"]],["consistency-latency-spread","Same prompt, 10 times: time per call","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-caching-consistency"]]]},[{"id":"caching-turns","columns":[],"rows":[]},{"id":"consistency-cells","columns":[],"rows":[]}],"\u0001"],["agent-memory","Does memory help Claude Code? 8 kinds of agent memory, tested","Does CLAUDE.md help? Agent memory tested on Claude Code","200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.","Does project memory make Claude Code do better work, and which kind of memory?","Claude Sonnet 5.5 followed almost every rule it could see without any memory (95% of code rules, 100% of folder rules), but only 6/15 of the team-knowledge checks; with an 11-line curated file it passed 15/15.","2026-10-06","2026-10-06",["claude-code","agent-memory","claude-md","hooks","context-engineering"],[],["agent-memory-study"],{"$k":["id","label","value","unit","display","n","note","ci"],"$r":[["memory-sessions","Claude Code sessions, every one graded (120 Sonnet 5.5, 80 Haiku 4.5)",200,"count","200",200,"0 planned sessions not run.","\u0001"],["memory-team-knowledge-none","Team-knowledge checks passed with no memory (Sonnet 5.5)",0.4,"rate","40% (6/15)",15,"\u0001",[0.1982,0.6425]],["memory-team-knowledge-curated","Team-knowledge checks passed with an 11-line curated file (Sonnet 5.5)",1,"rate","100% (15/15)",15,"\u0001",[0.7961,1]],["memory-late-fee-known","Late fee right when any memory file held the rate (Sonnet 5.5)",1,"rate","15/15",15,"\u0001",[0.7961,1]],["memory-late-fee-asked","Without the rate, sessions that asked for it and wrote no code (Sonnet 5.5)",0.7778,"rate","7/9",9,"The other 2 guessed a rate and said it was a placeholder.","\u0001"],["memory-late-fee-flagged-sonnet","Without the rate: sessions whose final message said the rate was unknown or a guess (Sonnet 5.5)",1,"rate","9/9",9,"\u0001","\u0001"],["memory-late-fee-flagged-haiku","Without the rate: sessions whose final message said the rate was unknown or a guess (Haiku 4.5)",0,"rate","0/6",6,"Every other Haiku session invented a rate and reported the task done.","\u0001"],["memory-init-broken-command","Sessions with the /init file that ran the broken test command it copied from the README (Sonnet 5.5)",0.8667,"rate","13/15",15,"\u0001","\u0001"],["memory-hook-input-ratio","Median input tokens per session, Stop hook only vs no memory (calculation, Sonnet 5.5)",1.6,"ratio","1.6×",15,"\u0001","\u0001"],["memory-haiku-raw-broken-command","Haiku 4.5 with the raw notes: sessions that ran the stale test command",1,"rate","10/10",10,"\u0001","\u0001"],["memory-haiku-dreamed-broken-command","Haiku 4.5 with the dreamed notes: sessions that ran the stale test command",0.1,"rate","1/10",10,"\u0001","\u0001"],["memory-total-cost","List-price estimate of every session (calculation, not an invoice)",17.49,"usd","$17.49",200,"\u0001","\u0001"]]},{"$k":["id","title","subtitle","kind","unit","whisker","series","note","sourceIds","polarity"],"$r":[["memory-full-pass","Full pass rate by kind of memory","Hidden tests pass and every convention check passes · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["No memory",0.6,0.3575,0.8018,15,"\u0001"],["/init CLAUDE.md",0.6,0.3575,0.8018,15,"\u0001"],["Curated, 11 lines",1,0.7961,1,15,true],["Raw notes, 60 lines",0.9333,0.7018,0.9881,15,"\u0001"],["Dreamed notes",1,0.7961,1,15,"\u0001"],["Handbook, 210 lines",1,0.7961,1,15,"\u0001"],["Stop hook only",0.8,0.5481,0.9295,15,"\u0001"],["Curated + hook",1,0.7961,1,15,"\u0001"]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.2,0.0567,0.5098,10],["/init CLAUDE.md",0.2,0.0567,0.5098,10],["Curated, 11 lines",0.7,0.3968,0.8922,10],["Raw notes, 60 lines",0.6,0.3127,0.8318,10],["Dreamed notes",0.7,0.3968,0.8922,10],["Handbook, 210 lines",0.3,0.1078,0.6032,10],["Stop hook only",0.8,0.4902,0.9433,10],["Curated + hook",0.9,0.5958,0.9821,10]]}}],"Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.",["agent-memory-study"],"\u0001"],["memory-knowledge-class","Where memory helps: what the repo shows vs what only the team knows","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-team-knowledge-by-model","Team knowledge followed, Sonnet vs Haiku","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-late-fee","\"Charge our standard late fee\": what Sonnet 5.5 did","\u0001","stacked-bar","\u0001","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-late-fee-haiku","\"Charge our standard late fee\": what Haiku 4.5 did","\u0001","stacked-bar","\u0001","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-broken-test-command","A stale README command: who still ran it?","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-memory-study"],"lower"],["memory-input-tokens","What memory costs in context","\u0001","bar","\u0001","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-cost-per-full-pass","List-price cost per fully correct result (calculation)","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-wall-time","Time per session","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-dreaming-scorecard","Dreaming: what one consolidation pass kept and dropped","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"]]},{"$k":["id","columns","rows"],"$r":[["memory-by-condition",[],[]],["memory-by-condition-haiku",[],[]],["memory-raw-notes",[],[]],["memory-by-task",[],[]]]},{"statIds":["memory-team-knowledge-none","memory-team-knowledge-curated"]}],["system-one-arena","System One arena: Jev vs Clef and five open decision models, head to head","Jev vs Clef: 7 decision models on 1,085 decisions","Jev 1.13, Clef 27B, Clef-Flash and four small open decision models on 1,085 checkable decisions and in head-to-head games, Pong included.","How good are typed-decision models at decisions with a checkable answer, and which one wins when they play each other?","Jev 1.13 answered 76.8% of 1,048 graded decisions correctly, clearly ahead of every other model (exact McNemar p < 0.001 for each pair).","2026-10-06","2026-10-06",["jev","clef","decision-models","system-one","llama-cpp","open-models","routing"],[],["system-one-arena"],{"$k":["id","label","value","unit","display","n","ci"],"$r":[["arena-items","Decision items, written by agents that never called a model (1,048 graded, 37 consistency-only)",1085,"count","1,085",1085,"\u0001"],["arena-calls","Decision calls made and counted (exam, latency and repeat passes)",23247,"calls","23,247",23247,"\u0001"],["arena-best","Highest accuracy on the graded items: Jev 1.13",0.7681,"rate","76.8% (805/1048)",1048,[0.7416,0.7927]],["arena-best-open","Best open model: Clef 27B",0.6994,"rate","69.9% (733/1048)",1048,[0.671,0.7264]],["arena-smallest","Smallest model: Julia-1 (0.17 GB file)",0.2586,"rate","25.9% (271/1048)",1048,[0.233,0.2859]],["arena-jev-cost","Jev 1.13 list-price cost per 1,000 decisions (calculation from its reported input tokens)",0.03465,"usd","$0.0347",3081,"\u0001"],["arena-jev-repeat","Same question asked twice: Jev 1.13 gave the same answer 58/60 times; the local models every time",0.9667,"rate","58/60",60,"\u0001"],["arena-jev-latency","Jev 1.13 hosted API: median call time from Houston, network included (not comparable with a model on another machine)",137,"ms","137 ms",120,"\u0001"],["arena-chance","Expected accuracy of picking an option at random on the same graded items (calculation)",0.224,"rate","22.4%",1048,"\u0001"],["arena-rank-agreement","Rank agreement between our accuracy order and the independent Decision Index 0.2.1 order, same seven models (Spearman, calculation)",0.93,"score","0.93",7,"\u0001"],["arena-games","Games played to the end and counted (round robin, featured series and speed ladder)",1620,"count","1,620",1620,"\u0001"]]},{"$k":["id","title","subtitle","kind","unit","whisker","series","note","sourceIds","polarity"],"$r":[["arena-accuracy","Who decides right? Accuracy on 1,000+ checkable decisions","Share of graded items answered correctly, first presentation · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13",0.7681,0.7416,0.7927,1048,true],["Clef 27B",0.6994,0.671,0.7264,1048,"\u0001"],["Clef-Flash 9B",0.6517,0.6224,0.68,1048,"\u0001"],["Kev 4B",0.6403,0.6107,0.6688,1048,"\u0001"],["lev 4B",0.5859,0.5558,0.6153,1048,"\u0001"],["Laya",0.2739,0.2477,0.3016,1048,"\u0001"],["Julia-1",0.2586,0.233,0.2859,1048,"\u0001"]]}}],"1048 graded items in five suites. An error, a timeout or a label outside the option set counts as wrong. Two models differ clearly only where the paired McNemar test says so (table below).",["system-one-arena"],"\u0001"],["arena-by-suite","Where each model is strong","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"],["arena-latency-this-mac","Speed on this Mac: the six open models","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["system-one-arena"],"lower"],["arena-latency-same-gpu","Speed on one GPU: third-party numbers","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["system-one-arena"],"lower"],["arena-speed-accuracy-gpu","Open models: accuracy against speed on the same GPU","\u0001","scatter","\u0001","\u0001",[],"\u0001",["system-one-arena"],"none"],["arena-latency-gateway","Speed through one gateway: OpenRouter's own numbers","\u0001","bar","\u0001","\u0001",[],"\u0001",["system-one-arena"],"lower"],["arena-cost-same-provider","Price per 1,000 decisions at one provider's list prices","\u0001","bar","\u0001","\u0001",[],"\u0001",["system-one-arena"],"lower"],["arena-size-accuracy","Does a bigger file decide better?","\u0001","scatter","\u0001","\u0001",[],"\u0001",["system-one-arena"],"none"],["arena-robust","Same question, different presentation","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"],["arena-flips","Decisions that changed when only the presentation changed","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["system-one-arena"],"lower"],["arena-escape","Knowing when to say \"none of these\"","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["system-one-arena"],"none"],["arena-injection","Prompt injection: does text in the state hijack the decision?","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"],["arena-by-length","Short inputs vs long inputs","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"],["arena-confident-wrong","Wrong and sure of it","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["system-one-arena"],"lower"],["arena-elo","Tournament rating across every game","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"],["arena-wdl","Wins, draws and losses in the round robin","\u0001","stacked-bar","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"],["arena-perfect-moves","How often a model found the perfect move","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"],["arena-pong","Pong as deployed: who won","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"],["arena-pong-quality","Pong decision quality, ignoring time","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"],["arena-pong-ladder","How much is a millisecond worth in Pong?","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["system-one-arena"],"\u0001"]]},{"$k":["id","columns","rows"],"$r":[["arena-fairness",[],[]],["arena-models",[],[]],["arena-pairwise",[],[]],["arena-examples",[],[]],["arena-families",[],[]],["arena-industries",[],[]],["arena-games-table",[],[]],["arena-pong-table",[],[]],["arena-pong-featured",[],[]],["arena-pong-sides",[],[]],["arena-c4-players",[],[]],["arena-pong-rally",[],[]],["arena-c4-game",[],[]],["arena-leagues-table",[],[]]]},"\u0001"],["routing-jev-vs-llm","Jev vs Claude as a router: accuracy and cost","Jev router vs Claude Haiku and Sonnet: routing and cost","Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.","Should a small dedicated router or a general LLM make the platform’s typed routing decisions?","Exact decisions: Jev 1.13 (TypeSafe) 221 of 246 live calls (90%; 74, 73 and 74 of 82 per repeat; case-level interval 82% to 95%); Claude Haiku 4.5 73 of 82 (89%, 80% to 94%); Claude Sonnet 5.5 77 of 82 (94%, 87% to 97%).","2026-10-05","2026-10-06",["routing","jev","claude-haiku","claude-sonnet","model-routing","thought-experiment"],[],["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["exact-jev","Jev 1.13 (TypeSafe): exact decisions",0.8984,"rate","90% (74/82)",82,[0.8191,0.9497],"Live run: 3 repeats of the same 82 decisions. 221 of 246 calls were exact (89.8%; 74, 73 and 74 of 82 per repeat). The count is shown on the 82-decision scale (89.8% of 82 is 74), the scale of the interval: repeats of one decision are not independent, so the interval is taken at n = 82, not 246."],["exact-claude-haiku","Claude Haiku 4.5: exact decisions",0.8902,"rate","89% (73/82)",82,[0.8044,0.9412],"\u0001"],["exact-claude-sonnet","Claude Sonnet 5.5: exact decisions",0.939,"rate","94% (77/82)",82,[0.8651,0.9737],"\u0001"],["jev-cost-per-1000","Jev cost per 1,000 decisions",0.0337,"usd","$0.0337",246,"\u0001","Calculation: mean reported input tokens per decision × the published input price."],["economics-policy-vs-all-sonnet","Thought experiment: policy (Opus strong, Haiku ancillary) vs all Sonnet 5.5",1.489,"ratio","$161.62 (1.49x)",2362,"\u0001","Calculation on recorded tokens, not a run."],["economics-split-vs-all-sonnet","Thought experiment: split (Sonnet main line, Haiku ancillary) vs all Sonnet 5.5",0.9723,"ratio","$105.53 (0.97x)",2362,"\u0001","Calculation on recorded tokens, not a run."],["economics-cache-share","Cache-read share of recorded input (economics data)",0.9304,"rate","93.0%",2362,"\u0001","\u0001"]]},{"$k":["id","title","subtitle","kind","unit","yLabel","whisker","series","note","sourceIds"],"$r":[["routing-exact-decisions","Typed routing decisions answered exactly right","Share of asked cases where every scored question was acceptable","dot-range","rate","Exact","ci95",[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8984,0.8191,0.9497,82,true],["Claude Haiku 4.5",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5",0.939,0.8651,0.9737,82,false]]}}],"Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.",["agent-routing","agent-jev-live"]],["routing-key-accuracy","Per-question accuracy","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing","agent-jev-live"]],["routing-exact-by-decision","Exact rate by decision type","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing","agent-jev-live"]],["routing-cost-per-1000","Cost per 1,000 routing decisions","\u0001","bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"]],["routing-decision-latency","Time per routing decision","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing","agent-jev-live"]],["routing-economics-scenarios","Thought experiment: recorded agent work under different model mixes","\u0001","bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing","calc-repricing","price-anthropic"]],["routing-economics-by-stage","Thought experiment: repriced cost by pipeline stage","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing","calc-repricing","price-anthropic"]]]},{"$k":["id","columns","rows"],"$r":[["routing-head-to-head",[],[]],["routing-question-accuracy",[],[]],["routing-routers",[],[]],["routing-jev-live-repeats",[],[]]]},"\u0001"],["routing-overhead","Routing overhead: deterministic policy vs LLM routers vs Jev","Routing overhead: rules vs LLM routers vs Jev","How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.","What delay and what cost does each kind of router add before the real work of a call starts?","The deterministic routing policy decided in a median 1.42 µs (p95 2.33 µs, 20,000 decisions, $0).","2026-10-06","2026-10-06",["routing","latency","overhead","jev","llm-router","cli","calculation"],[],["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"],{"$k":["id","label","value","unit","display","n","note"],"$r":[["router-overhead-policy-p50","Deterministic routing policy: median decision time",0.00142,"ms","1.42 µs (p95 2.33 µs, p99 3.04 µs)",20000,"Timer resolution 0.041 µs; 469,409 decisions per second in a 200,000-call batch."],["router-overhead-speedup","Median LLM router call ÷ median policy decision",1828873,"ratio","about 1.8 million times",82,"\u0001"],["router-overhead-system-one-rule-arm","Recorded rule-based System One decision, record write included",1,"ms","1 ms median, 2 ms p95 (millisecond resolution)",419,"\u0001"],["router-overhead-decisions-per-task","Model calls per task (each one a routing decision)",49.5,"calls","49.5 median (13 to 73)",48,"\u0001"],["router-overhead-jev-p50","Jev 1.13 (TypeSafe): median decision time over the API",136.5,"ms","137 ms (p95 196 ms)",246,"Client wall time from one Mac over a home network, 246 calls in a 35-second window. The API reports no server time."],["router-overhead-sonnet-vs-jev","Median Claude Sonnet 5.5 call through the CLI ÷ median Jev call over the API",19,"ratio","about 19 times",82,"A calculation across two routes (CLI vs a direct API call from one Mac), not a model-against-model comparison."],["router-overhead-haiku-vs-jev","Median Claude Haiku 4.5 call through the CLI ÷ median Jev call over the API",91.9,"ratio","about 92 times",82,"A calculation across two routes (CLI vs a direct API call from one Mac), not a model-against-model comparison."],["cli-startup-claude-harness-ms","Claude Code time outside the model on a one-word answer",1690,"ms","1,690 ms median (1,533 to 1,811)",5,"\u0001"]]},{"$k":["id","title","subtitle","kind","unit","yLabel","whisker","series","note","sourceIds"],"$r":[["router-overhead-decision-latency","Time to make one routing decision","Median; whiskers = median to 95th percentile","dot-range","ms","Time per decision","p50-p95",[{"name":"Decision time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",0.00142,0.00142,0.00233,20000,true],["Jev 1.13 (TypeSafe)",136.5,136.5,195.7,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",2597,2597,4298,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",12543,12543,34481,82,"\u0001"]]}}],"The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.",["agent-routing-overhead","agent-routing","agent-jev-live"]],["router-overhead-cli-vs-model-time","Where an LLM router’s time goes: model vs CLI","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","agent-routing"]],["router-overhead-completed","Routing calls that returned a decision","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","agent-routing","agent-jev-live"]],["router-overhead-cost-reported","Cost per 1,000 routing decisions: no model call vs provider-reported","\u0001","bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","agent-routing","price-jev"]],["router-overhead-cost-list-price","Cost per 1,000 routing decisions for the model routers (calculation)","\u0001","bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]],["router-overhead-cost-per-1000-tasks","Added routing cost per 1,000 tasks (calculation)","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]],["router-overhead-delay-per-task","Added routing delay per task (calculation)","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]],["cli-startup-tax","CLI start-up tax on a one-word answer","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead"]],["cli-startup-input-tokens","Input tokens a CLI sends for a one-word answer","\u0001","bar","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead"]]]},[{"id":"router-overhead-status","columns":[],"rows":[]},{"id":"router-overhead-jev-live","columns":[],"rows":[]}],"\u0001"],["inference-provider-index","Inference provider index: 27 models, 52 providers","Inference provider prices: OpenRouter vs direct","Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.","For the same model, how much do inference providers differ in price, and what does a gateway such as OpenRouter add over the first-party list price?","OpenRouter’s public API listed 265 endpoints from 52 providers for 27 models on 2026-10-06.","2026-10-06","2026-10-06",["inference","providers","openrouter","pricing","gateway","open-weight","third-party-reported"],[],["openrouter-api-snapshot","openrouter-fees","price-anthropic","price-openai","price-google"],{"$k":["id","label","value","unit","display","n","note"],"$r":[["provider-index-endpoints","Provider endpoints in the snapshot",265,"count","265 endpoints, 52 providers, 27 models",27,"\u0001"],["provider-index-max-spread","Largest standard-tier price spread (DeepSeek V4 Flash 0423)",12.57,"ratio","12.6x (most expensive: Cloudflare; cheapest: StreamLake (fp8))",15,"Calculation on reported prices, blended 3:1."],["gateway-zero-markup-models","Models where OpenRouter’s per-token price equals the first-party list price",10,"count","10 of 10",10,"\u0001"],["gateway-credit-fee","Credit-purchase fee on OpenRouter’s Standard plan (third-party-reported)",5.5,"percent","5.5% ($0.80 minimum by card)","\u0001","\u0001"],["provider-index-latency-reported","Endpoints with a latency figure in the keyless API",0,"count","0 of 265",265,"\u0001"]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["provider-index-spread","How much more the priciest provider charges than the cheapest","Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider","bar","ratio","Most expensive ÷ cheapest",[{"name":"Price spread","points":{"$k":["label","value","n","highlight"],"$r":[["DeepSeek V4 Flash 0423",12.57,15,true],["DeepSeek V4 Pro 0423",11.21,15,true],["gpt-oss-120b",6.92,20,true],["Llama 3.3 70B Instruct",6.71,10,true],["GLM 5.3",4.69,32,true],["Kimi K3",1.78,19,false],["Llama 4 Maverick",1.69,3,false],["Claude Haiku 4.5",1,4,false],["Claude Sonnet 5",1,5,false],["Claude Sonnet 5.5",1,5,false],["Claude Opus 4.8",1,5,false],["Claude Opus 5",1,5,false],["Claude Opus 5.5",1,5,false],["Claude Fable 5.1",1,4,false],["GPT-6 Sol",1,2,false],["GPT-6 Luna",1,2,false],["GPT-6 Astra",1,2,false],["GPT-5.5",1,2,false],["Gemini 3.8 Flash",1,2,false],["Gemini 3.5 Flash",1,2,false],["Gemini 3.5 Flash Lite",1,2,false],["Gemini 3.1 Pro Preview",1,2,false]]}}],"A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.",["openrouter-api-snapshot"]],["gateway-markup-vs-first-party","OpenRouter markup over the first-party list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","openrouter-fees","price-anthropic","price-openai","price-google"]],["gateway-vs-direct-claude-haiku-4-5","Claude Haiku 4.5: OpenRouter vs Anthropic list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-anthropic"]],["gateway-vs-direct-claude-sonnet-5","Claude Sonnet 5: OpenRouter vs Anthropic list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-anthropic"]],["gateway-vs-direct-claude-sonnet-5-5","Claude Sonnet 5.5: OpenRouter vs Anthropic list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-anthropic"]],["gateway-vs-direct-claude-opus-4-8","Claude Opus 4.8: OpenRouter vs Anthropic list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-anthropic"]],["gateway-vs-direct-claude-opus-5","Claude Opus 5: OpenRouter vs Anthropic list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-anthropic"]],["gateway-vs-direct-claude-opus-5-5","Claude Opus 5.5: OpenRouter vs Anthropic list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-anthropic"]],["gateway-vs-direct-claude-fable-5-1","Claude Fable 5.1: OpenRouter vs Anthropic list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-anthropic"]],["gateway-vs-direct-gpt-6-luna","GPT-6 Luna: OpenRouter vs OpenAI list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-openai"]],["gateway-vs-direct-gemini-3-8-flash","Gemini 3.8 Flash: OpenRouter vs Google list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-google"]],["gateway-vs-direct-gemini-3-5-flash","Gemini 3.5 Flash: OpenRouter vs Google list price","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot","price-google"]],["provider-prices-claude-haiku-4-5","Claude Haiku 4.5: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-claude-sonnet-5","Claude Sonnet 5: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-claude-opus-4-8","Claude Opus 4.8: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-claude-opus-5","Claude Opus 5: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-claude-opus-5-5","Claude Opus 5.5: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-claude-fable-5-1","Claude Fable 5.1: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-gpt-6-sol","GPT-6 Sol: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-gpt-6-luna","GPT-6 Luna: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-gpt-6-astra","GPT-6 Astra: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-gpt-5-5","GPT-5.5: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-gpt-oss-120b","gpt-oss-120b: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-gemini-3-8-flash","Gemini 3.8 Flash: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-gemini-3-5-flash","Gemini 3.5 Flash: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-gemini-3-5-flash-lite","Gemini 3.5 Flash Lite: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-gemini-3-1-pro-preview","Gemini 3.1 Pro Preview: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-llama-4-maverick","Llama 4 Maverick: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-deepseek-v4-pro","DeepSeek V4 Pro 0423: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-deepseek-v4-flash","DeepSeek V4 Flash 0423: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-kimi-k3","Kimi K3: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]],["provider-prices-glm-5-3","GLM 5.3: price per million tokens by provider","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["openrouter-api-snapshot"]]]},{"$k":["id","columns","rows"],"$r":[["provider-index-models",[],[]],["gateway-overhead-status",[],[]],["provider-index-endpoints",[],[]]]},"\u0001"],["cost-thought-experiments","What if every call ran on Opus? Repricing real agent tokens","Agent token costs repriced: Haiku vs Sonnet vs Opus vs Fable","Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.","Agent recorded every token it used on 33 SWE-bench instances. What would the same tokens cost at other models’ list prices, and what did caching save?","Calculation, not a run.","2026-10-05","2026-10-05",["thought-experiment","llm-pricing","prompt-caching","opus","sonnet","haiku","jev"],[],["calc-repricing","agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","price-anthropic","price-google","price-openai","price-jev"],{"$k":["id","label","value","unit","display","n"],"$r":[["tokens-input","Input tokens recorded",162861253,"tokens","162.9M",33],["tokens-output","Output tokens recorded",1760652,"tokens","1.8M",33],["cache-read-share","Share of input served from cache",0.9401,"rate","94.0%",33],["sonnet-repriced","Recorded tokens at Sonnet 5.5 list price",87.23,"usd","$87.23",33],["opus-repriced","Same tokens at Opus 5.5 list price (calculation)",143.83,"usd","$143.83",33],["haiku-repriced","Same tokens at Haiku 4.5 list price (calculation)",43.61,"usd","$43.61",33],["no-cache-sonnet","Sonnet 5.5 without caching (calculation)",343.33,"usd","$343.33",33],["panel-cost-per-resolved-mean","Public panel mean cost per resolved instance (recorded)",0.569,"usd","$0.57",11]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["repriced-cost-per-resolved","Thought experiment: the same tokens at other list prices","Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price","bar","usd","USD per resolved instance",[{"name":"Repriced cost per resolved instance","points":{"$k":["label","value","highlight"],"$r":[["Claude Fable 5.1",12.852,false],["Claude Opus 5",8.723,false],["Claude Opus 5.5",5.753,false],["Claude Sonnet 5.5",3.489,true],["GPT-6.1 Sol",2.097,false],["Claude Haiku 4.5",1.745,false],["Gemini 3.x Flash",1.016,false],["Jev 1.13 (router)",0.042,false]]}}],"Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.",["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic","price-google","price-openai","price-jev"]],["cost-per-resolved-agent-vs-panel","Recorded cost per resolved instance: Agent vs the public panel","\u0001","bar","\u0001","\u0001",[],"\u0001",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"]],["prompt-cache-savings","Thought experiment: what prompt caching saved","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic"]],["token-cost-mix","Where the token dollars go","\u0001","bar","\u0001","\u0001",[],"\u0001",["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic"]],["router-overhead-jev","Thought experiment: the price of a routing decision on every call","\u0001","bar","\u0001","\u0001",[],"\u0001",["calc-repricing","price-jev","agent-swebench-c1","agent-swebench-c2"]]]},[{"id":"repricing-table","columns":[],"rows":[]}],"\u0001"],["cli-model-latency-tokens","Claude Code CLI vs Codex CLI vs the API: latency and tokens","Claude Code vs Codex CLI vs API: latency and token overhead","194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.","How much time and how many tokens does a coding CLI add on top of the model, and how do Claude Code and Codex compare on the same repair task?","For a one-line answer, the Codex CLI took a median 3.5 times as long as the OpenAI API with the same model and effort, and it sent about 19,551 input tokens instead of 17.","2026-10-03","2026-10-05",["claude-code","codex","latency","tokens","cli-vs-api"],[],["agent-provider-explorer"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["explorer-pass-rate","Evaluated runs that passed their validator",1,"rate","100% (194/194)",194,[0.9806,1],"18 excluded and 18 diagnostic receipts are not counted."],["exact-reply-cli-over-api","Codex CLI vs OpenAI API, median total time for a one-line answer",3.49,"ratio","3.5x slower",30,"\u0001","Medians 3.9 s (CLI) vs 1.1 s (API), same models and efforts."],["scheduler-claude-cli-median","Claude Code CLI (Sonnet 5.5) median time to repair the scheduler",15,"seconds","15.0 s",3,"\u0001","\u0001"],["scheduler-codex-cli-median","Codex CLI (GPT-6.1 Sol) median time to repair the scheduler",61.2,"seconds","61.2 s",3,"\u0001","\u0001"],["cli-hidden-prompt","Median input tokens the Codex CLI sends for a one-line request",19551,"tokens","19,551",15,"\u0001","The API sends 17 tokens for the same request."]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["cli-vs-api-exact-reply-latency","CLI vs API: time for a one-line answer","Matched cohort, fixed exact reply, 5 runs per configuration","dot-range","seconds","Seconds",[{"name":"Total time","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.97,0.65,1.5,5],["OpenAI API · GPT-6.1 Sol · low",1.02,0.96,1.87,5],["OpenAI API · GPT-6.1 Sol · high",1.52,1.35,2.23,5],["Codex CLI · GPT-6 Luna · none",3.19,2.88,3.83,5],["Codex CLI · GPT-6.1 Sol · low",4.18,3.86,4.53,5],["Codex CLI · GPT-6.1 Sol · high",4.19,3.81,4.69,5]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.82,0.51,1.37,5],["OpenAI API · GPT-6.1 Sol · low",0.87,0.84,1.74,5],["OpenAI API · GPT-6.1 Sol · high",1.34,1.26,2.12,5],["Codex CLI · GPT-6 Luna · none",2.79,2.46,3.42,5],["Codex CLI · GPT-6.1 Sol · low",3.75,3.44,4.1,5],["Codex CLI · GPT-6.1 Sol · high",3.79,3.37,4.3,5]]}}],"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.",["agent-provider-explorer"]],["cli-vs-api-small-coding-latency","CLI vs API: time for a small coding task","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-provider-explorer"]],["cli-vs-api-prompt-overhead","Hidden prompt: input tokens for the same one-line request","\u0001","bar","\u0001","\u0001",[],"\u0001",["agent-provider-explorer"]],["scheduler-repair-claude-vs-codex","Repairing a scheduler: Claude Code vs Codex vs API","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["agent-provider-explorer"]],["scheduler-repair-output-tokens","Output tokens to repair the scheduler","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-provider-explorer"]]]},[{"id":"cli-api-configurations","columns":[],"rows":[]}],"\u0001"],["coding-calibration","Coding calibration: what broke on three real pull requests","AI coding agent calibration on fastify, h3 and uvicorn tasks","One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.","On three real upstream issues, does the AI worker deliver a verified change, and what stops it when it does not?","Across four platform builds, verified delivery went from 0 of 3 to 1 of 3; the latest build delivered fastify/session cleanly, but h3 regressed to an empty patch when a formatting failure was lost across a context fold, and uvicorn passed its full suite yet stayed unverified for two gate-classification reasons.","2026-10-02","2026-10-05",["calibration","coding-agents","failures","fastify","h3","uvicorn"],[],["agent-coding-calibration"],{"$k":["id","label","value","unit","display","n"],"$r":[["verified-latest","Verified deliveries, latest build",1,"count","1 of 3",3],["cost-latest","Notional cost, latest build, all 3 tasks",11.06,"usd","$11.06",3],["refusals-trend","Guardrail refusals, first vs latest slice",19,"count","26 → 19",3],["first-run-cost","First calibration run (capped, fastify/session)",4.89,"usd","$4.89, stopped at cap",1]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["calibration-outcomes-by-slice","Three real tasks, four platform builds","Tasks per slice: functional pass, verified delivery, pull request opened","grouped-bar","count","Tasks (of 3)",{"$k":["name","points"],"$r":[["Functional pass (offline gates)",{"$k":["label","value","n"],"$r":[["Baseline (capped)",2,3],["Fix wave 1 (capped)",2,3],["Uncapped, build f0ac3a8a",1,3],["Uncapped, build 236c0d3f",1,3]]}],["Verified delivery",{"$k":["label","value","n","highlight"],"$r":[["Baseline (capped)",0,3,false],["Fix wave 1 (capped)",0,3,false],["Uncapped, build f0ac3a8a",0,3,false],["Uncapped, build 236c0d3f",1,3,true]]}],["Pull request opened",{"$k":["label","value","n"],"$r":[["Baseline (capped)",1,3],["Fix wave 1 (capped)",2,3],["Uncapped, build f0ac3a8a",2,3],["Uncapped, build 236c0d3f",2,3]]}]]},"One attempt per task per slice. The two capped slices stopped at 20 minutes or $5; the uncapped slices had no ceiling. Each slice is a different platform build, so a change is not a matched improvement.",["agent-coding-calibration"]],["calibration-cost-by-task","Notional model cost per task, by slice","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-coding-calibration"]],["calibration-minutes-by-task","Wall time per task, by slice","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["agent-coding-calibration"]],["calibration-guardrail-refusals","Guardrail refusals per slice","\u0001","bar","\u0001","\u0001",[],"\u0001",["agent-coding-calibration"]]]},[{"id":"calibration-every-attempt","columns":[],"rows":[]}],"\u0001"],["single-call-vs-agent-loop","Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks","Agent loop vs single call: does tool use improve accuracy?","118 attempts on 8 hard tasks: one call vs an agent loop that runs code in a sandbox. Pass rate, time, tokens and cost.","On 8 hard tasks with strict validators, does an agent loop that may write and run code in a sandbox pass more often than one call with tools off, and what does the loop cost in time, tokens, tool calls and list price per pass?","Not clearly, on this set: no model's agent loop is ahead of its single call by the 95% intervals.","2026-10-06","2026-10-06",["agent-loop","tool-use","single-call","hard-tasks","claude-haiku","claude-sonnet","gpt-6-luna","claude-code","codex-cli","sandbox"],[],["agent-agent-loop","agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["agent-loop-pass-haiku","Claude Haiku 4.5 strict pass rate, agent loop",0.5417,"rate","54% (13/24)",24,[0.3507,0.7211],"Single call: 11/24 (28% to 65%); the intervals overlap, so there is no clear difference."],["agent-loop-pass-sonnet","Claude Sonnet 5.5 strict pass rate, agent loop",1,"rate","100% (16/16)",16,[0.8064,1],"Single call: 24/24 (86% to 100%); both at the ceiling, so the set cannot separate them."],["agent-loop-pass-luna","GPT-6 Luna strict pass rate, agent loop (Codex CLI)",0.8571,"rate","86% (12/14)",14,[0.6006,0.9599],"Single call: 10/16 (39% to 82%); the intervals overlap, so there is no clear difference."],["agent-loop-haiku-gain","Change in strict pass rate, agent loop minus single call, Claude Haiku 4.5 (calculation)",0.0833,"rate","+8 points",48,"\u0001","11/24 to 13/24; the intervals overlap."],["agent-loop-used-tools","Scored agent-loop attempts that ran at least one tool",29,"count","29 of 54",54,"\u0001","Haiku 4.5 24 of 24; Sonnet 5.5 3 of 16; GPT-6 Luna (Codex CLI) 2 of 14"],["agent-loop-outside-edits","Edits outside the work folder that ran, in agent-loop attempts",0,"count","0 in 56 attempts; 2 outside attempts were refused before they ran",56,"\u0001","\u0001"],["agent-loop-contaminated","Agent-loop attempts left out for reading outside the work folder",2,"count","2 of 56",56,"\u0001","\u0001"],["agent-loop-median-tools","Median tool calls per agent-loop attempt",1.5,"count","1.5",54,"\u0001","Haiku 4.5 3; Sonnet 5.5 0; GPT-6 Luna (Codex CLI) 0"]]},{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","sourceIds"],"$r":[["agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks","Same tasks and validators. Whiskers are 95% Wilson intervals","dot-range","rate","higher","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",0.4583,0.2789,0.6493,24],["Claude Haiku 4.5 (agent loop) · Claude Code",0.5417,0.3507,0.7211,24],["Claude Sonnet 5.5 (single call) · Claude Code",1,0.862,1,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",1,0.8064,1,16],["GPT-6 Luna (single call) · Codex CLI",0.625,0.3864,0.8152,16],["GPT-6 Luna (agent loop) · Codex CLI",0.8571,0.6006,0.9599,14]]}}],"Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.","ci95",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-by-task","Strict passes per task: single call vs agent loop","\u0001","grouped-bar","\u0001","higher","\u0001",[],"\u0001","\u0001",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-total-time","Total time per attempt: single call vs agent loop","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001","\u0001",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-tokens","Tokens per attempt: single call vs agent loop","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001","\u0001",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-tool-calls","Tool calls per agent-loop attempt","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001","\u0001",["agent-agent-loop"]],["agent-loop-cost-per-pass","List-price cost per strict pass: single call vs agent loop (calculation)","\u0001","bar","\u0001","\u0001","\u0001",[],"\u0001","\u0001",["agent-agent-loop","agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]]]},[{"id":"agent-loop-cells","columns":[],"rows":[]},{"id":"agent-loop-audit","columns":[],"rows":[]}],"\u0001"],["haiku-thinking-on-off","Does thinking pay for Claude Haiku 4.5? Thinking on vs off","Claude Haiku 4.5 thinking on vs off: benchmark results","Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.","Does extended thinking pay for Claude Haiku 4.5: what does it buy in accuracy, and what does it cost in time and money, on typed routing decisions and on hard tasks?","Routing accuracy is unresolved on this sample of 82 paired decisions.","2026-10-06","2026-10-06",["claude-haiku","extended-thinking","routing","hard-tasks","latency","cost"],[],["agent-haiku-thinking","agent-routing","calc-repricing","price-anthropic","agent-provider-h2h-hard"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["haiku-thinking-exact-off","Claude Haiku 4.5 (thinking off): exact routing decisions",0.8659,"rate","87% (71/82)",82,[0.7755,0.9234],"\u0001"],["haiku-thinking-exact-on","Claude Haiku 4.5 (thinking on): exact routing decisions",0.8902,"rate","89% (73/82)",82,[0.8044,0.9412],"\u0001"],["haiku-thinking-keys-off","Claude Haiku 4.5 (thinking off): per-question routing accuracy",0.9124,"rate","91% (177/194)",194,[0.8642,0.9446],"\u0001"],["haiku-thinking-keys-on","Claude Haiku 4.5 (thinking on): per-question routing accuracy",0.9433,"rate","94% (183/194)",194,[0.9013,0.968],"\u0001"],["haiku-thinking-mcnemar","Paired exact test, thinking off vs on (exact McNemar p)",0.754,"score","0.754",82,"\u0001","p = 0.754 (6 only on, 4 only off, 82 cases)"],["haiku-thinking-wall-off","Claude Haiku 4.5 (thinking off): median wall time per routing decision",4.66,"seconds","4.66 s",82,"\u0001","Range 2.20 s to 11.57 s; not a confidence interval. Nearest-rank p50."],["haiku-thinking-wall-on","Claude Haiku 4.5 (thinking on): median wall time per routing decision",12.54,"seconds","12.54 s",82,"\u0001","Range 5.86 s to 51.28 s; not a confidence interval. Nearest-rank p50."],["haiku-thinking-wall-ratio","Median wall time, thinking on ÷ thinking off (calculation)",2.69,"ratio","2.7x (12.54 s ÷ 4.66 s)",82,"\u0001","Calculation from two medians; the arms did not run at the same time."],["haiku-thinking-tokens-off","Claude Haiku 4.5 (thinking off): thinking tokens per routing decision (mean)",0,"tokens","0",82,"\u0001","\u0001"],["haiku-thinking-tokens-on","Claude Haiku 4.5 (thinking on): thinking tokens per routing decision (mean)",1100.8,"tokens","1,101 (range 285 to 4,443)",82,"\u0001","\u0001"],["haiku-thinking-cost-off","Claude Haiku 4.5 (thinking off): list-price cost per 1,000 routing decisions (calculation)",3.364,"usd","$3.364",82,"\u0001","\u0001"],["haiku-thinking-cost-on","Claude Haiku 4.5 (thinking on): list-price cost per 1,000 routing decisions (calculation)",8.924,"usd","$8.924",82,"\u0001","\u0001"],["haiku-thinking-hard-strict-off","Claude Haiku 4.5 (thinking off): strict passes on hard tasks",0.1667,"rate","17% (4/24)",24,[0.0668,0.3585],"\u0001"],["haiku-thinking-hard-strict-on","Claude Haiku 4.5 (thinking on): strict passes on hard tasks",0.4583,"rate","46% (11/24)",24,[0.2789,0.6493],"\u0001"],["haiku-thinking-hard-time-off","Claude Haiku 4.5 (thinking off): median total time per hard-task call",2.95,"seconds","2.9 s",24,"\u0001","Range 1.7 s to 13.0 s; not a confidence interval."],["haiku-thinking-hard-time-on","Claude Haiku 4.5 (thinking on): median total time per hard-task call",39.01,"seconds","39.0 s",24,"\u0001","Range 15.3 s to 75.1 s; not a confidence interval."],["haiku-thinking-hard-reasoning-on","Claude Haiku 4.5 (thinking on): median reasoning tokens per hard-task call",4556,"tokens","4,556",24,"\u0001","\u0001"],["haiku-thinking-hard-cost-per-pass-off","Claude Haiku 4.5 (thinking off): list-price cost per strict pass on hard tasks (calculation)",0.03654,"usd","$0.0365",24,"\u0001","\u0001"],["haiku-thinking-hard-cost-per-pass-on","Claude Haiku 4.5 (thinking on): list-price cost per strict pass on hard tasks (calculation)",0.0672,"usd","$0.0672",24,"\u0001","\u0001"],["haiku-thinking-new-calls","Counted model calls made for this study (thinking off)",106,"calls","106 (82 routing, 24 hard tasks; 2 more uncounted probes)",106,"\u0001","\u0001"],["haiku-thinking-off-check","Thinking-off calls that reported any thinking tokens",0,"count","0 of 106",106,"\u0001","Each call reports its thinking tokens in the CLI result; 0 means the setting held."]]},{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","factContext","sourceIds"],"$r":[["haiku-thinking-router-exact","Haiku thinking study: typed routing decisions answered exactly right","Claude Haiku 4.5 and Claude Sonnet 5.5 (low effort); 82 decisions, the same cases for every arm","dot-range","rate","higher","Correct",[{"name":"Exact decisions (every scored question right)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0.8659,0.7755,0.9234,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5 (low) · Claude Code",0.939,0.8651,0.9737,82,false]]}},{"name":"Per-question accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0.9124,0.8642,0.9446,194,true],["Claude Haiku 4.5 (thinking on) · Claude Code",0.9433,0.9013,0.968,194,false],["Claude Sonnet 5.5 (low) · Claude Code",0.9742,0.9411,0.9889,194,false]]}}],"Whiskers are 95% Wilson intervals. Questions within a decision are related; per-question intervals are descriptive, not an independent-question test. The thinking-on and Sonnet arms are the recorded 2026-10-05 routing run, reused, not rerun; the thinking-off arm ran later on another account. An unanswered question counts as wrong.","ci95","typed routing decisions, thinking on vs off",["agent-haiku-thinking","agent-routing"]],["haiku-thinking-router-latency","Haiku thinking study: time per routing decision","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001","\u0001","\u0001",["agent-haiku-thinking","agent-routing"]],["haiku-thinking-router-tokens","Haiku thinking study: thinking and visible output tokens per routing decision","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001","\u0001","\u0001",["agent-haiku-thinking","agent-routing"]],["haiku-thinking-router-cost","Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)","\u0001","bar","\u0001","\u0001","\u0001",[],"\u0001","\u0001","\u0001",["agent-haiku-thinking","agent-routing","calc-repricing","price-anthropic"]],["haiku-thinking-hard-pass","Haiku thinking study: pass rate on eight hard tasks","\u0001","dot-range","\u0001","higher","\u0001",[],"\u0001","\u0001","\u0001",["agent-haiku-thinking","agent-provider-h2h-hard"]],["haiku-thinking-hard-time","Haiku thinking study: total time per call on hard tasks","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001","\u0001","\u0001",["agent-haiku-thinking","agent-provider-h2h-hard"]]]},{"$k":["id","columns","rows"],"$r":[["haiku-thinking-paired-test",[],[]],["haiku-thinking-router-cells",[],[]],["haiku-thinking-by-decision-type",[],[]],["haiku-thinking-hard-cells",[],[]],["haiku-thinking-hard-matrix",[],[]]]},{"statIds":["haiku-thinking-exact-off","haiku-thinking-exact-on"],"testStatId":"haiku-thinking-mcnemar"}],["json-schema-vs-instructions","Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI","LLM structured output: JSON schema vs instructions","96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.","For the same extraction tasks, does enforcing a JSON schema through the CLI change the strict pass rate, the format misses and the wrong values, compared with asking for JSON in the prompt?","96 counted calls on three JSON extraction prompts, each asked with instructions only and with the CLI’s JSON schema mode.","2026-10-07","2026-10-07",["structured-output","json-schema","format-miss","claude-code","codex-cli","claude-haiku","claude-sonnet","gpt-6-1-sol"],[],["agent-structured-output"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["structured-output-calls","Counted calls in this study (every one counted)",96,"calls","96 (72 Claude Code, 24 Codex CLI)","\u0001","\u0001","\u0001"],["structured-output-format-miss-instructions","Format-miss rate with instructions only, all models",0.3542,"rate","35% (17/48)",48,[0.2343,0.4956],"A right answer in a code fence or prose. Every error counted as a call."],["structured-output-format-miss-schema","Format-miss rate with a JSON schema, all models",0,"rate","0% (0/48)",48,[0,0.0741],"A right answer in a code fence or prose. Every error counted as a call."],["structured-output-fenced-instructions","Replies in a code fence with instructions only, all models",0.5,"rate","50% (24/48)",48,[0.3639,0.6361],"The prompt said: no code fence. Completed replies only; a fence can hold a right or a wrong answer."],["structured-output-fenced-schema","Replies in a code fence with a JSON schema, all models",0,"rate","0% (0/48)",48,[0,0.0741],"Completed replies only. Claude Code uses its structured_output field. Codex CLI uses its final message."],["structured-output-wrong-values-instructions","Wrong-values rate with instructions only, all models",0.1458,"rate","15% (7/48)",48,[0.0725,0.2717],"A completed reply that is neither a strict pass nor a format miss. Denominator includes all attempts; errors are a separate outcome."],["structured-output-wrong-values-schema","Wrong-values rate with a JSON schema, all models",0.125,"rate","13% (6/48)",48,[0.0586,0.247],"A completed reply that is neither a strict pass nor a format miss. Denominator includes all attempts; errors are a separate outcome."],["structured-output-strict-instructions","Strict pass rate with instructions only, all models",0.5,"rate","50% (24/48)",48,[0.3639,0.6361],"Pooled over three models and three prompts; the models differ in n, so read the per-model chart."],["structured-output-strict-schema","Strict pass rate with a JSON schema, all models",0.875,"rate","88% (42/48)",48,[0.753,0.9414],"Pooled over three models and three prompts; the models differ in n, so read the per-model chart."]]},{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","sourceIds"],"$r":[["structured-output-pass-rate","Does a JSON schema raise the pass rate? Instructions vs schema mode","Three extraction prompts pooled; whiskers are 95% Wilson intervals","dot-range","rate","higher","Passed",[{"name":"Strict pass: the whole reply is the right JSON","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0,0,0.138,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}},{"name":"Right answer in any format (strict pass or format miss)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0.7083,0.5083,0.8509,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}}],"Whiskers are 95% Wilson intervals (a calculation) over 24 calls and 12 calls per configuration; every error counts as a fail. Strict: the whole reply parses as JSON and matches the expected answer exactly. A format miss is a right answer inside a code fence or prose, so it is never a strict pass.","ci95",["agent-structured-output"]],["structured-output-outcomes","What each call produced: strict pass, format miss, wrong values or error","\u0001","stacked-bar","\u0001","\u0001","\u0001",[],"\u0001","\u0001",["agent-structured-output"]],["structured-output-time","Time per call, instructions vs schema mode","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001","\u0001",["agent-structured-output"]],["structured-output-tokens","Output and reasoning tokens per call, instructions vs schema mode","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001","\u0001",["agent-structured-output"]]]},[{"id":"structured-output-cells","columns":[],"rows":[]},{"id":"structured-output-pairs","columns":[],"rows":[]}],"\u0001"],["prompt-cache-across-sessions","Does a new Claude Code session reuse the prompt cache of an earlier one?","Claude Code prompt caching across sessions, tested","30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.","In the earlier caching study, a new Claude Code session did not read the cache that an earlier session wrote. Does a fixed working folder change that, and does putting the ledger in the system prompt help?","Later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder.","2026-10-07","2026-10-07",["prompt-caching","cache-reuse","claude-code","codex-cli","claude-sonnet","gpt-6-1-sol","working-folder","calculation"],[],["agent-cache-sessions","calc-cache-pricing","price-anthropic","agent-caching-consistency"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["cache-sessions-later-reuse-a","Later sessions with at least 50% of turn-1 input cached, A: new folder each time",0,"rate","0 of 2 (95% interval 0% to 66%)",2,[0,0.6576],"Claude Sonnet 5.5 in Claude Code. A later session is session 2 or 3. Reused = turn-1 read share of 0.5 or more. The surviving protocol records this rule but dates from after the calls. No token-level trace identifies the ledger."],["cache-sessions-later-reuse-b","Later sessions with at least 50% of turn-1 input cached, B: fixed folder",1,"rate","2 of 2 (95% interval 34% to 100%)",2,[0.3424,1],"Claude Sonnet 5.5 in Claude Code. A later session is session 2 or 3. Reused = turn-1 read share of 0.5 or more. The surviving protocol records this rule but dates from after the calls. No token-level trace identifies the ledger."],["cache-sessions-later-reuse-c","Later sessions with at least 50% of turn-1 input cached, C: fixed folder, ledger in system prompt",1,"rate","2 of 2 (95% interval 34% to 100%)",2,[0.3424,1],"Claude Sonnet 5.5 in Claude Code. A later session is session 2 or 3. Reused = turn-1 read share of 0.5 or more. The surviving protocol records this rule but dates from after the calls. No token-level trace identifies the ledger."],["cache-sessions-pooled-new-folder","Later sessions with at least 50% of turn-1 input cached, new folders, pooled across studies",0,"rate","0 of 6 (95% interval 0% to 39%)",6,[0,0.3903],"Calculation, post hoc (Amendment 2): this study’s setup A plus the earlier study’s 4 later Claude Code sessions (2 Sonnet, 2 Opus; 5 turns per session; another ledger). The 4 earlier sessions ran under a different Claude login, with Sonnet and Opus sessions interleaved: 13 to 27 seconds between a session’s last call and the next same-model session’s first call (calculation), against under 1 second here."],["cache-sessions-pooled-fixed-folder","Later sessions with at least 50% of turn-1 input cached, fixed folder, B and C pooled",1,"rate","4 of 4 (95% interval 51% to 100%)",4,[0.5101,1],"Calculation, post hoc (Amendment 2): setups B and C differ in where the ledger sits."],["cache-sessions-turn1-usd-new-folder","Turn-1 list-price cost of a later session, new folder each time (calculation)",0.025875,"usd","$0.0259",2,"\u0001","Mean of sessions 2 and 3, setup A; range $0.025871 to $0.025879. Calculation from recorded tokens and list prices; not a bill."],["cache-sessions-turn1-usd-fixed-folder","Turn-1 list-price cost of a later session, fixed folder (calculation)",0.001599,"usd","$0.0016",4,"\u0001","Mean of sessions 2 and 3 in setups B and C; range $0.001597 to $0.001601. Calculation from recorded tokens and list prices; not a bill."],["cache-sessions-turn1-cost-ratio","Turn-1 cost of a later session: new folder as a multiple of fixed folder (calculation)",16.18,"ratio","16.2×",6,"\u0001","Calculation. Turn-1 input here is about 7.8k tokens, including the ledger. A different workload can change both token use and cache matches."],["cache-sessions-codex-declared-threshold","Later Codex CLI sessions with at least 50% of turn-1 input cached (protocol threshold)",0.5,"rate","2 of 4 (95% interval 15% to 85%)",4,[0.15,0.85],"Counter threshold only. A prefix read can meet it; it does not prove that the ledger was read. The surviving protocol was created after the calls."],["cache-sessions-codex-later-reuse","Later Codex CLI sessions with at least 90% of turn-1 input cached (post-hoc rule) (setups A and B)",0,"rate","0 of 4 (95% interval 0% to 49%)",4,[0,0.4899],"GPT-6.1 Sol at medium effort. Reused = cached input of 90% or more of the input (Amendment 1, set after the counts were seen); the highest turn-1 cached share was 56%. Under the 50% threshold we recorded in the protocol, 2 of 4 later Codex sessions would count (8,960 tokens cached in each, the count the Codex probe call showed, which we read as the Codex system prefix). The 90% rule, set after we saw the counts, gives 0 of 4. Their Wilson 95% intervals are 15% to 85% and 0% to 49%, respectively."],["cache-sessions-calls","Counted calls in this study (every one counted)",30,"calls","30 (18 Claude Code, 12 Codex CLI), plus 2 uncounted probe calls","\u0001","\u0001","\u0001"]]},{"$k":["id","title","subtitle","kind","unit","polarity","xLabel","yLabel","series","note","sourceIds"],"$r":[["cache-sessions-turn1-read-share","Claude Sonnet 5.5 · Claude Code: share of turn-1 input read from the cache, by setup and session","One bar per call: turn 1 of one session; 3 sessions per setup, run back to back; read-share calculation","grouped-bar","rate","none","Setup","Input tokens read from cache",{"$k":["name","points"],"$r":[["Session 1 (first in its setup)",{"$k":["label","value","n"],"$r":[["A: new folder each time",0.0676,1],["B: fixed folder",0.1867,1],["C: fixed folder, ledger in system prompt",0.0691,1]]}],["Session 2",{"$k":["label","value","n"],"$r":[["A: new folder each time",0.1863,1],["B: fixed folder",0.9997,1],["C: fixed folder, ledger in system prompt",0.9997,1]]}],["Session 3",{"$k":["label","value","n"],"$r":[["A: new folder each time",0.1863,1],["B: fixed folder",0.9997,1],["C: fixed folder, ledger in system prompt",0.9997,1]]}]]},"Session 1 was the first session to use its setup’s ledger. Sessions 2 and 3 used that same ledger. Different seeds prevent full ledger-prefix reuse between setups; shared CLI-prefix reads remain possible. We interpret session-1 reads as a shared CLI prefix; no token-level trace proves this. Read share (calculation) = cache reads ÷ (uncached input + cache reads + cache writes), as the provider reports them. Each bar is one call, not a rate; the counts per setup are in the table. 2 later sessions per setup is a small number.",["agent-cache-sessions"]],["cache-sessions-turn1-cost","Claude Sonnet 5.5 · Claude Code: list-price cost of turn 1, by setup and session (calculation)","\u0001","grouped-bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-cache-sessions","calc-cache-pricing","price-anthropic"]],["cache-sessions-turn-time","Claude Sonnet 5.5 · Claude Code: time per turn, by setup","\u0001","dot-range","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-cache-sessions"]],["cache-sessions-codex-turn1-cached","GPT-6.1 Sol (medium) · Codex CLI: cached input tokens on turn 1, by setup and session","\u0001","grouped-bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-cache-sessions"]]]},[{"id":"cache-sessions-turns","columns":[],"rows":[]}],{"statIds":["cache-sessions-later-reuse-a","cache-sessions-later-reuse-b"]}],["prompt-cache-break-even","Prompt cache break-even: after how many reuses does a cached prefix cost less?","Prompt cache break-even by model and session length","A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.","With the cache-write surcharge Anthropic lists, after how many reuses does a cached prompt prefix cost less than no cache? What does that mean for a session of 1 to 20 turns, for each model’s price, and for a workload split across sessions?","Calculation: a wholly new prefix with a 1-hour write costs less from the 3rd request, after 2 reuses.","2026-10-06","2026-10-06",["prompt-caching","break-even","thought-experiment","calculation","llm-pricing","claude-haiku","claude-sonnet","claude-opus","claude-fable","gpt-6-1-sol","gpt-6-luna"],[],["calc-cache-pricing","agent-caching-consistency","price-anthropic","price-openai"],{"$k":["id","label","value","unit","display","note","n"],"$r":[["cache-break-even-1h-reuses","Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation)",2,"count","2 reuses (the 3rd request)","Same for Haiku 4.5, Sonnet 5.5, Opus 5.5 and Fable 5.1. Exact break-even 1.03 to 1.11 reuses.","\u0001"],["cache-break-even-1h-reuses-recorded","Reuses before a 1-hour cached prefix costs less, 19% of it already cached as a pooled recorded share (calculation)",1,"count","1 reuse (the 2nd request)","Exact break-even 0.65 to 0.72 reuses. The recorded Claude Code sessions paid back on turn 2 (see the recorded-payback stat).","\u0001"],["cache-break-even-5m-reuses","Reuses before a 5-minute cached prefix costs less, whole prefix new (calculation, assumed 1.25× write)",1,"count","1 reuse (the 2nd request)","Exact break-even 0.26 to 0.28 reuses. The 1.25 multiplier is an assumption.","\u0001"],["cache-break-even-no-surcharge-reuses","Reuses before a cached prefix costs less when the list price has no write surcharge (GPT-6.1 Sol and GPT-6 Luna; calculation)",1,"count","1 reuse (saves from the first read)","Exact break-even 0 reuses: a write is plain input, so the first read is already cheaper than sending the prefix again. Another table of the product lists a write price of 1.25× input for both models. At that price the exact break-even is 0.26 reuses (GPT-6.1 Sol) and 0.28 reuses (GPT-6 Luna), so 1 reuse still pays. We did not check the vendor price.","\u0001"],["cache-break-even-prefix-sonnet","Mean turn-1 input in the Sonnet 5.5 sessions (calculation)",7831,"tokens","7,831 tokens","Rounded mean calculation from 3 sessions (range 7,831 to 7,832): uncached input + cache reads + cache writes on turn 1.",3],["cache-break-even-prefix-opus","Mean turn-1 input in the Opus 5.5 sessions (calculation)",7828,"tokens","7,828 tokens","Rounded mean calculation from 3 sessions (the same in every session).",3],["cache-break-even-precached-share","Share of turn-1 input read from the cache (calculation; origin not isolated)",0.1869,"rate","19% (8,778 of 46,978 turn-1 tokens)","Calculation across 6 sessions; session share range 18.680% to 18.689%. This is not a binomial pass rate. This part cost the read price on turn 1, not the write price.",6],["cache-break-even-recorded-payback","Turn at which a recorded Claude Code session’s total input cost with the cache first fell below its cost with no cache (calculation on recorded tokens)",2,"count","2 turns (6 of 6 sessions)","Input side only, 1-hour writes at the list price, output left out. After turn 1 the cache had cost more, as the formula says for a prefix used once.",6],["cache-break-even-once-penalty","Cost of caching a prefix that is used once, as a multiple of no cache (1-hour write; calculation)",2,"ratio","2.0x","A 5-minute write: 1.3x (assumption). Same for every Anthropic model in the price list.","\u0001"],["cache-break-even-sonnet-10-turns","Saving from a 1-hour cache over 10 turns, Sonnet 5.5 prefix (calculation)",0.71,"rate","71.0% ($45.42 vs $156.62 per 1,000 sessions)","Prefix of 7,831 tokens; whole prefix new; output left out.","\u0001"],["cache-break-even-opus-10-turns","Saving from a 1-hour cache over 10 turns, Opus 5.5 prefix, cache read $0.2 per M (calculation)",0.755,"rate","75.5% ($76.71 vs $313.12 per 1,000 sessions)","Prefix of 7,828 tokens; whole prefix new; output left out.","\u0001"],["cache-break-even-opus-read-price","Saving from a 1-hour cache over 10 turns, Opus 5.5 prefix, cache read $0.4 per M (calculation)",0.71,"rate","71.0% ($90.80 vs $313.12 per 1,000 sessions)","Break-even 1.11 reuses at $0.4 per M against 1.05 at $0.2 per M. The vendor price was not checked.","\u0001"],["cache-break-even-split-sonnet","10 one-turn sessions with no shared reuse with a 1-hour cache, as a multiple of one 10-turn session, Sonnet 5.5 prefix (calculation)",6.9,"ratio","6.9x ($313.24 vs $45.42 per 1,000 workloads)","With no cache the same 10 requests cost $156.62. Each recorded session ran in a new temporary folder, and 0 of 4 met the read-more-than-write proxy. The split case assumes a wholly new prefix with no shared reuse.","\u0001"],["cache-break-even-cross-session","Later sessions whose first turn read more tokens than it wrote (reuse proxy)",0,"count","0 of 4","95% Wilson 0% to 49%, n = 4 later sessions. Same proxy as the caching study: reads exceed writes. It cannot identify cache provenance. The cause was not tested.",4],["cache-break-even-5m-writes-recorded","Recorded Claude turns that wrote a 5-minute cache entry",0,"count","0 of 30","Every recorded write was a 1-hour write, so the 1.25 multiplier of the 5-minute rows was not tested.",30],["cache-break-even-check-sonnet","Saving over 5 turns on the input side, recorded vs the formula (calculation), Sonnet 5.5",0.5535,"rate","55.4% recorded (formula 52.0% with a new prefix, 59.1% with 19% already cached)","Calculation on recorded tokens and list prices; session saving range 54.3% to 57.1%, not a confidence interval. The recorded figure includes new turn tokens; the formula does not.",3],["cache-break-even-check-opus","Saving over 5 turns on the input side, recorded vs the formula (calculation), Opus 5.5",0.5918,"rate","59.2% recorded (formula 56.0% with a new prefix, 63.3% with 19% already cached)","Calculation on recorded tokens and list prices; session saving range 58.5% to 59.7%, not a confidence interval. Opus 5.5 cache read at $0.2 per M.",3]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["cache-break-even-reads","Reuses before a cached prefix costs less, by model and write type (calculation)","The break-even point: reuses at which the cached and the uncached cost are equal, at list prices","grouped-bar","score","Reuses at break-even",{"$k":["name","points"],"$r":[["1-hour write (2× input), whole prefix new",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",1.11],["Claude Sonnet 5.5",1.11],["Claude Opus 5.5 (cache read $0.2 per M)",1.05],["Claude Opus 5.5 (cache read $0.4 per M)",1.11],["Claude Fable 5.1",1.03]]}],["1-hour write, pooled n = 6 session share, 19% already cached (as recorded)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0.72],["Claude Sonnet 5.5",0.72],["Claude Opus 5.5 (cache read $0.2 per M)",0.67],["Claude Opus 5.5 (cache read $0.4 per M)",0.72],["Claude Fable 5.1",0.65]]}],["5-minute write (1.25× input, an assumption)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0.28],["Claude Sonnet 5.5",0.28],["Claude Opus 5.5 (cache read $0.2 per M)",0.26],["Claude Opus 5.5 (cache read $0.4 per M)",0.28],["Claude Fable 5.1",0.26]]}]]},"Calculation, not a run: break-even reuses = (write price − input price) ÷ (input price − cache-read price). A cached prefix costs less once the reuses pass that point. On a new prefix, a 1-hour write needs 2 reuses (the 3rd request). With 19% already cached it needs 1 reuse, and a 5-minute write needs 1 reuse. The 1.25 of the 5-minute write is an assumption. The caching study protocol states it as Anthropic’s published figure. No 5-minute write occurred in the recorded sessions. We recorded the 19% on Sonnet 5.5 and Opus 5.5 sessions. For Haiku 4.5 and Fable 5.1 it is a what-if. GPT-6.1 Sol and GPT-6 Luna list no write surcharge, so their break-even is 0 reuses and the chart leaves them out. The price list gives Opus 5.5 a cache read of $0.2 per million. Another table of the product lists $0.4. We did not check the vendor price, so the chart shows both.",["calc-cache-pricing","price-anthropic","price-openai"]],["cache-break-even-cost-curve","Cost of a reused prefix with and without the cache, by session length (calculation)","\u0001","line","\u0001","\u0001",[],"\u0001",["calc-cache-pricing","agent-caching-consistency","price-anthropic","price-openai"]],["cache-break-even-session-split","One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation)","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["calc-cache-pricing","agent-caching-consistency","price-anthropic","price-openai"]]]},{"$k":["id","columns","rows"],"$r":[["cache-break-even-models",[],[]],["cache-break-even-saving-by-turns",[],[]],["cache-break-even-cost-per-1000",[],[]],["cache-break-even-split",[],[]],["cache-break-even-inputs",[],[]]]},"\u0001"],["routing-holdout","Jev vs Claude routers on unseen decisions: a blind holdout","Jev vs Claude router accuracy on unseen decisions","Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.","Does a router keep its accuracy on typed routing decisions that nobody tuned against its answers?","The holdout does not establish a router ranking.","2026-10-06","2026-10-06",["routing","jev","claude-haiku","claude-sonnet","model-routing","holdout","benchmark-method"],[],["agent-routing-holdout","agent-routing","calc-repricing","price-jev","price-anthropic"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["holdout-exact-jev","Jev 1.13 (TypeSafe): exact on unseen decisions",0.8214,"rate","82% (46/56)",56,[0.7016,0.9],"\u0001"],["holdout-exact-claude-haiku","Claude Haiku 4.5 · Claude Code: exact on unseen decisions",0.7857,"rate","79% (44/56)",56,[0.6618,0.8729],"\u0001"],["holdout-exact-claude-sonnet","Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions",0.875,"rate","88% (49/56)",56,[0.7637,0.9381],"\u0001"],["holdout-jev-stability","Jev 1.13 (TypeSafe): same answers on every key in 3 repetitions",0.9464,"rate","95% (53/56)",56,[0.8539,0.9816],"\u0001"],["holdout-jev-reps-range","Jev 1.13 (TypeSafe): exact in each of 3 repetitions",0.8214,"rate","46, 46, 47 of 56",56,"\u0001","Range 46 to 47 of 56; not a confidence interval. The value is the first repetition."],["holdout-label-agreement","Second labeller (GPT-6.1 Sol): inside our acceptable label sets",0.92,"rate","92% (115/125)",125,[0.859,0.956],"Per labelled question, across all four decision types. 24 of the 125 questions accept two or three labels, so this counts a second label that is inside our set, not only our first label. First label only: 100 of 125 (80%)."],["holdout-label-agreement-first","Second labeller (GPT-6.1 Sol): equal to our first label",0.8,"rate","80% (100/125)",125,[0.7214,0.8607],"Per labelled question. Our first label is the least inclusive one in our set. The inside-our-set figure is holdout-label-agreement."],["holdout-gap-jev","Jev 1.13 (TypeSafe): holdout minus tuned-set exact rate",-0.081,"rate","−8.1 points",56,"\u0001","Calculation. Tuned 90% (82% to 95%, n = 82); holdout 82% (70% to 90%, n = 56). The intervals overlap."],["holdout-gap-claude-haiku","Claude Haiku 4.5 · Claude Code: holdout minus tuned-set exact rate",-0.1045,"rate","−10.5 points",56,"\u0001","Calculation. Tuned 89% (80% to 94%, n = 82); holdout 79% (66% to 87%, n = 56). The intervals overlap."],["holdout-gap-claude-sonnet","Claude Sonnet 5.5 (low) · Claude Code: holdout minus tuned-set exact rate",-0.064,"rate","−6.4 points",56,"\u0001","Calculation. Tuned 94% (87% to 97%, n = 82); holdout 88% (76% to 94%, n = 56). The intervals overlap."],["holdout-jev-cost-per-1000","Jev 1.13 (TypeSafe): cost per 1,000 unseen decisions",0.03065,"usd","$0.0307",168,"\u0001","Calculation: reported input tokens × $0.042 per million."]]},{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","sourceIds"],"$r":[["routing-holdout-exact","Unseen routing decisions answered exactly right","Share of the 56 holdout cases where every scored question was acceptable","dot-range","rate","higher","Exact",[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.","ci95",["agent-routing-holdout"]],["routing-holdout-key-accuracy","Per-question accuracy on unseen decisions","\u0001","dot-range","\u0001","higher","\u0001",[],"\u0001","\u0001",["agent-routing-holdout"]],["routing-holdout-by-purpose","Exact rate on unseen decisions, by decision type","\u0001","grouped-bar","\u0001","higher","\u0001",[],"\u0001","\u0001",["agent-routing-holdout"]],["routing-holdout-tuned-vs-unseen","Tuned case set vs unseen holdout: exact rate per router","\u0001","grouped-bar","\u0001","higher","\u0001",[],"\u0001","\u0001",["agent-routing-holdout","agent-routing"]],["routing-holdout-latency","Time per routing decision, by route","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001","\u0001",["agent-routing-holdout"]],["routing-holdout-cost-per-1000","Cost per 1,000 unseen routing decisions","\u0001","bar","\u0001","\u0001","\u0001",[],"\u0001","\u0001",["agent-routing-holdout","calc-repricing","price-jev","price-anthropic"]]]},{"$k":["id","columns","rows"],"$r":[["routing-holdout-pairs",[],[]],["routing-holdout-label-agreement",[],[]],["routing-holdout-tuned-by-type",[],[]],["routing-holdout-cases",[],[]],["routing-holdout-routers",[],[]]]},"\u0001"],["thinking-token-bill","How much of an AI bill is thinking? Reasoning tokens by model and effort","Thinking tokens by model and effort: share and cost","Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.","Across 378 recorded calls of Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol, how many output tokens come from reasoning (thinking)? What do they cost at list price per call and per strict pass, and do they track time? The calls cover eight hard tasks, five short tasks and several efforts.","On eight hard tasks, reasoning tokens were 46% to 92% of the output tokens of a median call (16 to 24 calls per configuration).","2026-10-06","2026-10-06",["thought-experiment","calculation","reasoning-tokens","thinking-tokens","effort","llm-pricing","claude-haiku","claude-sonnet","claude-opus","claude-fable","gpt-6-1-sol","claude-code","codex-cli","latency"],[],["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],{"$k":["id","label","value","unit","display","n","note"],"$r":[["thinking-bill-calls","Recorded calls analysed (no new calls)",378,"count","378 (152 hard head-to-head, 96 new effort-ladder, 130 five-task head-to-head)",378,"\u0001"],["thinking-bill-zero-reasoning","Calls that reported 0 reasoning tokens (kept as 0, not \"not reported\")",98,"count","98 of 378; 0 calls had no usable count",378,"\u0001"],["thinking-bill-sum-claude","Sum check, Claude Code: output tokens gained per extra reasoning token, same task and model (calculation)",0.993,"ratio","0.99 (leave-one-task-out range 0.98 to 1.01; 290 calls)",290,"About 1 is consistent with reasoning inside output. This correlation does not prove how a CLI counts tokens."],["thinking-bill-sum-codex","Sum check, Codex CLI: output tokens gained per extra reasoning token, same task and model (calculation)",0.967,"ratio","0.97 (leave-one-task-out range 0.97 to 0.99; 88 calls)",88,"About 1 is consistent with reasoning inside output. This correlation does not prove how a CLI counts tokens."],["thinking-bill-share-highest","Highest median reasoning share of output tokens, hard tasks (calculation)",0.9168,"rate","92% (Claude Haiku 4.5 · Claude Code; range 76% to 99%)",24,"\u0001"],["thinking-bill-share-lowest","Lowest median reasoning share of output tokens, hard tasks (calculation)",0.4633,"rate","46% (GPT-6.1 Sol (medium) · Codex CLI; range 11% to 87%)",16,"\u0001"],["thinking-bill-reasoning-cost-highest","Highest mean list-price cost of reasoning per call, hard tasks (calculation)",0.053696,"usd","$0.0537 (Claude Fable 5.1 · Claude Code; call range $0.0037 to $0.2944)",24,"\u0001"],["thinking-bill-reasoning-cost-lowest","Lowest mean list-price cost of reasoning per call, hard tasks (calculation)",0.002273,"usd","$0.0023 (GPT-6.1 Sol (medium) · Codex CLI; call range $0.0006 to $0.0084)",16,"\u0001"],["thinking-bill-cost-share-highest","Highest pooled reasoning share of total list-price cost, hard tasks (calculation)",0.7953,"rate","80% (Claude Haiku 4.5 · Claude Code; call range 40% to 90%)",24,"\u0001"],["thinking-bill-cost-share-lowest","Lowest pooled reasoning share of total list-price cost, hard tasks (calculation)",0.0887,"rate","9% (GPT-6.1 Sol (medium) · Codex CLI; call range 2% to 20%)",16,"\u0001"],["thinking-bill-fable-vs-sonnet","Reasoning cost per call, Fable 5.1 ÷ Sonnet 5.5, hard tasks (calculation, means)",8.056,"ratio","8.1x ($0.0537 vs $0.0067)",24,"\u0001"],["thinking-bill-haiku-failed-share","Share of Haiku 4.5 reasoning cost spent on calls that did not pass (calculation)",0.6114,"rate","61% (13 of 24 calls did not pass strictly)",24,"\u0001"],["thinking-bill-effort-tasks-sonnet-5-5","Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Sonnet 5.5 · Claude Code (paired, same tasks; calculation)",7,"count","7 of 8 tasks",8,"Each task ran 2 times at each effort. Eight tasks is a small sample; this counts tasks, it is not an interval."],["thinking-bill-effort-tasks-opus-5-5","Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Opus 5.5 · Claude Code (paired, same tasks; calculation)",8,"count","8 of 8 tasks",8,"Each task ran 2 times at each effort. Eight tasks is a small sample; this counts tasks, it is not an interval."],["thinking-bill-effort-tasks-gpt-6-1-sol","Tasks where mean reasoning tokens per call were higher at high than at low effort, GPT-6.1 Sol · Codex CLI (paired, same tasks; calculation)",8,"count","8 of 8 tasks",8,"Each task ran 2 times at each effort. Eight tasks is a small sample; this counts tasks, it is not an interval."],["thinking-bill-effort-ratio-sonnet-5-5","Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means)",2.17,"ratio","2.2x ($0.0043 at low, $0.0094 at high)",16,"\u0001"],["thinking-bill-effort-ratio-opus-5-5","Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means)",3.584,"ratio","3.6x ($0.0050 at low, $0.0180 at high)",16,"\u0001"],["thinking-bill-effort-ratio-gpt-6-1-sol","Reasoning cost per call, high ÷ low effort, GPT-6.1 Sol · Codex CLI (calculation, means)",3.366,"ratio","3.4x ($0.0012 at low, $0.0041 at high)",16,"\u0001"],["thinking-bill-time-rho-claude","Spearman, reasoning tokens vs total time, per Claude model: lowest (calculation)",0.852,"score","0.85 to 0.98 across 4 Claude models",200,"The value is the lowest of the per-model correlations; the display gives the range."],["thinking-bill-time-slope-claude","Seconds per 1,000 reasoning tokens on the same task, per Claude model: lowest (calculation)",7.649,"seconds","7.6 to 14.5 s across 4 Claude models",200,"The value is the lowest of the per-model slopes; the display gives the range."],["thinking-bill-time-rho-codex","Spearman, reasoning tokens vs total time, GPT-6.1 Sol in Codex CLI (calculation)",0.556,"score","0.56 (leave-one-task-out range 0.34 to 0.65; 48 calls)",48,"The range omits one whole task at a time. It is a sensitivity check, not a confidence interval."]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","whisker","polarity","sourceIds"],"$r":[["thinking-bill-share","Reasoning share of output tokens per call on hard tasks (calculation)","Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)","bar","percent","Reasoning share of output tokens (%)",[{"name":"Median call","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16]]}}],"Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.","minmax","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]],["thinking-bill-cost-per-call","List-price cost per call: reasoning, remaining output and input (calculation)","\u0001","stacked-bar","\u0001","\u0001",[],"\u0001","\u0001","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]],["thinking-bill-by-effort","Reasoning cost per strict pass by effort, with the total (calculation)","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001","\u0001","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]],["thinking-bill-vs-time","Reasoning tokens vs total time per call (calculation)","\u0001","scatter","\u0001","\u0001",[],"\u0001","\u0001","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]],["thinking-bill-short-vs-hard","Reasoning share on short tasks vs hard tasks (calculation)","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001","\u0001","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]]]},{"$k":["id","columns","rows"],"$r":[["thinking-bill-hard-cells",[],[]],["thinking-bill-effort-cells",[],[]],["thinking-bill-short-cells",[],[]],["thinking-bill-time-link",[],[]],["thinking-bill-sum-check",[],[]]]},{"statIds":["thinking-bill-share-highest","thinking-bill-share-lowest"]}],["llm-speed-anatomy","Where the seconds go: first text, output speed and prompt size for 6 LLMs","LLM speed: first text and tokens per second through CLIs","60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.","For 6 models run through their own coding CLIs (Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low)), how long until the first text, how fast does text stream after it, and what does a longer prompt add?","We measured CLI start-up, first text and total time on one 250-line answer and a ledger lookup at three prompt sizes.","2026-10-07","2026-10-07",["latency","time-to-first-token","tokens-per-second","output-speed","prompt-size","claude-code","codex-cli","claude-haiku","claude-sonnet","claude-opus","claude-fable","gpt-6-1-sol","gpt-6-luna"],[],["agent-speed-anatomy"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["speed-anatomy-calls","Counted calls in the speed anatomy study",60,"calls","60 calls (24 in part A, 36 in part B; 43 through Claude Code, 17 through Codex CLI)",60,"\u0001","\u0001"],["speed-anatomy-part-a-exact","Part A replies that matched all 250 lines exactly (strict)",0.9583,"rate","96% (23/24)",24,[0.7976,0.9926],"0 more were right but wrapped or laid out differently (format misses); 1 was wrong."],["speed-anatomy-lookup-exact","Part B lookups answered exactly (strict)",0.8889,"rate","89% (32/36)",36,[0.7469,0.9559],"0 more were right with extra words (format misses); 4 were wrong."],["speed-anatomy-token-count-ratio","The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation)",1.8,"ratio","1.8x",23,"\u0001","Median visible tokens of exact replies per model (observed token range; n): Haiku 1,210 (1,210 to 1,210; n = 3). Sonnet 1,941 (1,941 to 1,941; n = 4). Opus 1,941 (1,941 to 1,941; n = 4). Fable 1,941 (1,941 to 1,941; n = 4). Sol (low) 1,066 (1,066 to 1,066; n = 4). Luna (low) 1,066 (1,066 to 1,066; n = 4). Each vendor counts the same text differently, so tokens per second do not compare across vendors. A calculation, not a run."],["speed-anatomy-lookup-last-line","Part B replies whose last line was exactly the right number (post-hoc reading, not a pass)",1,"rate","100% (36/36)",36,[0.9036,1],"A reading declared after the Claude part B batch (protocol Amendment 2), derived from the stored replies. The strict score above stays the headline; this one is never counted as a pass or a format miss."],["speed-anatomy-cache-reads-grew","Models whose cache reads grew with the ledger size",0,"count","0 of 4",36,"\u0001","A new ledger for every call. The descriptive threshold is the 1k-prompt maximum plus 10% and 128 tokens. Growth does not prove ledger reuse; the cache table lists every cell."],["speed-anatomy-size-cost-haiku","Haiku: extra time to first text at 64k vs 1k (calculation)",0.86,"seconds","+0.9 s",6,"\u0001","Difference of two medians (2.78 s minus 1.93 s); the ranges are 1.85 to 2.04 s and 2.45 to 2.89 s. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."],["speed-anatomy-size-cost-sonnet","Sonnet: extra time to first text at 64k vs 1k (calculation)",1.62,"seconds","+1.6 s",6,"\u0001","Difference of two medians (3.07 s minus 1.45 s); the ranges are 1.23 to 1.72 s and 1.38 to 3.61 s and overlap. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."],["speed-anatomy-size-cost-opus","Opus: extra time to first text at 64k vs 1k (calculation)",0.28,"seconds","+0.3 s",6,"\u0001","Difference of two medians (1.79 s minus 1.51 s); the ranges are 1.46 to 2.01 s and 1.72 to 3.72 s and overlap. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."],["speed-anatomy-size-cost-sol-low","Sol (low): extra time to first text at 64k vs 1k (calculation)",0.57,"seconds","+0.6 s",6,"\u0001","Difference of two medians (3.93 s minus 3.36 s); the ranges are 3.36 to 4.75 s and 3.42 to 4.38 s and overlap. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","whisker","sourceIds","polarity"],"$r":[["speed-anatomy-first-text","Time to first text: a 250-line answer, six models","Median of 4 calls per model; whiskers = fastest and slowest call","dot-range","seconds","Seconds to first text",[{"name":"Time to first text","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",4,2.84,6.38,4],["Claude Sonnet 5.5 · Claude Code",1.96,0.88,4.09,4],["Claude Opus 5.5 · Claude Code",1.97,1.7,2.35,4],["Claude Fable 5.1 · Claude Code",4.43,2.27,4.64,4],["GPT-6.1 Sol (low) · Codex CLI",3.52,2.75,4.42,4],["GPT-6 Luna (low) · Codex CLI",3.3,3.19,3.47,4]]}}],"Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.","minmax",["agent-speed-anatomy"],"\u0001"],["speed-anatomy-output-speed","Output speed after the first text: visible tokens per second (calculation)","\u0001","dot-range","\u0001","\u0001",[],"\u0001","\u0001",["agent-speed-anatomy"],"\u0001"],["speed-anatomy-chars-per-second","Output speed in characters per second after the first text (calculation)","\u0001","dot-range","\u0001","\u0001",[],"\u0001","\u0001",["agent-speed-anatomy"],"\u0001"],["speed-anatomy-prompt-size","Time to first text as the prompt grows","\u0001","line","\u0001","\u0001",[],"\u0001","\u0001",["agent-speed-anatomy"],"\u0001"],["speed-anatomy-total-by-size","Total time per call by prompt size","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001","\u0001",["agent-speed-anatomy"],"\u0001"],["speed-anatomy-lookup-correct","Exact lookup answers at the 1k, 16k and 64k prompt-size targets","\u0001","dot-range","\u0001","\u0001",[],"\u0001","\u0001",["agent-speed-anatomy"],"higher"]]},[{"id":"speed-anatomy-cells","columns":[],"rows":[]},{"id":"speed-anatomy-cache-reads","columns":[],"rows":[]}],{"statIds":["speed-anatomy-token-count-ratio"]}],["haiku-retry-or-escalate","Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts","Haiku retry and escalate: cost per correct answer","Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.","On the 8 hard tasks, what does one correct answer cost, and how long does it take, if you try Claude Haiku 4.5 first and retry or escalate, against Claude Sonnet 5.5 every time?","The calculated Haiku-first policies cost more on these 8 hard tasks.","2026-10-07","2026-10-07",["thought-experiment","claude-haiku","claude-sonnet","llm-pricing","cost-per-correct-answer","retry","escalation","hard-tasks"],[],["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["retry-escalate-haiku-passes","Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks",0.4583,"rate","46% (11/24)",24,[0.2789,0.6493],"\u0001"],["retry-escalate-sonnet-passes","Strict passes, Claude Sonnet 5.5 · Claude Code, eight hard tasks",1,"rate","100% (24/24)",24,[0.862,1],"\u0001"],["retry-escalate-sonnet-low-passes","Strict passes, Claude Sonnet 5.5 (low) · Claude Code, eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"\u0001"],["retry-escalate-baseline-cost","Cost per correct answer, Sonnet 5.5 every time (calculation)",0.01322,"usd","$0.0132",24,"\u0001","\u0001"],["retry-escalate-baseline-time","Wall time per correct answer, Sonnet 5.5 every time (calculation)",8.24,"seconds","8.2 s",24,"\u0001","\u0001"],["retry-escalate-haiku-once-cost","Cost per correct answer, Haiku once (calculation)",0.06772,"usd","$0.0677",24,"\u0001","\u0001"],["retry-escalate-haiku-once-time","Wall time per correct answer, Haiku once (calculation)",91.47,"seconds","91.5 s",24,"\u0001","\u0001"],["retry-escalate-retry-cost","Cost per correct answer, Haiku with one retry then Sonnet (calculation)",0.05664,"usd","$0.0566",48,"\u0001","\u0001"],["retry-escalate-retry-time","Wall time per correct answer, Haiku with one retry then Sonnet (calculation)",71.54,"seconds","71.5 s",48,"\u0001","\u0001"],["retry-escalate-retry-cost-ratio","Haiku with one retry then Sonnet vs Sonnet every time: cost per correct answer (calculation)",4.28,"ratio","4.3x",48,"\u0001","\u0001"],["retry-escalate-retry-time-ratio","Haiku with one retry then Sonnet vs Sonnet every time: wall time per correct answer (calculation)",8.68,"ratio","8.7x",48,"\u0001","\u0001"],["retry-escalate-three-tries-cost-ratio","Haiku up to 3 tries then Sonnet vs Sonnet every time: cost per correct answer (calculation)",5.44,"ratio","5.4x",48,"\u0001","\u0001"],["retry-escalate-sonnet-low-cost","Cost per correct answer, Sonnet 5.5 low effort once (calculation)",0.01219,"usd","$0.0122",16,"\u0001","\u0001"],["retry-escalate-sonnet-low-time","Wall time per correct answer, Sonnet 5.5 low effort once (calculation)",7.07,"seconds","7.1 s",16,"\u0001","\u0001"],["retry-escalate-haiku-dearer-tasks","Tasks where Haiku’s median call cost more than Sonnet’s (calculation)",8,"count","8 of 8",8,"\u0001","\u0001"],["retry-escalate-haiku-slower-tasks","Tasks where Haiku’s median call took longer than Sonnet’s",8,"count","8 of 8",8,"\u0001","\u0001"],["retry-escalate-haiku-output-tokens","Median output tokens per call, Haiku 4.5 (Claude Code)",5064,"tokens","5,064",24,"\u0001","Call range 1,899 to 9,321 tokens; not a confidence interval."],["retry-escalate-sonnet-output-tokens","Median output tokens per call, Sonnet 5.5 (Claude Code)",1050,"tokens","1,050",24,"\u0001","Call range 176 to 3,895 tokens; not a confidence interval."],["retry-escalate-breakeven-cost","Haiku call cost, as a share of measured, at which one retry then Sonnet matches Sonnet every time (calculation)",0.1177,"ratio","11.8% of measured",8,"\u0001","All pass rates kept as observed. Arithmetic, not a prediction: a Haiku that thinks less may pass at a different rate."],["retry-escalate-breakeven-time","Haiku call time, as a share of measured, at which one retry then Sonnet matches Sonnet every time (calculation)",0.0543,"ratio","5.4% of measured",8,"\u0001","All pass rates kept as observed. Arithmetic, not a prediction."],["retry-escalate-sensitivity-lowest-ratio","Lowest cost ratio of any Haiku-first policy to Sonnet every time, across the four Haiku pass-rate settings (sensitivity range, calculation)",2.96,"ratio","3.0x",8,"\u0001","Haiku once, with Haiku’s pass rate on every task set to its Wilson upper bound."],["retry-escalate-joint-extreme-ratio","Lowest cost ratio of any Haiku-first policy to Sonnet every time, with Haiku’s pass rate at its Wilson upper bound and Sonnet’s at its Wilson lower bound on every task at once (joint extreme, calculation)",1.3,"ratio","1.3x",8,"\u0001","Haiku once. The lowest time ratio is 2.8x (Haiku once). In this setting Sonnet every time succeeds on 44% of the tasks, which no run showed (Sonnet passed 24/24). A joint extreme on 3 calls per task, not a forecast."],["retry-escalate-lenient-retry-cost","Cost per correct answer, Haiku with one retry then Sonnet, format misses counted as passes (calculation)",0.04591,"usd","$0.0459",48,"\u0001","\u0001"]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds","polarity"],"$r":[["retry-escalate-cost-per-correct","Expected list-price cost per correct answer, by retry policy (calculation)","Eight hard tasks, each equally likely. A failed try is paid for. Calls recorded in Claude Code","bar","usd","USD per correct answer",[{"name":"Cost per correct answer (calculation)","points":{"$k":["label","value","n","highlight"],"$r":[["Sonnet 5.5 low effort once (calculation)",0.01219,16,false],["Sonnet 5.5 every time (calculation)",0.01322,24,true],["Haiku, one retry, then Sonnet (calculation)",0.05664,48,false],["Haiku once (calculation)",0.06772,24,false],["Haiku up to 3 tries, then Sonnet (calculation)",0.07193,48,false]]}}],"Calculation, not a run. These are median-input scenarios, not measured mean costs. Total modelled list-price cost of all tries over the 8 tasks, divided by the expected number of correct answers. Per-call cost is each task’s median call at list price (reported tokens × list price); the calls ran on a flat subscription. A try runs only when the earlier tries failed the validator, and tries are independent at each task’s observed pass rate. n is the number of recorded calls behind each bar. A calculation has no confidence interval; the sensitivity chart shows a range. Highlighted: the baseline, Sonnet 5.5 every time.",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"\u0001"],["retry-escalate-time-per-correct","Expected wall time per correct answer, by retry policy (calculation)","\u0001","bar","\u0001","\u0001",[],"\u0001",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"\u0001"],["retry-escalate-success","Expected success rate by retry policy (calculation)","\u0001","bar","\u0001","\u0001",[],"\u0001",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"higher"],["retry-escalate-haiku-by-task","Strict pass rate of Haiku 4.5 on each hard task","\u0001","bar","\u0001","\u0001",[],"\u0001",["agent-provider-h2h-hard"],"higher"],["retry-escalate-call-cost-by-task","List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation)","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"\u0001"],["retry-escalate-sensitivity","Sensitivity: cost per correct answer if Haiku’s pass rate is higher or lower (calculation)","\u0001","dot-range","\u0001","\u0001",[],"\u0001",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"\u0001"]]},{"$k":["id","columns","rows"],"$r":[["retry-escalate-policies",[],[]],["retry-escalate-task-inputs",[],[]],["retry-escalate-sensitivity-table",[],[]]]},"\u0001"],["harder-tasks-head-to-head","GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks","GPT-6.1 Sol vs Claude Opus 5.5 on harder tasks","56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.","On a task set built so that Claude Sonnet 5.5 did not pass it every time, does pass rate separate GPT-6.1 Sol (Codex CLI) from Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 (Claude Code)?","22 of 56 counted calls passed strictly (39%, 95% Wilson interval 28% to 52%) on 4 tasks.","2026-10-07","2026-10-07",["head-to-head","harder-tasks","gpt-6-1-sol","codex-cli","claude-sonnet","claude-opus","claude-haiku","reasoning","selection-effect","format-misses","latency"],[],["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"],{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["harder-h2h-pass-all","Counted calls that passed strictly (harder set)",0.3929,"rate","39% (22/56)",56,[0.2758,0.5237],"\u0001"],["harder-h2h-correct-all","Counted calls with a correct answer, format misses included (lenient reading)",0.4107,"rate","41% (23/56)",56,[0.2917,0.5412],"\u0001"],["harder-h2h-pass-sol","GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks",0.6875,"rate","69% (11/16)",16,[0.444,0.8584],"\u0001"],["harder-h2h-pass-opus","Opus 5.5 (Claude Code): strict pass rate on the harder tasks",0.4167,"rate","42% (5/12)",12,[0.1933,0.6805],"\u0001"],["harder-h2h-pass-sonnet","Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks",0.375,"rate","38% (6/16)",16,[0.1848,0.6136],"\u0001"],["harder-h2h-pass-haiku","Haiku 4.5 (Claude Code): strict pass rate on the harder tasks",0,"rate","0% (0/12)",12,[0,0.2425],"\u0001"],["harder-h2h-format-misses","Non-passes that were format misses, not wrong answers",1,"count","1 of 34 non-passes (21 wrong answers, 12 no answer)",34,"\u0001","\u0001"],["harder-h2h-tool-attempts","Counted calls that tried a tool although tools were off (a behaviour, not a quality score)",0.1964,"rate","20% (11/56)",56,[0.1134,0.3184],"Per configuration: GPT-6.1 Sol (medium) 0 of 16, Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12. None of these calls passed. The flag comes from the reply text: tool-call markup, or a CLI message that a tool call could not be parsed. The Codex runner also tells the model not to call tools; the Claude Code runner does not."],["harder-h2h-timeouts","Counted calls that ran past the 300 s limit and gave no answer",0.1786,"rate","18% (10/56)",56,[0.1,0.2984],"Per configuration: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12, Sonnet 5.5 4 of 16 and Haiku 4.5 2 of 12. A timeout counts as a non-pass."],["harder-h2h-pilot-sonnet","Pilot (not scored): strict passes of Claude Sonnet 5.5 on all candidate tasks",0.8125,"rate","81% (26/32)",32,[0.6469,0.9111],"Two calls per candidate. Its rates have a selection effect and repeated-task dependence. The pilot decided which tasks were kept; it is not part of any counted cell."],["harder-h2h-pilot-sonnet-kept","Pilot (not scored): strict passes of Claude Sonnet 5.5 on the tasks that were kept",0.25,"rate","25% (2/8)",8,[0.0715,0.5907],"The same four tasks the counted calls use. The counted Sonnet rate is in harder-h2h-pass-sonnet. Selection against Sonnet predicts a lower pilot rate than a fresh rate."],["harder-h2h-tasks-kept","Candidate tasks kept for the counted set",4,"count","4 of 16 (Sonnet passed 12 candidates 2 of 2, including 10 of 10 code, SQL, spec, numeric and simulation tasks)",16,"\u0001","\u0001"],["harder-h2h-cheapest-per-pass","Lowest recorded cost lower bound per strict pass (calculation)",0.08293,"usd","GPT-6.1 Sol (medium) · Codex CLI: $0.0829 (lower bound)",16,"\u0001","All passing configurations have unpriced timeout calls. Unknown costs can change their order. This is a recorded lower bound, not a cost ranking."]]},{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","whisker","series","note","sourceIds"],"$r":[["harder-h2h-pass-rate","Pass rate on 4 harder tasks","Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","dot-range","rate","higher","Passed","ci95",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.5,0.2538,0.7462,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}}],"Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row.\n\nCounted calls are new calls.",["agent-harder-tasks"]],["harder-h2h-outcomes","What happened on every call","\u0001","stacked-bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-harder-tasks"]],["harder-h2h-tool-attempts","Calls that tried a tool although tools were off","\u0001","dot-range","\u0001","none","\u0001","\u0001",[],"\u0001",["agent-harder-tasks"]],["harder-h2h-pass-by-task","Strict pass rate by task","\u0001","grouped-bar","\u0001","higher","\u0001","\u0001",[],"\u0001",["agent-harder-tasks"]],["harder-h2h-total-latency","Total time per call on harder tasks","\u0001","dot-range","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-harder-tasks"]],["harder-h2h-output-tokens","Output tokens per call on harder tasks","\u0001","grouped-bar","\u0001","none","\u0001","\u0001",[],"\u0001",["agent-harder-tasks"]],["harder-h2h-cost-per-pass","List-price cost per strict pass on harder tasks (calculation)","\u0001","bar","\u0001","\u0001","\u0001","\u0001",[],"\u0001",["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"]],["harder-h2h-frontier","Observed quality vs cost frontier (calculation)","\u0001","scatter","\u0001","higher","\u0001","\u0001",[],"\u0001",["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"]]]},[{"id":"harder-h2h-cells","columns":[],"rows":[]},{"id":"harder-h2h-tasks","columns":[],"rows":[]}],"\u0001"],["voice-agent-latency-budget","Voice agent latency budget: component calculations, not a measured turn","Voice agent latency budget: calculated component times","Calculation: component times against assumed 300 ms, 800 ms and 1,500 ms budgets. No voice turn or audio was measured.","Which decision steps fit inside a voice agent’s turn budget, using only measured times?","Rules and Jev 1.13 fit all three assumed budgets at the median and p95.","2026-10-06","2026-10-06",["voice-agent","latency","latency-budget","routing","jev","thought-experiment","calculation"],[],["calc-latency-budget","agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"],{"$k":["id","label","value","unit","display","n","note"],"$r":[["latency-budget-steps-counted","Decision and first-output steps compared",14,"count","14 steps",14,"5 under 1,000 ms at the median, 9 at 1,000 ms or more. Every time is a measured median from another study."],["latency-budget-fit-300-slow","Steps that fit a 300 ms budget at the slow end (calculation)",3,"count","3 of 14 (median: 3 of 14)",14,"Budget (assumption: a budget a voice team might set). Slow end: the 95th percentile for steps with 30 or more runs, the slowest run for the rest."],["latency-budget-fit-800-slow","Steps that fit an 800 ms budget at the slow end (calculation)",3,"count","3 of 14 (median: 3 of 14)",14,"Budget (assumption: a budget a voice team might set). Slow end: the 95th percentile for steps with 30 or more runs, the slowest run for the rest."],["latency-budget-fit-1500-slow","Steps that fit a 1,500 ms budget at the slow end (calculation)",4,"count","4 of 14 (median: 8 of 14)",14,"Budget (assumption: a budget a voice team might set). Slow end: the 95th percentile for steps with 30 or more runs, the slowest run for the rest."],["latency-budget-jev-share-300","Share of a 300 ms budget that one Jev 1.13 decision uses at the median (calculation)",45.5,"percent","46% at the median, 65% at p95",246,"Median 136.5 ms, p95 195.7 ms over HTTPS from one Mac on a home network. Budget (assumption: a budget a voice team might set)."],["latency-budget-jev-in-sequence-800","Back-to-back Jev 1.13 decisions that fit an 800 ms budget at the slow end (calculation)",4,"count","4 (5 at the median)",246,"Budget (assumption: a budget a voice team might set)."],["latency-budget-sonnet-router-share-1500","Share of a 1,500 ms budget that one Sonnet 5.5 router decision uses at the median (calculation)",173,"percent","173% at the median, 287% at p95",82,"Median 2,598 ms, p95 4,298 ms through the Claude Code CLI. Budget (assumption: a budget a voice team might set)."],["latency-budget-fastest-model-first-output","Lowest observed API median time to the first output of a one-line answer",823,"ms","823 ms (slowest run 1,371 ms)",5,"GPT-6 Luna through the OpenAI API. Range 506 ms to 1,371 ms, n = 5; not a confidence interval. API ranges overlap, so this is not a ranking."],["latency-budget-codex-cli-added","Difference in median first-output time across 3 Codex model setups: Codex CLI minus API (calculation)",1964,"ms","+1,964 to +2,880 ms",5,"Median through the Codex CLI minus median through the OpenAI API, for the same model, effort and one-line prompt (GPT-6 Luna: +1,964 ms; GPT-6.1 Sol (effort low): +2,880 ms; GPT-6.1 Sol (effort high): +2,445 ms). For each setup the slowest API run is below the CLI median. A calculation across runs of 5 per setup; the samples are small."],["latency-budget-jev-model-left-1500","Time left of a 1,500 ms budget after Jev 1.13 and the API cell with the lowest observed median, at the median (calculation)",540.5,"ms","540.5 ms left (sum of medians 959.5 ms)",5,"Separate samples: Jev n = 246; API n = 5. Jev median 136.5 ms plus GPT-6 Luna median 823 ms through the OpenAI API. At the slow ends the sum is 1,566.7 ms, 66.7 ms over. This arithmetic remainder is not measured audio latency. API timings already include the client-to-provider network. Budget (assumption: a budget a voice team might set). A calculation across runs that were not made together."],["latency-budget-jev-cold-call","First Jev call minus the later-call median (calculation)",88.3,"ms","88.3 ms once",1,"The first call used a fresh connection and took 224.7 ms; the later-call median was 136.4 ms (n = 245). One first call, so no range. This difference does not isolate connection cost or server warmth."]]},{"$k":["id","title","subtitle","kind","unit","yLabel","whisker","series","note","sourceIds"],"$r":[["latency-budget-fast-steps","Steps that take under a second, against three assumed budgets (median to p95)","Median per step; whiskers = median to p95; budget dots are assumptions","dot-range","ms","Time per step","p50-p95",[{"name":"Time per step","points":{"$k":["label","value","lo","hi","n"],"$r":[["Deterministic routing policy (Agent, in process)",0.00142,0.00142,0.00233,20000],["Rule-based System One decision (record write excluded)",1,1,2,419],["Jev 1.13 (TypeSafe, direct HTTPS)",136.5,136.5,195.7,246]]}},{"name":"Budget (assumption: a budget a voice team might set)","points":{"$k":["label","value"],"$r":[["Budget 300 ms",300],["Budget 800 ms",800],["Budget 1,500 ms",1500]]}}],"These 5 steps have a median under 1,000 ms. Decisions are whole calls; model steps are the time to the first useful output of a one-line answer. The dot is the median. The whisker runs from the median to the 95th percentile (30 or more runs per step). Observed limits stay in the measured-step table; some minima were not retained. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]],["latency-budget-fast-steps-ranges","Steps that take under a second, against three assumed budgets (observed ranges)","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]],["latency-budget-slow-steps","Steps that take a second or more, against three assumed budgets (median to p95)","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]],["latency-budget-slow-steps-ranges","Steps that take a second or more, against three assumed budgets (observed ranges)","\u0001","dot-range","\u0001","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]],["latency-budget-fit","How many of the steps fit each assumed budget (calculation)","\u0001","grouped-bar","\u0001","\u0001","\u0001",[],"\u0001",["calc-latency-budget","agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]]]},{"$k":["id","columns","rows"],"$r":[["latency-budget-measured-steps",[],[]],["latency-budget-fits",[],[]],["latency-budget-shares",[],[]],["latency-budget-in-sequence",[],[]],["latency-budget-router-plus-model",[],[]]]},"\u0001"],["routing-at-scale","What does routing a million AI requests a day cost? A calculation from measured runs","LLM router cost at scale: 1 million decisions a day","A calculation from measured runs: what 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude, with median/p95 time scenarios.","At 10,000 to 10 million routing decisions a day, what does each router cost per day, what in-flight and waiting scenarios do median and p95 times give?","This is a calculation, not a run.","2026-10-06","2026-10-06",["thought-experiment","calculation","routing","llm-router","jev","latency","cost","scale","littles-law"],[],["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"],{"$k":["id","label","value","unit","display","n","note"],"$r":[["routing-at-scale-cost-per-1000-jev","Cost per 1,000 routing decisions, Jev (calculation)",0.0337,"usd","$0.0337",82,"Calculation: provider-reported total cost ÷ counted calls × 1,000. This fixes the recorded token mix; no cost interval is available."],["routing-at-scale-cost-per-1000-sonnet","Cost per 1,000 routing decisions, Sonnet 5.5 router (calculation)",7.324,"usd","$7.32",82,"A list-price calculation on the tokens the CLI reported. The calls ran on a subscription."],["routing-at-scale-cost-per-1000-haiku","Cost per 1,000 routing decisions, Haiku 4.5 router (calculation)",8.924,"usd","$8.92",82,"A list-price calculation on the tokens the CLI reported. The calls ran on a subscription."],["routing-at-scale-system-one-share","System One decisions as a share of model calls (ratio of the two medians, calculation)",0.1414,"rate","14.1%",48,"7 ÷ 49.5. A ratio of medians, not a sample proportion, so it carries no interval."],["routing-at-scale-rate-1m","Decisions per second at 1 million a day (calculation)",11.5741,"count","11.57 per second","\u0001","A steady rate over 24 hours."],["routing-at-scale-cost-1m-policy","Daily cost at 1 million decisions, rule-based policy (calculation)",0,"usd","$0",20000,"$0 a year at a flat volume."],["routing-at-scale-cost-1m-jev","Daily cost at 1 million decisions, Jev (calculation)",33.7,"usd","$33.70",82,"$12,301 a year at a flat volume."],["routing-at-scale-cost-1m-sonnet","Daily cost at 1 million decisions, Sonnet 5.5 router (calculation)",7324,"usd","$7,324",82,"$2,673,260 a year at a flat volume."],["routing-at-scale-cost-1m-haiku","Daily cost at 1 million decisions, Haiku 4.5 router (calculation)",8924,"usd","$8,924",82,"$3,257,260 a year at a flat volume."],["routing-at-scale-flight-1m-jev","Decisions in flight at 1 million a day, Jev (calculation)",1.5799,"calls","1.58 (2.27 at the 95th percentile time)",246,"Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."],["routing-at-scale-flight-1m-sonnet","Decisions in flight at 1 million a day, Sonnet 5.5 router (calculation)",30.069,"calls","30.1 (49.7 at the 95th percentile time)",82,"Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."],["routing-at-scale-flight-1m-haiku","Decisions in flight at 1 million a day, Haiku 4.5 router (calculation)",146.69,"calls","146.7 (398.3 at the 95th percentile time)",82,"Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."],["routing-at-scale-hours-1m-policy","Waiting hours a day at 1 million decisions, rule-based policy (calculation)",0.00039444,"count","1.42 s in total",20000,"Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."],["routing-at-scale-hours-1m-jev","Waiting hours a day at 1 million decisions, Jev (calculation)",37.917,"count","37.9 hours (54.4 at the 95th percentile time)",246,"Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."],["routing-at-scale-hours-1m-sonnet","Waiting hours a day at 1 million decisions, Sonnet 5.5 router (calculation)",721.67,"count","721.7 hours (1,194 at the 95th percentile time)",82,"Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."]]},{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["routing-at-scale-daily-cost","Daily cost of routing at 10,000 to 10 million decisions a day (calculation)","Decisions a day × cost per decision, USD per day, every decision routed. The rule-based policy costs $0 and is not on the log axis","grouped-bar","usd","USD per day",{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","n"],"$r":[["10,000 a day",0.337,82],["100,000 a day",3.37,82],["1 million a day",33.7,82],["10 million a day",337,82]]}],["Claude Sonnet 5.5 (low) · Claude Code",{"$k":["label","value","n"],"$r":[["10,000 a day",73.24,82],["100,000 a day",732.4,82],["1 million a day",7324,82],["10 million a day",73240,82]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","n"],"$r":[["10,000 a day",89.24,82],["100,000 a day",892.4,82],["1 million a day",8924,82],["10 million a day",89240,82]]}]]},"This is a calculation, not a run. Daily cost = decisions a day × cost per decision. Cost per 1,000 decisions: Jev $0.0337 (calculation from provider-reported run cost, n = 82), Sonnet 5.5 $7.32, Haiku 4.5 $8.92. The Claude figures use one-hour cache-write rates and list price × the tokens the CLI reported (Sonnet n = 82, Haiku n = 82). \n\nThe rule-based policy makes no model call, so it costs $0 at every volume and has no bar: a log axis cannot show zero. The chart prices model calls only, not servers, retries or the work itself.",["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev"]],["routing-at-scale-in-flight","In-flight scenarios at 10,000 to 10 million a day (calculation)","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["calc-routing-at-scale","agent-routing-overhead","agent-routing","agent-jev-live"]],["routing-at-scale-waiting-hours","Waiting scenarios per day at 1 million decisions a day (calculation)","\u0001","bar","\u0001","\u0001",[],"\u0001",["calc-routing-at-scale","agent-routing-overhead","agent-routing","agent-jev-live"]],["routing-at-scale-scenarios","Daily cost at 1 million model calls a day: route every call or only System One decisions (calculation)","\u0001","grouped-bar","\u0001","\u0001",[],"\u0001",["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev"]]]},{"$k":["id","columns","rows"],"$r":[["routing-at-scale-inputs",[],[]],["routing-at-scale-formulas",[],[]],["routing-at-scale-context",[],[]],["routing-at-scale-grid",[],[]]]},"\u0001"]]}