{"schema":"agent-public-bench@1","generatedAt":"2026-10-07T00:00:00.000Z","sources":{"$k":["id","title","kind","date","note","data","url"],"$r":[["agent-swebench-c1","Agent on SWE-bench Verified, campaign 1 (25 instances)","run","2026-10-04","Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.",["/benchmarks/raw/swebench/attempts.json","/benchmarks/raw/swebench/exclusions.json"],"\u0001"],["agent-swebench-c2","Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)","run","2026-10-05","The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.",["/benchmarks/raw/swebench/attempts.json"],"\u0001"],["swebench-leaderboard","SWE-bench Verified leaderboard, mini-SWE-agent v2 runs","public-leaderboard","2026-02-17","Public per-instance results of 11 models under mini-SWE-agent 2.0.0 (bash only, one attempt). Costs are API list prices as published.",["/benchmarks/raw/swebench/panel.json"],"https://www.swebench.com"],["swebench-protocol","SWE-bench campaign rules and sample design","protocol","2026-10-04","Rules declared before the first run: escalations are graded as delivered, gold must resolve on the host, blocked instances are replaced in the same difficulty band, no second attempts. The excluded and replaced instances are listed in the exclusions extract.",["/benchmarks/raw/swebench/exclusions.json"],"\u0001"],["agent-provider-explorer","Provider explorer receipts: CLI vs API","run","2026-10-03","230 imported receipts for short fixed tasks over Claude Code CLI, Codex CLI and the OpenAI API, with time to first useful output, total time, tokens and validation.",["/benchmarks/raw/provider-explorer/runs.json"],"\u0001"],["agent-provider-h2h-hard","Provider head-to-head, hard set: eight hard tasks with strict validators","run","2026-10-06","Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.",["/benchmarks/raw/provider-h2h-hard/receipts.json"],"\u0001"],["agent-effort-ladder","Effort ladder: the hard task set at each effort level","run","2026-10-06","The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.",["/benchmarks/raw/effort-ladder/receipts.json","/benchmarks/raw/provider-h2h-hard/receipts.json"],"\u0001"],["agent-memory-study","Agent memory study: 8 kinds of project memory on Claude Code","run","2026-10-06","One small Node.js repository, 5 tasks, 8 memory conditions (none, /init, curated, raw notes, dreamed notes, long handbook, Stop hook, curated + hook). Claude Code 2.1.286 headless: Sonnet lane 3 repetitions, Haiku lane 2. Hidden tests and deterministic convention checks; protocol declared before the first session; every attempt kept. The exact memory files and three dreaming passes are published.",["/benchmarks/raw/agent-memory/attempts-sonnet.json","/benchmarks/raw/agent-memory/attempts-haiku.json","/benchmarks/raw/agent-memory/dreams.json","/benchmarks/raw/agent-memory/memory-files.json"],"\u0001"],["agent-caching-consistency","Caching sessions and repeated prompts (Claude Code and Codex CLI)","run","2026-10-06","Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.",["/benchmarks/raw/caching-consistency/caching.json","/benchmarks/raw/caching-consistency/consistency.json"],"\u0001"],["agent-routing","Routing runs: Jev router vs LLM routing","run","2026-10-05","Routing decisions recorded per case and arm.",["/benchmarks/raw/routing/receipts.json"],"\u0001"],["agent-jev-live","Jev live run: 246 timed calls on the 82 routing decisions","run","2026-10-06","Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.",["/benchmarks/raw/jev-live/summary.json","/benchmarks/raw/jev-live/calls.json"],"\u0001"],["price-anthropic","Anthropic list prices (Claude models)","price-list","2026-09-21","Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.","\u0001","https://platform.claude.com/docs/en/about-claude/pricing"],["price-openai","OpenAI list prices","price-list","2026-10-03","Token prices as listed by the vendor on 2026-10-03.","\u0001","https://developers.openai.com/api/docs/pricing"],["price-jev","Jev 1.13 list price","price-list","2026-09-23","Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.","\u0001","https://docs.typesafe.ai/models"],["agent-routing-overhead","Routing overhead runs: policy microbenchmark and CLI start-up","run","2026-10-06","In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.",["/benchmarks/raw/routing-overhead/results.json"],"\u0001"],["calc-routing-overhead","Routing overhead per 1,000 tasks (calculation)","calculation","2026-10-06","Decisions per task from recorded bench runs multiplied by the cost and the median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound. A calculation, not a run.",["/benchmarks/raw/routing-overhead/results.json"],"\u0001"],["calc-cache-pricing","Cost with and without the prompt cache (calculation)","calculation","2026-10-06","Recorded tokens per turn × Anthropic list prices. With the cache: uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at twice the input price, 5-minute writes at 1.25 times (an assumption; none occurred). Without a cache: every input token at the input price. Output is priced the same in both. Not a bill.",["/benchmarks/raw/caching-consistency/caching.json"],"\u0001"],["calc-repricing","Repricing calculation","calculation","2026-10-05","Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.","\u0001","\u0001"]]},"studies":{"$k":["slug","title","seoTitle","description","question","answer","date","updated","tags","method","caveats","sourceIds","stats","charts","tables","related","hero"],"$r":[["swe-bench-verified","Agent on SWE-bench Verified vs 11 public models","SWE-bench Verified: Agent vs GPT, Claude and Gemini","","","","2026-10-05","2026-10-05",[],[],[],["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol","price-anthropic"],[{"id":"agent-rate-33","label":"Agent resolved, all 33 attempted instances","value":0.7576,"unit":"rate","display":"76% (25/33)","n":33,"ci":[0.5898,0.8717],"note":"Both campaigns, one attempt each, failures and empty patches included."}],[],[],[],"\u0001"],["hard-model-head-to-head","Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks","Hard tasks: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol","","","","2026-10-06","2026-10-06",[],[],[],["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],[{"id":"hard-h2h-pass-all","label":"Calls that passed strictly (hard set)","value":0.9145,"unit":"rate","display":"91% (139/152)","n":152,"ci":[0.8592,0.9493]},{"id":"hard-h2h-correct-all","label":"Calls with a correct answer, format misses included (lenient reading)","value":0.9474,"unit":"rate","display":"95% (144/152)","n":152,"ci":[0.8996,0.9731]}],[{"id":"hard-h2h-pass-rate","title":"Pass rate on eight hard tasks","subtitle":"Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.4583,0.2789,0.6493,24]]}}],"note":"Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.","sourceIds":["agent-provider-h2h-hard"]},{"id":"hard-h2h-cost-per-pass","title":"List-price cost per strict pass on hard tasks (calculation)","subtitle":"All calls in a configuration, failures and format misses included, divided by its strict passes","kind":"bar","unit":"usd","yLabel":"USD per strict pass","series":[{"name":"Cost per strict pass","points":[{"label":"Claude Sonnet 5.5 · Claude Code","value":0.01435,"n":24,"highlight":true},{"label":"Claude Opus 5.5 · Claude Code","value":0.02824,"n":24,"highlight":false}]}],"note":"Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.","sourceIds":["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]}],[],[],"\u0001"],["effort-ladder","Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks","Effort ladder: does more AI effort buy quality?","","","","2026-10-06","2026-10-06",[],[],[],["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"],[],[{"id":"effort-ladder-pass-rate","title":"Strict pass rate by effort on eight hard tasks","subtitle":"Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","yLabel":"Passed","series":[{"name":"Strict pass","points":[{"label":"Claude Sonnet 5.5 (low) · Claude Code","value":1,"lo":0.8064,"hi":1,"n":16},{"label":"Claude Sonnet 5.5 (high) · Claude Code","value":1,"lo":0.8064,"hi":1,"n":16}]}],"note":"Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.","whisker":"ci95","sourceIds":["agent-effort-ladder"]}],[],[],"\u0001"],["caching-consistency","Prompt caching and run-to-run consistency in Claude Code and Codex CLI","Prompt caching savings and LLM consistency, measured","","","","2026-10-06","2026-10-06",[],[],[],["agent-caching-consistency","calc-cache-pricing","price-anthropic"],[],{"$k":["id","title","subtitle","kind","unit","xLabel","yLabel","series","note","sourceIds","whisker"],"$r":[["caching-read-share-by-turn","Share of input read from the cache, by turn in a session","Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each","line","rate","Turn in the session","Input tokens read from cache",[{"name":"Claude Sonnet 5.5 · Claude Code","points":[{"label":"Turn 2","value":0.9924,"n":3},{"label":"Turn 5","value":0.9074,"n":3}]}],"Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.",["agent-caching-consistency"],"\u0001"],["caching-latency-first-vs-later","Time per turn: first turn vs later turns in a cached session","Median; whiskers = fastest and slowest turn","dot-range","seconds","\u0001","Seconds",[{"name":"Turn 1 (writes the ledger to the cache)","points":[{"label":"Claude Sonnet 5.5 · Claude Code","value":1.64,"lo":1.58,"hi":1.79,"n":3}]},{"name":"Turns 2-5 (read the ledger from the cache)","points":[{"label":"Claude Sonnet 5.5 · Claude Code","value":1.61,"lo":1.35,"hi":5.63,"n":12}]}],"Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.",["agent-caching-consistency"],"minmax"],["consistency-pass-rate","Same prompt, 10 times: strict pass rate","One series per prompt; whiskers are 95% Wilson intervals","dot-range","rate","\u0001","Passed",[{"name":"Exact number","points":[{"label":"Claude Haiku 4.5 · Claude Code","value":0,"lo":0,"hi":0.2775,"n":10},{"label":"Claude Sonnet 5.5 · Claude Code","value":1,"lo":0.7225,"hi":1,"n":10}]}],"Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.",["agent-caching-consistency"],"ci95"]]},[],[],"\u0001"],["agent-memory","Does memory help Claude Code? 8 kinds of agent memory, tested","Does CLAUDE.md help? Agent memory tested on Claude Code","","","","2026-10-06","2026-10-06",[],[],[],["agent-memory-study"],[{"id":"memory-team-knowledge-none","label":"Team-knowledge checks passed with no memory (Sonnet 5.5)","value":0.4,"unit":"rate","display":"40% (6/15)","n":15,"ci":[0.1982,0.6425]},{"id":"memory-team-knowledge-curated","label":"Team-knowledge checks passed with an 11-line curated file (Sonnet 5.5)","value":1,"unit":"rate","display":"100% (15/15)","n":15,"ci":[0.7961,1]}],[{"id":"memory-knowledge-class","title":"Where memory helps: what the repo shows vs what only the team knows","subtitle":"Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)","kind":"grouped-bar","unit":"rate","series":[{"name":"Rules the code already shows","points":[{"label":"Stop hook only","value":1,"lo":0.9398,"hi":1,"n":60}]},{"name":"Team knowledge only","points":[{"label":"Stop hook only","value":0.6667,"lo":0.4171,"hi":0.8482,"n":15}]}],"note":"Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.","sourceIds":["agent-memory-study"]},{"id":"memory-broken-test-command","title":"A stale README command: who still ran it?","polarity":"lower","subtitle":"Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.8,0.5481,0.9295,15],["/init CLAUDE.md",0.8667,0.6212,0.9626,15],["Curated, 11 lines",0,0,0.2039,15]]}}],"note":"The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old \"run npm test\" note and a later correction. The /init file repeats the README.","sourceIds":["agent-memory-study"]}],[],[],{"statIds":["memory-team-knowledge-none","memory-team-knowledge-curated"]}],["routing-overhead","Routing overhead: deterministic policy vs LLM routers vs Jev","Routing overhead: rules vs LLM routers vs Jev","","","","2026-10-06","2026-10-06",[],[],[],["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"],[],[{"id":"router-overhead-decision-latency","title":"Time to make one routing decision","subtitle":"Median; whiskers = median to 95th percentile","kind":"dot-range","unit":"ms","yLabel":"Time per decision","whisker":"p50-p95","series":[{"name":"Decision time","points":[{"label":"Deterministic routing policy (Agent, in process)","value":0.00142,"lo":0.00142,"hi":0.00233,"n":20000,"highlight":true},{"label":"Claude Sonnet 5.5 (effort low, via Claude Code)","value":2597,"lo":2597,"hi":4298,"n":82}]}],"note":"The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.","sourceIds":["agent-routing-overhead","agent-routing","agent-jev-live"]}],[],[],"\u0001"],["cli-model-latency-tokens","Claude Code CLI vs Codex CLI vs the API: latency and tokens","Claude Code vs Codex CLI vs API: latency and token overhead","","","","2026-10-03","2026-10-05",[],[],[],["agent-provider-explorer"],[],[{"id":"cli-vs-api-exact-reply-latency","title":"CLI vs API: time for a one-line answer","subtitle":"Matched cohort, fixed exact reply, 5 runs per configuration","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Total time","points":[{"label":"OpenAI API · GPT-6.1 Sol · low","value":1.02,"lo":0.96,"hi":1.87,"n":5},{"label":"Codex CLI · GPT-6.1 Sol · low","value":4.18,"lo":3.86,"hi":4.53,"n":5}]}],"note":"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.","sourceIds":["agent-provider-explorer"]}],[],[],"\u0001"]]}}