{"schema":"agent-public-bench@1","generatedAt":"2026-10-07T00:00:00.000Z","sources":{"$k":["id","title","kind","date","note","data","url"],"$r":[["agent-swebench-c1","Agent on SWE-bench Verified, campaign 1 (25 instances)","run","2026-10-04","Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.",["/benchmarks/raw/swebench/attempts.json","/benchmarks/raw/swebench/exclusions.json"],"\u0001"],["agent-swebench-c2","Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)","run","2026-10-05","The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.",["/benchmarks/raw/swebench/attempts.json"],"\u0001"],["swebench-leaderboard","SWE-bench Verified leaderboard, mini-SWE-agent v2 runs","public-leaderboard","2026-02-17","Public per-instance results of 11 models under mini-SWE-agent 2.0.0 (bash only, one attempt). Costs are API list prices as published.",["/benchmarks/raw/swebench/panel.json"],"https://www.swebench.com"],["swebench-protocol","SWE-bench campaign rules and sample design","protocol","2026-10-04","Rules declared before the first run: escalations are graded as delivered, gold must resolve on the host, blocked instances are replaced in the same difficulty band, no second attempts. The excluded and replaced instances are listed in the exclusions extract.",["/benchmarks/raw/swebench/exclusions.json"],"\u0001"],["agent-blind-review","Blind review panel: AI worker change vs merged human change","run","2026-09-28","Each pair is judged by 3 or 4 critic models without labels, in both orders. 5 public OSS tasks and 7 private tasks. Private task names are replaced by neutral labels.",["/benchmarks/raw/blind-review/attempts.json"],"\u0001"],["agent-provider-explorer","Provider explorer receipts: CLI vs API","run","2026-10-03","230 imported receipts for short fixed tasks over Claude Code CLI, Codex CLI and the OpenAI API, with time to first useful output, total time, tokens and validation.",["/benchmarks/raw/provider-explorer/runs.json"],"\u0001"],["agent-coding-calibration","Coding calibration: fastify/session, h3, uvicorn","run","2026-10-05","Three real upstream tasks, one attempt per task per platform slice, offline gates against the merged reference.",["/benchmarks/raw/calibration/slices.json"],"\u0001"],["agent-provider-h2h","Provider head-to-head: Claude Code models vs Codex efforts","run","2026-10-05","Five short tasks with deterministic validators, declared protocol, every attempt kept.",["/benchmarks/raw/provider-h2h/receipts.json"],"\u0001"],["agent-provider-h2h-hard","Provider head-to-head, hard set: eight hard tasks with strict validators","run","2026-10-06","Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.",["/benchmarks/raw/provider-h2h-hard/receipts.json"],"\u0001"],["agent-effort-ladder","Effort ladder: the hard task set at each effort level","run","2026-10-06","The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.",["/benchmarks/raw/effort-ladder/receipts.json","/benchmarks/raw/provider-h2h-hard/receipts.json"],"\u0001"],["system-one-arena","System One arena: typed-decision models on checkable decisions and in head-to-head games","run","2026-10-06","Jev 1.13 (TypeSafe API) and six open System One models in llama.cpp 0.6.0 on one Mac Studio (Clef 27B and Clef-Flash 9B Q4_K_M, lev 4B and Kev 4B Q4_K_M, Laya and Julia-1 Q8_0). 1,085 model-blind items in five suites with independent verifiers; exam, latency and repeat passes. Round-robin games (tic-tac-toe, Connect Four, Nim, Dots and Boxes, real-time Pong) against each other and a random and a perfect player, with the featured series and a speed ladder; declared amendments 1 and 2 in the published protocol. Protocol, item hashes and game counts declared before the first counted call; every call and game kept.",["/benchmarks/raw/system-one-arena/items.json","/benchmarks/raw/system-one-arena/exam.json","/benchmarks/raw/system-one-arena/latency.json","/benchmarks/raw/system-one-arena/repeat.json","/benchmarks/raw/system-one-arena/same-host-reference.json","/benchmarks/raw/system-one-arena/matches.json","/benchmarks/raw/system-one-arena/match-summary.json","/benchmarks/raw/system-one-arena/featured.json","/benchmarks/raw/system-one-arena/leagues.json","/benchmarks/raw/system-one-arena/speed-lab.json","/benchmarks/raw/system-one-arena/clock-registry.json"],"\u0001"],["agent-memory-study","Agent memory study: 8 kinds of project memory on Claude Code","run","2026-10-06","One small Node.js repository, 5 tasks, 8 memory conditions (none, /init, curated, raw notes, dreamed notes, long handbook, Stop hook, curated + hook). Claude Code 2.1.286 headless: Sonnet lane 3 repetitions, Haiku lane 2. Hidden tests and deterministic convention checks; protocol declared before the first session; every attempt kept. The exact memory files and three dreaming passes are published.",["/benchmarks/raw/agent-memory/attempts-sonnet.json","/benchmarks/raw/agent-memory/attempts-haiku.json","/benchmarks/raw/agent-memory/dreams.json","/benchmarks/raw/agent-memory/memory-files.json"],"\u0001"],["agent-coding-agents","Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks","run","2026-10-06","Six small Node.js repositories with hidden tests; controls before the first session (every base fails, every reference passes). Claude Code 2.1.286 with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol at medium effort, 2 repetitions per task, OS sandboxes without network, protocol declared before the first session, every session kept. Gemini CLI was probed and not run (browser login).",["/benchmarks/raw/coding-agents/sessions.json"],"\u0001"],["agent-swebench-opus","Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim)","run","2026-10-06","8 Verified instances declared before the first run, 2 per difficulty band. One Opus attempt each on platform build 4f6f4027, paired with the earlier Sonnet attempt on the same instance; official grader. Interim: a usage gate stopped the campaign after 3 instances; the other 5 resume after the reset on 2026-10-09.",["/benchmarks/raw/swebench-opus/instances.json","/benchmarks/raw/swebench/attempts.json"],"\u0001"],["agent-caching-consistency","Caching sessions and repeated prompts (Claude Code and Codex CLI)","run","2026-10-06","Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.",["/benchmarks/raw/caching-consistency/caching.json","/benchmarks/raw/caching-consistency/consistency.json"],"\u0001"],["agent-routing","Routing runs: Jev router vs LLM routing","run","2026-10-05","Routing decisions recorded per case and arm.",["/benchmarks/raw/routing/receipts.json"],"\u0001"],["agent-jev-live","Jev live run: 246 timed calls on the 82 routing decisions","run","2026-10-06","Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.",["/benchmarks/raw/jev-live/summary.json","/benchmarks/raw/jev-live/calls.json"],"\u0001"],["price-anthropic","Anthropic list prices (Claude models)","price-list","2026-09-21","Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.","\u0001","https://platform.claude.com/docs/en/about-claude/pricing"],["price-google","Google Gemini list prices","price-list","2026-09-21","Gemini 3.x Flash prices as listed by the vendor on 2026-09-21. The vendor announced a doubling from 2027-01-01.","\u0001","https://ai.google.dev/pricing"],["price-openai","OpenAI list prices","price-list","2026-10-03","Token prices as listed by the vendor on 2026-10-03.","\u0001","https://developers.openai.com/api/docs/pricing"],["price-jev","Jev 1.13 list price","price-list","2026-09-23","Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.","\u0001","https://docs.typesafe.ai/models"],["agent-routing-overhead","Routing overhead runs: policy microbenchmark and CLI start-up","run","2026-10-06","In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.",["/benchmarks/raw/routing-overhead/results.json"],"\u0001"],["calc-routing-overhead","Routing overhead per 1,000 tasks (calculation)","calculation","2026-10-06","Decisions per task from recorded bench runs multiplied by the cost and the median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound. A calculation, not a run.",["/benchmarks/raw/routing-overhead/results.json"],"\u0001"],["openrouter-api-snapshot","OpenRouter public API: models and provider endpoints (snapshot)","price-list","2026-10-06","Prices, context, quantization and uptime per provider endpoint as reported by OpenRouter’s public, keyless API on 2026-10-06. Third-party-reported, not measured by Agent. Latency and throughput were not returned.",["/benchmarks/raw/provider-index/index.json"],"https://openrouter.ai/docs/api-reference/list-endpoints-for-a-model"],["openrouter-fees","OpenRouter pricing and fees","price-list","2026-10-06","OpenRouter states that inference is billed at the provider list price and that its fee is charged when credits are bought (5.5% on Standard by card, $0.80 minimum; 8% on Business; 5% by crypto). Page fetched 2026-10-06.",["/benchmarks/raw/provider-index/index.json"],"https://openrouter.ai/pricing"],["calc-cache-pricing","Cost with and without the prompt cache (calculation)","calculation","2026-10-06","Recorded tokens per turn × Anthropic list prices. With the cache: uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at twice the input price, 5-minute writes at 1.25 times (an assumption; none occurred). Without a cache: every input token at the input price. Output is priced the same in both. Not a bill.",["/benchmarks/raw/caching-consistency/caching.json"],"\u0001"],["calc-repricing","Repricing calculation","calculation","2026-10-05","Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.","\u0001","\u0001"],["agent-agent-loop","Single call vs agent loop","run","2026-10-06","Receipts of the single call vs agent loop study (public-runs/single-call-vs-agent-loop). Every attempt is kept, failures and contaminated attempts included. Reference single-call cells (Claude Haiku 4.5 and Claude Sonnet 5.5, default effort) are read from raw/provider-h2h-hard.",["/benchmarks/raw/agent-loop/receipts.json"],"\u0001"],["agent-haiku-thinking","Haiku thinking on vs off","run","2026-10-06","Receipts of the Haiku thinking study: Claude Haiku 4.5 through Claude Code with thinking off (MAX_THINKING_TOKENS=0) against the recorded thinking-on arms. Routing arms carry per-arm totals computed from each arm’s per-call log and per-case results; hard-task receipts are every attempt of the thinking-off run. Every attempt is kept, failures included. The thinking-on hard-task receipts live in the hard head-to-head extract.",["/benchmarks/raw/haiku-thinking/receipts.json"],"\u0001"],["agent-structured-output","JSON schema vs instructions","run","2026-10-07","Receipts copied from a run of the same three JSON extraction prompts, each asked with instructions only (mode I) and with the CLI’s JSON schema mode (mode S). Prompts, model output and failure reasons are not published; a failed check is named, never quoted. Every attempt is kept. Both routes ran one call at a time. The current protocol file does not verify pre-call registration; see protocolAudit. Calls repeat three fixed hand-made prompts, including a known format-miss case. Both routes used a shared Mac.",["/benchmarks/raw/structured-output/receipts.json"],"\u0001"],["agent-cache-sessions","Prompt cache across sessions","run","2026-10-07","Sanitized cache-session receipts: one CLI process per session, 2 turns each, a seeded synthetic ledger (a different seed per condition) and 2 short questions with exact answers. Conditions: A a new temporary working folder per session, B one fixed folder, C a fixed folder with the ledger in the system prompt (Claude Code only). The Codex app-server reports cached input only, so its rows have no cache-write count. Probes are single uncounted calls with their own ledger seed. The answers and the ledger itself are not published. Session 1 was the first use of each setup’s ledger; sessions 2 and 3 reused that ledger. Different seeds prevent full ledger-prefix reuse between setups, but shared CLI-prefix reads remain possible. Correctness uses the reference checker after it trims spaces, surrounding quotes and backticks, a final period and currency units. The surviving protocol file dates from after the counted calls; pre-call declaration is not verified.",["/benchmarks/raw/cache-sessions/sessions.json"],"\u0001"],["agent-routing-holdout","Routing on unseen holdout decisions","run","2026-10-06","Frozen unseen routing decisions; every router call retained. Costs are calculations from recorded tokens and list prices.",["/benchmarks/raw/routing-holdout/results.json"],"\u0001"],["agent-speed-anatomy","LLM speed anatomy","run","2026-10-07","Receipts of the speed anatomy study: one prompt that asks for 250 numbers in words (output speed, 24 calls) and a seeded synthetic ledger at three sizes with one lookup question (prompt size, 36 calls). Every attempt is kept, failures included. A new seed for every ledger call; no ledger text, prompt or model output is copied.",["/benchmarks/raw/speed-anatomy/receipts.json"],"\u0001"],["calc-thinking-bill","Reasoning token bill (calculation)","calculation","2026-10-07","Reported reasoning tokens priced at the recorded list prices. A calculation, not a new run.",["/benchmarks/raw/provider-h2h-hard/receipts.json","/benchmarks/raw/effort-ladder/receipts.json","/benchmarks/raw/provider-h2h/receipts.json"],"\u0001"],["agent-harder-tasks","Harder tasks head-to-head","run","2026-10-07","Frozen tasks selected with a Sonnet pilot. Fresh counted calls retain failures, format misses and timeouts. Replies and expected answers are not published.",["/benchmarks/raw/harder-tasks/receipts.json"],"\u0001"],["calc-latency-budget","Voice-agent latency budget calculation","calculation","2026-10-07","Recorded decision and first-output times compared with assumed 300, 800 and 1,500 ms budgets. No complete voice turn ran.","\u0001","\u0001"],["calc-routing-at-scale","Routing at scale calculation","calculation","2026-10-07","Recorded routing costs and times scaled to assumed daily volumes. No load test ran; median and p95 scenarios are not measured mean concurrency.","\u0001","\u0001"]]},"articleCharts":{"/blog/ai-agent-cost-per-task-claims-vs-receipts":["swebench-cost-by-stage","swebench-cost-vs-resolved","cost-per-resolved-agent-vs-panel"],"/blog/ai-agent-says-done-but-tests-fail":["memory-late-fee","memory-late-fee-haiku","memory-broken-test-command","calibration-outcomes-by-slice"],"/blog/ai-coding-agent-best-practices-backed-by-data":["consistency-pass-rate","caching-read-share-by-turn","memory-knowledge-class","router-overhead-decision-latency","swebench-same-instance-leaderboard"],"/blog/ai-coding-benchmarks-roundup-october-2026":["swebench-same-instance-leaderboard","hard-h2h-pass-rate","router-overhead-cost-per-1000-tasks","provider-index-spread","effort-ladder-pass-rate","caching-cost-with-without","memory-knowledge-class","coding-agents-wall-time","swebench-opus-sonnet-resolved","arena-accuracy"],"/blog/ai-coding-cost-per-developer-formula":["hard-h2h-cost-per-pass","prompt-cache-savings","repriced-cost-per-resolved","cost-per-resolved-agent-vs-panel"],"/blog/ai-model-leaderboard-without-composite-score":["effort-ladder-pass-rate"],"/blog/ai-subscription-usage-meter":[],"/blog/ai-week-october-2-8-2026":[],"/blog/best-ai-model-for-coding-october-2026":["hard-h2h-pass-rate","swebench-same-instance-leaderboard","hard-h2h-cost-per-pass","hard-h2h-frontier"],"/blog/best-llm-for-json-output-haiku-sonnet-gpt-6-1-sol":["consistency-pass-rate","consistency-distinct-answers","consistency-latency-spread"],"/blog/blind-review-ai-vs-human-pull-requests":["blind-review-first-vs-latest","blind-review-votes-by-task","blind-review-dimension-scores","blind-review-critic-agreement"],"/blog/cheapest-llm-for-classification-at-scale":["routing-exact-decisions","router-overhead-decision-latency","routing-exact-by-decision"],"/blog/cheapest-open-model-inference-providers-october-2026":["provider-index-spread","provider-prices-deepseek-v4-flash","provider-prices-deepseek-v4-pro","provider-prices-gpt-oss-120b","provider-prices-llama-3-3-70b-instruct","provider-prices-glm-5-3","provider-prices-kimi-k3"],"/blog/cheapest-way-to-run-an-ai-coding-agent":["memory-cost-per-full-pass","provider-index-spread"],"/blog/claude-code-cost-per-task-at-api-prices":["h2h-list-price-per-call","hard-h2h-cost-per-pass","caching-cost-with-without","memory-cost-per-full-pass","swebench-cost-by-stage"],"/blog/claude-code-stop-hook-tested":["memory-knowledge-class","memory-input-tokens","memory-wall-time","memory-late-fee","memory-broken-test-command","memory-full-pass"],"/blog/claude-code-vs-codex-cli-coding-agents-hidden-tests":["coding-agents-wall-time","coding-agents-pass-rate","coding-agents-time-by-task","coding-agents-tool-calls","coding-agents-diff-lines","coding-agents-token-mix","coding-agents-cost-per-pass"],"/blog/claude-code-vs-codex-cli-context-tax":["cli-vs-api-prompt-overhead","h2h-input-tokens","scheduler-repair-claude-vs-codex"],"/blog/claude-code-vs-codex-cli-vs-api-latency":["cli-vs-api-exact-reply-latency","cli-vs-api-prompt-overhead","scheduler-repair-claude-vs-codex","h2h-first-useful-latency"],"/blog/claude-fable-5-1-vs-opus-5-5-vs-sonnet-5-5":["h2h-total-latency","hard-h2h-total-latency","h2h-input-tokens","provider-prices-claude-fable-5-1","hard-h2h-cost-per-pass"],"/blog/claude-frontier-academy-production-ai":[],"/blog/claude-haiku-4-5-vs-sonnet-5-5-every-measured-row":["hard-h2h-pass-rate","h2h-pass-rate","consistency-pass-rate","memory-team-knowledge-by-model","memory-broken-test-command","router-overhead-decision-latency","hard-h2h-output-tokens","hard-h2h-cost-per-pass","h2h-cost-per-pass","routing-cost-per-1000"],"/blog/claude-haiku-5-5-subagents":[],"/blog/claude-haiku-thinking-on-vs-off-speed-and-accuracy":["haiku-thinking-router-exact","haiku-thinking-router-latency","router-overhead-decision-latency","haiku-thinking-router-tokens","haiku-thinking-router-cost","haiku-thinking-hard-pass","haiku-thinking-hard-time"],"/blog/claude-haiku-vs-sonnet-vs-opus-vs-fable-vs-codex":["h2h-total-latency","h2h-input-tokens","h2h-cost-per-pass","h2h-frontier"],"/blog/claude-max-included-api-credits":[],"/blog/claude-opus-vs-sonnet-swe-bench-verified-early-look":["swebench-opus-sonnet-resolved","swebench-opus-sonnet-cost-by-instance","swebench-opus-sonnet-cost-per-attempt","swebench-opus-sonnet-stage-cost","swebench-opus-sonnet-minutes"],"/blog/claude-sonnet-vs-opus-when-is-opus-worth-it":["h2h-list-price-per-call","hard-h2h-cost-per-pass","repriced-cost-per-resolved","coding-agents-cost-per-pass","routing-economics-scenarios"],"/blog/claude-vs-codex-hard-tasks-gpt-6-1-sol":["hard-h2h-pass-rate","hard-h2h-total-latency","hard-h2h-first-useful-latency","cli-startup-tax","hard-h2h-output-tokens","hard-h2h-cost-per-pass","hard-h2h-frontier"],"/blog/claude-vs-codex-personal-workflow":[],"/blog/decision-models-fair-mode-othello-tron-snake-pong":[],"/blog/decision-models-play-connect-four-nim-pong":["arena-elo","arena-pong-ladder"],"/blog/does-claude-code-reuse-prompt-cache-across-sessions":["caching-read-share-by-turn","cache-sessions-turn1-read-share","cache-sessions-turn1-cost","cache-sessions-turn-time","cache-sessions-codex-turn1-cached"],"/blog/does-claude-md-help-agent-memory-tested":["memory-knowledge-class","memory-late-fee","memory-late-fee-haiku","memory-broken-test-command","memory-team-knowledge-by-model","memory-dreaming-scorecard","memory-input-tokens","memory-cost-per-full-pass","memory-full-pass"],"/blog/does-letting-ai-run-code-help-single-call-vs-agent-loop":["agent-loop-pass-rate","agent-loop-tool-calls","agent-loop-by-task","agent-loop-total-time","agent-loop-tokens","agent-loop-cost-per-pass"],"/blog/does-llm-routing-save-money-net-of-router-cost":["routing-economics-scenarios","routing-cost-per-1000","router-overhead-cost-per-1000-tasks","routing-economics-by-stage","hard-h2h-pass-rate"],"/blog/does-reasoning-effort-buy-quality-claude-codex":["effort-ladder-pass-rate","effort-ladder-time-by-effort","effort-ladder-total-latency","effort-ladder-output-tokens","effort-ladder-cost-per-pass"],"/blog/fastest-claude-model-haiku-sonnet-opus-fable-timed":["h2h-total-latency","hard-h2h-total-latency","routing-decision-latency","hard-h2h-output-tokens","effort-ladder-total-latency"],"/blog/gpt-6-1-sol-vs-claude-sonnet-5-5-vs-opus-5-5":["hard-h2h-pass-rate","hard-h2h-total-latency","consistency-latency-spread","scheduler-repair-claude-vs-codex","hard-h2h-cost-per-pass"],"/blog/haiku-thinking-on-off":["haiku-thinking-router-tokens","haiku-thinking-router-exact","haiku-thinking-router-latency","haiku-thinking-router-cost","haiku-thinking-hard-pass","haiku-thinking-hard-time"],"/blog/hard-tasks-claude-haiku-sonnet-opus-fable-codex":["hard-h2h-pass-rate","hard-h2h-outcomes","hard-h2h-total-latency","hard-h2h-output-tokens","hard-h2h-cost-per-pass"],"/blog/harder-coding-tasks-gpt-6-1-sol-vs-claude-sonnet-vs-opus":["harder-h2h-pass-rate","hard-h2h-pass-rate","harder-h2h-tool-attempts","harder-h2h-pass-by-task","harder-h2h-cost-per-pass"],"/blog/harness-vs-model-where-agent-gains-come-from":["swebench-cost-vs-resolved","cli-vs-api-small-coding-latency","swebench-by-difficulty-band","calibration-cost-by-task"],"/blog/how-fast-is-jev-router-latency-measured":["router-overhead-decision-latency","routing-exact-decisions"],"/blog/how-long-does-an-ai-coding-agent-take-per-task":["calibration-minutes-by-task","swebench-model-calls"],"/blog/how-many-runs-to-compare-two-llms":["consistency-pass-rate","swebench-same-instance-leaderboard","effort-ladder-pass-rate"],"/blog/how-much-does-prompt-caching-save":["caching-read-share-by-turn","caching-cost-with-without","prompt-cache-savings","caching-latency-first-vs-later"],"/blog/how-much-of-your-ai-bill-is-thinking-tokens":["thinking-bill-share","thinking-bill-short-vs-hard","thinking-bill-cost-per-call","effort-ladder-output-tokens","thinking-bill-by-effort","thinking-bill-vs-time"],"/blog/how-to-estimate-your-ai-coding-bill":["token-cost-mix","hard-h2h-cost-per-pass","routing-cost-per-1000","prompt-cache-savings"],"/blog/intelligent-ui-houston-hackathon":[],"/blog/is-claude-haiku-cheaper-retry-and-escalate":["retry-escalate-haiku-by-task","retry-escalate-cost-per-correct","retry-escalate-time-per-correct","retry-escalate-success","retry-escalate-call-cost-by-task","hard-h2h-cost-per-pass","retry-escalate-sensitivity"],"/blog/is-claude-haiku-cheaper-than-sonnet-cost-per-correct-answer":["hard-h2h-cost-per-pass","hard-h2h-output-tokens","h2h-cost-per-pass","memory-cost-per-full-pass","routing-cost-per-1000"],"/blog/jev-vs-claude-haiku-vs-sonnet-router-compared":["routing-exact-decisions","routing-key-accuracy","routing-cost-per-1000","router-overhead-cost-per-1000-tasks","router-overhead-decision-latency"],"/blog/jev-vs-claude-routers-on-unseen-decisions":["routing-holdout-exact","routing-holdout-key-accuracy","routing-holdout-tuned-vs-unseen","routing-exact-decisions","routing-holdout-by-purpose","routing-holdout-latency","routing-holdout-cost-per-1000"],"/blog/jev-vs-clef-open-decision-models-tested":["arena-accuracy","arena-escape"],"/blog/jev-vs-llm-routers-routing-decisions":["routing-exact-decisions","routing-exact-by-decision","routing-cost-per-1000","routing-decision-latency"],"/blog/json-schema-vs-prompt-instructions-format-misses":["consistency-pass-rate","structured-output-pass-rate","structured-output-outcomes","structured-output-time","structured-output-tokens"],"/blog/liquid-ai-d1-decision-models":[],"/blog/llm-api-pricing-comparison-claude-gpt-gemini-october-2026":["provider-prices-claude-sonnet-5-5","provider-prices-gemini-3-8-flash","provider-prices-gpt-6-luna","hard-h2h-cost-per-pass","repriced-cost-per-resolved"],"/blog/llm-eval-format-misses-vs-wrong-answers":["hard-h2h-outcomes","h2h-pass-rate","consistency-pass-rate"],"/blog/llm-tail-latency-p95-not-median":["router-overhead-decision-latency","consistency-latency-spread"],"/blog/llm-time-to-first-token-and-tokens-per-second":["cli-startup-tax","speed-anatomy-first-text","speed-anatomy-output-speed","speed-anatomy-chars-per-second","speed-anatomy-prompt-size","speed-anatomy-total-by-size","speed-anatomy-lookup-correct"],"/blog/mac-studio-96gb-ai-memory":[],"/blog/mcnemar-test-compare-two-llms":["routing-exact-decisions","routing-key-accuracy"],"/blog/most-ai-model-comparisons-are-ties":["cli-startup-tax","router-overhead-decision-latency","hard-h2h-pass-rate","effort-ladder-pass-rate"],"/blog/openai-math-repository-verification":[],"/blog/openrouter-vs-direct-api-gateway-cost":["gateway-markup-vs-first-party","gateway-vs-direct-claude-sonnet-5-5","gateway-vs-direct-gemini-3-8-flash","provider-index-spread","provider-prices-claude-opus-5-5"],"/blog/opus-low-effort-vs-sonnet-high-effort":["effort-ladder-pass-rate","effort-ladder-total-latency","effort-ladder-output-tokens","effort-ladder-cost-per-pass"],"/blog/reasoning-tokens-cost-claude-gpt-6-1-sol":["hard-h2h-output-tokens","effort-ladder-output-tokens"],"/blog/same-prompt-ten-answers-claude-codex-consistency":["consistency-pass-rate","consistency-distinct-answers","consistency-latency-spread"],"/blog/self-consistency-majority-voting-thought-experiment":["consistency-pass-rate","consistency-distinct-answers","consistency-latency-spread"],"/blog/single-call-vs-agent-loop":["agent-loop-tokens","agent-loop-total-time","agent-loop-cost-per-pass","agent-loop-pass-rate"],"/blog/swe-bench-verified-agent-vs-public-models":["swebench-same-instance-leaderboard","swebench-by-difficulty-band","swebench-views","swebench-model-calls"],"/blog/voice-agent-latency-budget-llm-steps":["router-overhead-decision-latency","routing-decision-latency","cli-vs-api-exact-reply-latency","cli-startup-tax"],"/blog/voice-agent-latency-budget-what-fits-in-one-turn":["latency-budget-fit","router-overhead-decision-latency","latency-budget-fast-steps","latency-budget-slow-steps","latency-budget-slow-steps-ranges"],"/blog/what-a-resolved-swe-bench-task-costs":["cost-per-resolved-agent-vs-panel","swebench-cost-by-stage","token-cost-mix","prompt-cache-savings"],"/blog/what-does-an-llm-router-cost":["router-overhead-decision-latency","router-overhead-cli-vs-model-time","router-overhead-cost-reported","router-overhead-cost-list-price","router-overhead-cost-per-1000-tasks","router-overhead-delay-per-task"],"/blog/what-does-routing-a-million-ai-requests-a-day-cost":["routing-at-scale-daily-cost","routing-at-scale-in-flight","routing-at-scale-waiting-hours","routing-at-scale-scenarios"],"/blog/what-if-every-call-ran-on-opus":["repriced-cost-per-resolved","prompt-cache-savings","routing-economics-scenarios","routing-economics-by-stage"],"/blog/when-does-prompt-caching-pay-off":["cache-break-even-reads","cache-break-even-cost-curve","cache-break-even-session-split","caching-cost-with-without"],"/blog/which-claude-model-should-i-use":["hard-h2h-frontier","h2h-total-latency","consistency-pass-rate","memory-team-knowledge-by-model","effort-ladder-cost-per-pass"],"/blog/why-ai-coding-agents-fail-on-real-pull-requests":["calibration-outcomes-by-slice","hard-h2h-outcomes","calibration-cost-by-task","calibration-guardrail-refusals"],"/blog/why-is-claude-code-slow-where-the-seconds-go":["cli-startup-tax","cli-startup-input-tokens","router-overhead-cli-vs-model-time","hard-h2h-output-tokens","effort-ladder-time-by-effort","h2h-total-latency","cli-vs-api-exact-reply-latency"],"/blog/why-we-count-every-failed-attempt":["swebench-views","calibration-outcomes-by-slice","blind-review-first-vs-latest","calibration-guardrail-refusals","coding-agents-pass-rate"],"/learn/agent-memory-consolidation-dreaming":["memory-broken-test-command","memory-dreaming-scorecard"],"/learn/benchmark-saturation-ceiling-effect":["h2h-pass-rate","hard-h2h-pass-rate","effort-ladder-pass-rate"],"/learn/claude-code-hooks-explained":["memory-team-knowledge-by-model","memory-wall-time"],"/learn/claude-md-and-agent-memory-files":["memory-knowledge-class","memory-full-pass","memory-input-tokens"],"/learn/cost-per-correct-answer":["h2h-list-price-per-call","hard-h2h-cost-per-pass"],"/learn/how-many-runs-llm-eval-sample-size":["swebench-views","hard-h2h-pass-rate","routing-exact-decisions"],"/learn/how-to-read-ai-benchmarks-honestly":["hard-h2h-outcomes","h2h-total-latency","repriced-cost-per-resolved"],"/learn/inference-gateways-and-openrouter":["gateway-markup-vs-first-party","provider-index-spread"],"/learn/inference-providers-bedrock-vertex-azure":["provider-prices-claude-sonnet-5-5","gateway-markup-vs-first-party","provider-index-spread"],"/learn/llm-as-judge-blind-review":["blind-review-first-vs-latest","blind-review-critic-agreement"],"/learn/llm-latency-median-p95-percentiles":["router-overhead-decision-latency","consistency-latency-spread"],"/learn/llm-nondeterminism-same-prompt-different-answers":["consistency-distinct-answers","consistency-pass-rate","consistency-latency-spread"],"/learn/llm-pricing-per-million-tokens":["token-cost-mix","repriced-cost-per-resolved","provider-prices-claude-sonnet-5-5"],"/learn/mcnemar-test-for-llm-comparisons":["routing-exact-decisions","routing-exact-by-decision"],"/learn/open-weight-models-and-provider-prices":["provider-index-spread","provider-prices-gpt-oss-120b","provider-prices-llama-3-3-70b-instruct"],"/learn/pareto-frontier-llm-cost-quality":["hard-h2h-frontier","h2h-frontier"],"/learn/pass-at-k-explained":["hard-h2h-pass-rate","consistency-pass-rate","hard-h2h-outcomes"],"/learn/prompt-caching-explained":["caching-read-share-by-turn","caching-cost-with-without","prompt-cache-savings"],"/learn/quantization-and-cheap-inference":["provider-prices-gpt-oss-120b","provider-prices-deepseek-v4-flash"],"/learn/reasoning-tokens-explained":["hard-h2h-output-tokens","effort-ladder-output-tokens","scheduler-repair-output-tokens"],"/learn/strict-vs-lenient-grading-format-misses":["hard-h2h-pass-rate","consistency-pass-rate"],"/learn/swe-bench-verified-explained":["swebench-same-instance-leaderboard","swebench-by-difficulty-band"],"/learn/time-to-first-token-explained":["cli-startup-tax","cli-vs-api-exact-reply-latency","h2h-first-useful-latency"],"/learn/tokens-per-call-and-the-cli-context-tax":["cli-vs-api-prompt-overhead","h2h-input-tokens"],"/learn/what-is-an-agent-harness":["scheduler-repair-claude-vs-codex","swebench-cost-vs-resolved","swebench-cost-by-stage"],"/learn/what-is-an-llm-router":["routing-exact-decisions","router-overhead-decision-latency","routing-cost-per-1000"],"/learn/what-is-context-engineering":["cli-startup-input-tokens","memory-knowledge-class","caching-read-share-by-turn"],"/learn/what-is-reasoning-effort":["effort-ladder-pass-rate","effort-ladder-total-latency","effort-ladder-cost-per-pass"],"/learn/wilson-confidence-intervals-for-ai-benchmarks":["hard-h2h-pass-rate","consistency-pass-rate"]}}