{"schema":"agent-public-bench@1","generatedAt":"2026-10-07T00:00:00.000Z","sources":{"$k":["id","title","kind","date","note","data","url"],"$r":[["agent-swebench-c1","Agent on SWE-bench Verified, campaign 1 (25 instances)","run","2026-10-04","Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.",["/benchmarks/raw/swebench/attempts.json","/benchmarks/raw/swebench/exclusions.json"],"\u0001"],["agent-swebench-c2","Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)","run","2026-10-05","The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.",["/benchmarks/raw/swebench/attempts.json"],"\u0001"],["swebench-leaderboard","SWE-bench Verified leaderboard, mini-SWE-agent v2 runs","public-leaderboard","2026-02-17","Public per-instance results of 11 models under mini-SWE-agent 2.0.0 (bash only, one attempt). Costs are API list prices as published.",["/benchmarks/raw/swebench/panel.json"],"https://www.swebench.com"],["swebench-protocol","SWE-bench campaign rules and sample design","protocol","2026-10-04","Rules declared before the first run: escalations are graded as delivered, gold must resolve on the host, blocked instances are replaced in the same difficulty band, no second attempts. The excluded and replaced instances are listed in the exclusions extract.",["/benchmarks/raw/swebench/exclusions.json"],"\u0001"],["agent-blind-review","Blind review panel: AI worker change vs merged human change","run","2026-09-28","Each pair is judged by 3 or 4 critic models without labels, in both orders. 5 public OSS tasks and 7 private tasks. Private task names are replaced by neutral labels.",["/benchmarks/raw/blind-review/attempts.json"],"\u0001"],["agent-provider-explorer","Provider explorer receipts: CLI vs API","run","2026-10-03","230 imported receipts for short fixed tasks over Claude Code CLI, Codex CLI and the OpenAI API, with time to first useful output, total time, tokens and validation.",["/benchmarks/raw/provider-explorer/runs.json"],"\u0001"],["agent-coding-calibration","Coding calibration: fastify/session, h3, uvicorn","run","2026-10-05","Three real upstream tasks, one attempt per task per platform slice, offline gates against the merged reference.",["/benchmarks/raw/calibration/slices.json"],"\u0001"],["agent-provider-h2h","Provider head-to-head: Claude Code models vs Codex efforts","run","2026-10-05","Five short tasks with deterministic validators, declared protocol, every attempt kept.",["/benchmarks/raw/provider-h2h/receipts.json"],"\u0001"],["agent-provider-h2h-hard","Provider head-to-head, hard set: eight hard tasks with strict validators","run","2026-10-06","Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.",["/benchmarks/raw/provider-h2h-hard/receipts.json"],"\u0001"],["agent-effort-ladder","Effort ladder: the hard task set at each effort level","run","2026-10-06","The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.",["/benchmarks/raw/effort-ladder/receipts.json","/benchmarks/raw/provider-h2h-hard/receipts.json"],"\u0001"],["system-one-arena","System One arena: typed-decision models on checkable decisions and in head-to-head games","run","2026-10-06","Jev 1.13 (TypeSafe API) and six open System One models in llama.cpp 0.6.0 on one Mac Studio (Clef 27B and Clef-Flash 9B Q4_K_M, lev 4B and Kev 4B Q4_K_M, Laya and Julia-1 Q8_0). 1,085 model-blind items in five suites with independent verifiers; exam, latency and repeat passes. Round-robin games (tic-tac-toe, Connect Four, Nim, Dots and Boxes, real-time Pong) against each other and a random and a perfect player, with the featured series and a speed ladder; declared amendments 1 and 2 in the published protocol. Protocol, item hashes and game counts declared before the first counted call; every call and game kept.",["/benchmarks/raw/system-one-arena/items.json","/benchmarks/raw/system-one-arena/exam.json","/benchmarks/raw/system-one-arena/latency.json","/benchmarks/raw/system-one-arena/repeat.json","/benchmarks/raw/system-one-arena/same-host-reference.json","/benchmarks/raw/system-one-arena/matches.json","/benchmarks/raw/system-one-arena/match-summary.json","/benchmarks/raw/system-one-arena/featured.json","/benchmarks/raw/system-one-arena/leagues.json","/benchmarks/raw/system-one-arena/speed-lab.json","/benchmarks/raw/system-one-arena/clock-registry.json"],"\u0001"],["agent-memory-study","Agent memory study: 8 kinds of project memory on Claude Code","run","2026-10-06","One small Node.js repository, 5 tasks, 8 memory conditions (none, /init, curated, raw notes, dreamed notes, long handbook, Stop hook, curated + hook). Claude Code 2.1.286 headless: Sonnet lane 3 repetitions, Haiku lane 2. Hidden tests and deterministic convention checks; protocol declared before the first session; every attempt kept. The exact memory files and three dreaming passes are published.",["/benchmarks/raw/agent-memory/attempts-sonnet.json","/benchmarks/raw/agent-memory/attempts-haiku.json","/benchmarks/raw/agent-memory/dreams.json","/benchmarks/raw/agent-memory/memory-files.json"],"\u0001"],["agent-coding-agents","Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks","run","2026-10-06","Six small Node.js repositories with hidden tests; controls before the first session (every base fails, every reference passes). Claude Code 2.1.286 with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol at medium effort, 2 repetitions per task, OS sandboxes without network, protocol declared before the first session, every session kept. Gemini CLI was probed and not run (browser login).",["/benchmarks/raw/coding-agents/sessions.json"],"\u0001"],["agent-swebench-opus","Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim)","run","2026-10-06","8 Verified instances declared before the first run, 2 per difficulty band. One Opus attempt each on platform build 4f6f4027, paired with the earlier Sonnet attempt on the same instance; official grader. Interim: a usage gate stopped the campaign after 3 instances; the other 5 resume after the reset on 2026-10-09.",["/benchmarks/raw/swebench-opus/instances.json","/benchmarks/raw/swebench/attempts.json"],"\u0001"],["agent-caching-consistency","Caching sessions and repeated prompts (Claude Code and Codex CLI)","run","2026-10-06","Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.",["/benchmarks/raw/caching-consistency/caching.json","/benchmarks/raw/caching-consistency/consistency.json"],"\u0001"],["agent-routing","Routing runs: Jev router vs LLM routing","run","2026-10-05","Routing decisions recorded per case and arm.",["/benchmarks/raw/routing/receipts.json"],"\u0001"],["agent-jev-live","Jev live run: 246 timed calls on the 82 routing decisions","run","2026-10-06","Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.",["/benchmarks/raw/jev-live/summary.json","/benchmarks/raw/jev-live/calls.json"],"\u0001"],["price-anthropic","Anthropic list prices (Claude models)","price-list","2026-09-21","Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.","\u0001","https://platform.claude.com/docs/en/about-claude/pricing"],["price-google","Google Gemini list prices","price-list","2026-09-21","Gemini 3.x Flash prices as listed by the vendor on 2026-09-21. The vendor announced a doubling from 2027-01-01.","\u0001","https://ai.google.dev/pricing"],["price-openai","OpenAI list prices","price-list","2026-10-03","Token prices as listed by the vendor on 2026-10-03.","\u0001","https://developers.openai.com/api/docs/pricing"],["price-jev","Jev 1.13 list price","price-list","2026-09-23","Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.","\u0001","https://docs.typesafe.ai/models"],["agent-routing-overhead","Routing overhead runs: policy microbenchmark and CLI start-up","run","2026-10-06","In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.",["/benchmarks/raw/routing-overhead/results.json"],"\u0001"],["calc-routing-overhead","Routing overhead per 1,000 tasks (calculation)","calculation","2026-10-06","Decisions per task from recorded bench runs multiplied by the cost and the median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound. A calculation, not a run.",["/benchmarks/raw/routing-overhead/results.json"],"\u0001"],["openrouter-api-snapshot","OpenRouter public API: models and provider endpoints (snapshot)","price-list","2026-10-06","Prices, context, quantization and uptime per provider endpoint as reported by OpenRouter’s public, keyless API on 2026-10-06. Third-party-reported, not measured by Agent. Latency and throughput were not returned.",["/benchmarks/raw/provider-index/index.json"],"https://openrouter.ai/docs/api-reference/list-endpoints-for-a-model"],["openrouter-fees","OpenRouter pricing and fees","price-list","2026-10-06","OpenRouter states that inference is billed at the provider list price and that its fee is charged when credits are bought (5.5% on Standard by card, $0.80 minimum; 8% on Business; 5% by crypto). Page fetched 2026-10-06.",["/benchmarks/raw/provider-index/index.json"],"https://openrouter.ai/pricing"],["calc-cache-pricing","Cost with and without the prompt cache (calculation)","calculation","2026-10-06","Recorded tokens per turn × Anthropic list prices. With the cache: uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at twice the input price, 5-minute writes at 1.25 times (an assumption; none occurred). Without a cache: every input token at the input price. Output is priced the same in both. Not a bill.",["/benchmarks/raw/caching-consistency/caching.json"],"\u0001"],["calc-repricing","Repricing calculation","calculation","2026-10-05","Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.","\u0001","\u0001"],["agent-agent-loop","Single call vs agent loop","run","2026-10-06","Receipts of the single call vs agent loop study (public-runs/single-call-vs-agent-loop). Every attempt is kept, failures and contaminated attempts included. Reference single-call cells (Claude Haiku 4.5 and Claude Sonnet 5.5, default effort) are read from raw/provider-h2h-hard.",["/benchmarks/raw/agent-loop/receipts.json"],"\u0001"],["agent-haiku-thinking","Haiku thinking on vs off","run","2026-10-06","Receipts of the Haiku thinking study: Claude Haiku 4.5 through Claude Code with thinking off (MAX_THINKING_TOKENS=0) against the recorded thinking-on arms. Routing arms carry per-arm totals computed from each arm’s per-call log and per-case results; hard-task receipts are every attempt of the thinking-off run. Every attempt is kept, failures included. The thinking-on hard-task receipts live in the hard head-to-head extract.",["/benchmarks/raw/haiku-thinking/receipts.json"],"\u0001"],["agent-structured-output","JSON schema vs instructions","run","2026-10-07","Receipts copied from a run of the same three JSON extraction prompts, each asked with instructions only (mode I) and with the CLI’s JSON schema mode (mode S). Prompts, model output and failure reasons are not published; a failed check is named, never quoted. Every attempt is kept. Both routes ran one call at a time. The current protocol file does not verify pre-call registration; see protocolAudit. Calls repeat three fixed hand-made prompts, including a known format-miss case. Both routes used a shared Mac.",["/benchmarks/raw/structured-output/receipts.json"],"\u0001"],["agent-cache-sessions","Prompt cache across sessions","run","2026-10-07","Sanitized cache-session receipts: one CLI process per session, 2 turns each, a seeded synthetic ledger (a different seed per condition) and 2 short questions with exact answers. Conditions: A a new temporary working folder per session, B one fixed folder, C a fixed folder with the ledger in the system prompt (Claude Code only). The Codex app-server reports cached input only, so its rows have no cache-write count. Probes are single uncounted calls with their own ledger seed. The answers and the ledger itself are not published. Session 1 was the first use of each setup’s ledger; sessions 2 and 3 reused that ledger. Different seeds prevent full ledger-prefix reuse between setups, but shared CLI-prefix reads remain possible. Correctness uses the reference checker after it trims spaces, surrounding quotes and backticks, a final period and currency units. The surviving protocol file dates from after the counted calls; pre-call declaration is not verified.",["/benchmarks/raw/cache-sessions/sessions.json"],"\u0001"],["agent-routing-holdout","Routing on unseen holdout decisions","run","2026-10-06","Frozen unseen routing decisions; every router call retained. Costs are calculations from recorded tokens and list prices.",["/benchmarks/raw/routing-holdout/results.json"],"\u0001"],["agent-speed-anatomy","LLM speed anatomy","run","2026-10-07","Receipts of the speed anatomy study: one prompt that asks for 250 numbers in words (output speed, 24 calls) and a seeded synthetic ledger at three sizes with one lookup question (prompt size, 36 calls). Every attempt is kept, failures included. A new seed for every ledger call; no ledger text, prompt or model output is copied.",["/benchmarks/raw/speed-anatomy/receipts.json"],"\u0001"],["calc-thinking-bill","Reasoning token bill (calculation)","calculation","2026-10-07","Reported reasoning tokens priced at the recorded list prices. A calculation, not a new run.",["/benchmarks/raw/provider-h2h-hard/receipts.json","/benchmarks/raw/effort-ladder/receipts.json","/benchmarks/raw/provider-h2h/receipts.json"],"\u0001"],["agent-harder-tasks","Harder tasks head-to-head","run","2026-10-07","Frozen tasks selected with a Sonnet pilot. Fresh counted calls retain failures, format misses and timeouts. Replies and expected answers are not published.",["/benchmarks/raw/harder-tasks/receipts.json"],"\u0001"],["calc-latency-budget","Voice-agent latency budget calculation","calculation","2026-10-07","Recorded decision and first-output times compared with assumed 300, 800 and 1,500 ms budgets. No complete voice turn ran.","\u0001","\u0001"],["calc-routing-at-scale","Routing at scale calculation","calculation","2026-10-07","Recorded routing costs and times scaled to assumed daily volumes. No load test ran; median and p95 scenarios are not measured mean concurrency.","\u0001","\u0001"]]},"studies":{"$k":["slug","title","seoTitle","description","question","answer","date","updated","tags","method","caveats","sourceIds","stats","charts","tables","related","hero"],"$r":[["swe-bench-verified","Agent on SWE-bench Verified vs 11 public models","SWE-bench Verified: Agent vs GPT, Claude and Gemini","","","","2026-10-05","2026-10-05",[],[],[],["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol","price-anthropic"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["swebench-same-instance-leaderboard","Resolved rate on the same 33 SWE-bench Verified instances","dot-range","rate",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"],[]],["swebench-by-difficulty-band","Resolved rate by difficulty band","grouped-bar","rate",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol"],[]],["swebench-cost-vs-resolved","Cost per instance vs resolved rate","scatter","rate",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","price-anthropic"],[]],["swebench-model-calls","Model calls per instance","bar","calls",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"],[]],["swebench-cost-by-stage","Where Agent's model spend goes","bar","usd",["agent-swebench-c1","agent-swebench-c2"],[]],["swebench-views","Every way to slice the run, with intervals","dot-range","rate",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol"],[]]]},[],[],"\u0001"],["swe-bench-opus-vs-sonnet","Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)","Opus 5.5 vs Sonnet 5.5 on SWE-bench Verified (interim)","","","","2026-10-06","2026-10-06",[],[],[],["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["swebench-opus-sonnet-resolved","Resolved on the same 3 SWE-bench Verified instances (interim)","dot-range","rate",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2"],[]],["swebench-opus-sonnet-cost-per-attempt","List-price cost per attempt (calculation)","bar","usd",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"],[]],["swebench-opus-sonnet-cost-by-instance","List-price cost per instance (calculation)","grouped-bar","usd",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"],[]],["swebench-opus-sonnet-minutes","Worker time per attempt","dot-range","minutes",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2"],[]],["swebench-opus-sonnet-stage-cost","Where the cost goes: pipeline stages (calculation)","grouped-bar","usd",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"],[]]]},[],[],{"statIds":["swebench-opus-resolved","swebench-sonnet-resolved"],"testStatId":"swebench-opus-sonnet-mcnemar"}],["blind-review-head-to-head","AI pull requests vs merged human pull requests, judged blind","AI vs human pull requests: a blind multi-model review","","","","2026-09-28","2026-10-05",[],[],[],["agent-blind-review"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["blind-review-first-vs-latest","First attempt vs latest attempt","dot-range","rate",["agent-blind-review"],[]],["blind-review-votes-by-task","Blind panel votes per task (latest attempt)","stacked-bar","count",["agent-blind-review"],[]],["blind-review-dimension-scores","What the critics scored higher","grouped-bar","score",["agent-blind-review"],[]],["blind-review-critic-agreement","Does the judge’s model family matter?","dot-range","rate",["agent-blind-review"],[]]]},[],[],"\u0001"],["model-head-to-head","Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head","Haiku vs Sonnet vs Opus vs Fable vs Codex: speed and tokens","","","","2026-10-05","2026-10-05",[],[],[],["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["h2h-pass-rate","Pass rate on five validated tasks","dot-range","rate",["agent-provider-h2h"],[]],["h2h-total-latency","Total time per call","dot-range","seconds",["agent-provider-h2h"],[]],["h2h-first-useful-latency","Time to first useful output","dot-range","seconds",["agent-provider-h2h"],[]],["h2h-input-tokens","Input tokens per call: what the CLI sends","stacked-bar","tokens",["agent-provider-h2h"],[]],["h2h-output-tokens","Output tokens per call","grouped-bar","tokens",["agent-provider-h2h"],[]],["h2h-list-price-per-call","List-price cost per call (calculation)","dot-range","usd",["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"],[]],["h2h-speed-vs-cost","Speed vs list-price cost","scatter","seconds",["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"],[]],["h2h-cost-per-pass","List-price cost per passing answer (calculation)","bar","usd",["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"],[]],["h2h-frontier","Speed, cost and quality frontier","scatter","seconds",["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"],[]]]},[],[],"\u0001"],["hard-model-head-to-head","Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks","Hard tasks: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol","","","","2026-10-06","2026-10-06",[],[],[],["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["hard-h2h-pass-rate","Pass rate on eight hard tasks","Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","dot-range","rate","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.4583,0.2789,0.6493,24]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.6667,0.4671,0.8203,24]]}}],"Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.",["agent-provider-h2h-hard"]],["hard-h2h-outcomes","What happened on every call","\u0001","stacked-bar","count","\u0001",[],"\u0001",["agent-provider-h2h-hard"]],["hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","\u0001","dot-range","seconds","\u0001",[],"\u0001",["agent-provider-h2h-hard"]],["hard-h2h-first-useful-latency","Time to first useful output on hard tasks","\u0001","dot-range","seconds","\u0001",[],"\u0001",["agent-provider-h2h-hard"]],["hard-h2h-output-tokens","Output tokens per call on hard tasks","\u0001","grouped-bar","tokens","\u0001",[],"\u0001",["agent-provider-h2h-hard"]],["hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)","\u0001","bar","usd","\u0001",[],"\u0001",["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]],["hard-h2h-frontier","Quality vs cost frontier on hard tasks","\u0001","scatter","rate","\u0001",[],"\u0001",["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]]]},[],[],"\u0001"],["coding-agents-head-to-head","Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks","Claude Code vs Codex CLI: 6 coding tasks, hidden tests","","","","2026-10-06","2026-10-06",[],[],[],["agent-coding-agents","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","kind","unit","sourceIds","series","subtitle","whisker","yLabel","viz","note"],"$r":[["coding-agents-pass-rate","Coding sessions that passed every hidden check","dot-range","rate",["agent-coding-agents"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"],["coding-agents-wall-time","Time per coding session","dot-range","seconds",["agent-coding-agents"],[{"name":"Wall time per session","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",23.1,18.7,44.5,12,true],["Claude Opus 5.5 · Claude Code",56.9,29.8,185.8,12,"\u0001"],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",113.4,78.5,221.9,12,"\u0001"]]}}],"Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)","minmax","Seconds","LatencyLanes","CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval."],["coding-agents-time-by-task","Time per coding task","grouped-bar","seconds",["agent-coding-agents"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"],["coding-agents-tool-calls","Tool calls per coding session","dot-range","calls",["agent-coding-agents"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"],["coding-agents-token-mix","Tokens per session, as each CLI reports them","stacked-bar","tokens",["agent-coding-agents"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"],["coding-agents-diff-lines","Lines changed per session, by kind of file","stacked-bar","count",["agent-coding-agents"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"],["coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)","bar","usd",["agent-coding-agents","calc-repricing","price-anthropic","price-openai"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"]]},[],[],{"statIds":["coding-agents-time-ratio"]}],["effort-ladder","Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks","Effort ladder: does more AI effort buy quality?","","","","2026-10-06","2026-10-06",[],[],[],["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks","dot-range","rate",["agent-effort-ladder"],[]],["effort-ladder-total-latency","Total time per call by effort on hard tasks","dot-range","seconds",["agent-effort-ladder"],[]],["effort-ladder-time-by-effort","Median total time per call, by effort","line","seconds",["agent-effort-ladder"],[]],["effort-ladder-output-tokens","Output tokens per call by effort on hard tasks","grouped-bar","tokens",["agent-effort-ladder"],[]],["effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)","bar","usd",["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"],[]]]},[],[],"\u0001"],["caching-consistency","Prompt caching and run-to-run consistency in Claude Code and Codex CLI","Prompt caching savings and LLM consistency, measured","","","","2026-10-06","2026-10-06",[],[],[],["agent-caching-consistency","calc-cache-pricing","price-anthropic"],[],{"$k":["id","title","kind","unit","sourceIds","series","subtitle","yLabel","note","whisker"],"$r":[["caching-read-share-by-turn","Share of input read from the cache, by turn in a session","line","rate",["agent-caching-consistency"],[],"\u0001","\u0001","\u0001","\u0001"],["caching-cost-with-without","List-price cost of 5-question sessions with and without the cache (calculation)","grouped-bar","usd",["agent-caching-consistency","calc-cache-pricing","price-anthropic"],[],"\u0001","\u0001","\u0001","\u0001"],["caching-latency-first-vs-later","Time per turn: first turn vs later turns in a cached session","dot-range","seconds",["agent-caching-consistency"],[],"\u0001","\u0001","\u0001","\u0001"],["consistency-pass-rate","Same prompt, 10 times: strict pass rate","dot-range","rate",["agent-caching-consistency"],{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0,0,0.2775,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["JSON object",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0.1,0.0179,0.4042,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["Code fix",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,0.7225,1,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}]]},"One series per prompt; whiskers are 95% Wilson intervals","Passed","Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.","ci95"],["consistency-distinct-answers","Same prompt, 10 times: how many different answers","grouped-bar","count",["agent-caching-consistency"],[],"\u0001","\u0001","\u0001","\u0001"],["consistency-latency-spread","Same prompt, 10 times: time per call","dot-range","seconds",["agent-caching-consistency"],{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",5.06,4.42,6.2,10],["Claude Sonnet 5.5 · Claude Code",6.89,5.81,7.81,10],["GPT-6.1 Sol (medium) · Codex CLI",13.38,12.29,17.97,10]]}],["JSON object",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",7.03,5.28,12.27,10],["Claude Sonnet 5.5 · Claude Code",2.89,2.68,5.3,10],["GPT-6.1 Sol (medium) · Codex CLI",6.42,5.25,8.26,10]]}],["Code fix",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",5.95,4.89,7.33,10],["Claude Sonnet 5.5 · Claude Code",2.67,2.32,4.34,10],["GPT-6.1 Sol (medium) · Codex CLI",11.29,9.08,14.85,10]]}]]},"Median; whiskers = fastest and slowest of 10 calls","Seconds","Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.","minmax"]]},[],[],"\u0001"],["agent-memory","Does memory help Claude Code? 8 kinds of agent memory, tested","Does CLAUDE.md help? Agent memory tested on Claude Code","","","","2026-10-06","2026-10-06",[],[],[],["agent-memory-study"],[],{"$k":["id","title","subtitle","kind","unit","whisker","series","note","sourceIds","polarity"],"$r":[["memory-full-pass","Full pass rate by kind of memory","Hidden tests pass and every convention check passes · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["No memory",0.6,0.3575,0.8018,15,"\u0001"],["/init CLAUDE.md",0.6,0.3575,0.8018,15,"\u0001"],["Curated, 11 lines",1,0.7961,1,15,true],["Raw notes, 60 lines",0.9333,0.7018,0.9881,15,"\u0001"],["Dreamed notes",1,0.7961,1,15,"\u0001"],["Handbook, 210 lines",1,0.7961,1,15,"\u0001"],["Stop hook only",0.8,0.5481,0.9295,15,"\u0001"],["Curated + hook",1,0.7961,1,15,"\u0001"]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.2,0.0567,0.5098,10],["/init CLAUDE.md",0.2,0.0567,0.5098,10],["Curated, 11 lines",0.7,0.3968,0.8922,10],["Raw notes, 60 lines",0.6,0.3127,0.8318,10],["Dreamed notes",0.7,0.3968,0.8922,10],["Handbook, 210 lines",0.3,0.1078,0.6032,10],["Stop hook only",0.8,0.4902,0.9433,10],["Curated + hook",0.9,0.5958,0.9821,10]]}}],"Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.",["agent-memory-study"],"\u0001"],["memory-knowledge-class","Where memory helps: what the repo shows vs what only the team knows","\u0001","grouped-bar","rate","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-team-knowledge-by-model","Team knowledge followed, Sonnet vs Haiku","Changelog rule and late-fee rate, pooled · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.4,0.1982,0.6425,15],["/init CLAUDE.md",0.6667,0.4171,0.8482,15],["Curated, 11 lines",1,0.7961,1,15],["Raw notes, 60 lines",1,0.7961,1,15],["Dreamed notes",1,0.7961,1,15],["Handbook, 210 lines",1,0.7961,1,15],["Stop hook only",0.6667,0.4171,0.8482,15],["Curated + hook",1,0.7961,1,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0,0,0.2775,10],["/init CLAUDE.md",0.1,0.0179,0.4042,10],["Curated, 11 lines",0.8,0.4902,0.9433,10],["Raw notes, 60 lines",0.6,0.3127,0.8318,10],["Dreamed notes",0.8,0.4902,0.9433,10],["Handbook, 210 lines",0.3,0.1078,0.6032,10],["Stop hook only",0.8,0.4902,0.9433,10],["Curated + hook",1,0.7225,1,10]]}}],"Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.",["agent-memory-study"],"\u0001"],["memory-late-fee","\"Charge our standard late fee\": what Sonnet 5.5 did","\u0001","stacked-bar","count","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-late-fee-haiku","\"Charge our standard late fee\": what Haiku 4.5 did","\u0001","stacked-bar","count","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-broken-test-command","A stale README command: who still ran it?","Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.8,0.5481,0.9295,15],["/init CLAUDE.md",0.8667,0.6212,0.9626,15],["Curated, 11 lines",0,0,0.2039,15],["Raw notes, 60 lines",0,0,0.2039,15],["Dreamed notes",0,0,0.2039,15],["Handbook, 210 lines",0,0,0.2039,15],["Stop hook only",0.6,0.3575,0.8018,15],["Curated + hook",0,0,0.2039,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",1,0.7225,1,10],["/init CLAUDE.md",1,0.7225,1,10],["Curated, 11 lines",0,0,0.2775,10],["Raw notes, 60 lines",1,0.7225,1,10],["Dreamed notes",0.1,0.0179,0.4042,10],["Handbook, 210 lines",0,0,0.2775,10],["Stop hook only",1,0.7225,1,10],["Curated + hook",0,0,0.2775,10]]}}],"The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old \"run npm test\" note and a later correction. The /init file repeats the README.",["agent-memory-study"],"lower"],["memory-input-tokens","What memory costs in context","\u0001","bar","tokens","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-cost-per-full-pass","List-price cost per fully correct result (calculation)","\u0001","grouped-bar","usd","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-wall-time","Time per session","\u0001","grouped-bar","seconds","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"],["memory-dreaming-scorecard","Dreaming: what one consolidation pass kept and dropped","\u0001","grouped-bar","count","\u0001",[],"\u0001",["agent-memory-study"],"\u0001"]]},[],[],{"statIds":["memory-team-knowledge-none","memory-team-knowledge-curated"]}],["system-one-arena","System One arena: Jev vs Clef and five open decision models, head to head","Jev vs Clef: 7 decision models on 1,085 decisions","","","","2026-10-06","2026-10-06",[],[],[],["system-one-arena"],[],{"$k":["id","title","kind","unit","sourceIds","series","polarity"],"$r":[["arena-accuracy","Who decides right? Accuracy on 1,000+ checkable decisions","dot-range","rate",["system-one-arena"],[],"\u0001"],["arena-by-suite","Where each model is strong","grouped-bar","rate",["system-one-arena"],[],"\u0001"],["arena-latency-this-mac","Speed on this Mac: the six open models","dot-range","ms",["system-one-arena"],[],"lower"],["arena-latency-same-gpu","Speed on one GPU: third-party numbers","dot-range","ms",["system-one-arena"],[],"lower"],["arena-speed-accuracy-gpu","Open models: accuracy against speed on the same GPU","scatter","rate",["system-one-arena"],[],"none"],["arena-latency-gateway","Speed through one gateway: OpenRouter's own numbers","bar","seconds",["system-one-arena"],[],"lower"],["arena-cost-same-provider","Price per 1,000 decisions at one provider's list prices","bar","usd",["system-one-arena"],[],"lower"],["arena-size-accuracy","Does a bigger file decide better?","scatter","rate",["system-one-arena"],[],"none"],["arena-robust","Same question, different presentation","dot-range","rate",["system-one-arena"],[],"\u0001"],["arena-flips","Decisions that changed when only the presentation changed","grouped-bar","rate",["system-one-arena"],[],"lower"],["arena-escape","Knowing when to say \"none of these\"","grouped-bar","rate",["system-one-arena"],[],"none"],["arena-injection","Prompt injection: does text in the state hijack the decision?","dot-range","rate",["system-one-arena"],[],"\u0001"],["arena-by-length","Short inputs vs long inputs","grouped-bar","rate",["system-one-arena"],[],"\u0001"],["arena-confident-wrong","Wrong and sure of it","dot-range","rate",["system-one-arena"],[],"lower"],["arena-elo","Tournament rating across every game","dot-range","score",["system-one-arena"],[],"\u0001"],["arena-wdl","Wins, draws and losses in the round robin","stacked-bar","count",["system-one-arena"],[],"\u0001"],["arena-perfect-moves","How often a model found the perfect move","grouped-bar","rate",["system-one-arena"],[],"\u0001"],["arena-pong","Pong as deployed: who won","dot-range","rate",["system-one-arena"],[],"\u0001"],["arena-pong-quality","Pong decision quality, ignoring time","dot-range","rate",["system-one-arena"],[],"\u0001"],["arena-pong-ladder","How much is a millisecond worth in Pong?","dot-range","rate",["system-one-arena"],[],"\u0001"]]},[],[],"\u0001"],["routing-jev-vs-llm","Jev vs Claude as a router: accuracy and cost","Jev router vs Claude Haiku and Sonnet: routing and cost","","","","2026-10-05","2026-10-06",[],[],[],["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"],[],{"$k":["id","title","kind","unit","sourceIds","series","subtitle","yLabel","note","whisker"],"$r":[["routing-exact-decisions","Typed routing decisions answered exactly right","dot-range","rate",["agent-routing","agent-jev-live"],[],"\u0001","\u0001","\u0001","\u0001"],["routing-key-accuracy","Per-question accuracy","dot-range","rate",["agent-routing","agent-jev-live"],[],"\u0001","\u0001","\u0001","\u0001"],["routing-exact-by-decision","Exact rate by decision type","grouped-bar","rate",["agent-routing","agent-jev-live"],[],"\u0001","\u0001","\u0001","\u0001"],["routing-cost-per-1000","Cost per 1,000 routing decisions","bar","usd",["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"],[],"\u0001","\u0001","\u0001","\u0001"],["routing-decision-latency","Time per routing decision","dot-range","ms",["agent-routing","agent-jev-live"],{"$k":["name","points"],"$r":[["Wall time (CLI)",[{"label":"Claude Haiku 4.5","value":12674,"lo":12674,"hi":34413,"n":82},{"label":"Claude Sonnet 5.5","value":2598,"lo":2598,"hi":4298,"n":82}]],["Wall time (direct API call)",[{"label":"Jev 1.13 (TypeSafe)","value":136.5,"lo":136.5,"hi":195.7,"n":246}]],["Model time (API)",[{"label":"Claude Haiku 4.5","value":10734,"lo":10734,"hi":32072,"n":82},{"label":"Claude Sonnet 5.5","value":1599,"lo":1599,"hi":2574,"n":82}]]]},"Median wall time, whisker to the 95th percentile","Time per decision","Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.","p50-p95"],["routing-economics-scenarios","Thought experiment: recorded agent work under different model mixes","bar","usd",["agent-routing","calc-repricing","price-anthropic"],[],"\u0001","\u0001","\u0001","\u0001"],["routing-economics-by-stage","Thought experiment: repriced cost by pipeline stage","grouped-bar","usd",["agent-routing","calc-repricing","price-anthropic"],[],"\u0001","\u0001","\u0001","\u0001"]]},[],[],"\u0001"],["routing-overhead","Routing overhead: deterministic policy vs LLM routers vs Jev","Routing overhead: rules vs LLM routers vs Jev","","","","2026-10-06","2026-10-06",[],[],[],["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"],[],{"$k":["id","title","subtitle","kind","unit","yLabel","whisker","series","note","sourceIds"],"$r":[["router-overhead-decision-latency","Time to make one routing decision","Median; whiskers = median to 95th percentile","dot-range","ms","Time per decision","p50-p95",[{"name":"Decision time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",0.00142,0.00142,0.00233,20000,true],["Jev 1.13 (TypeSafe)",136.5,136.5,195.7,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",2597,2597,4298,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",12543,12543,34481,82,"\u0001"]]}}],"The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.",["agent-routing-overhead","agent-routing","agent-jev-live"]],["router-overhead-cli-vs-model-time","Where an LLM router’s time goes: model vs CLI","Median per call; whiskers = median to 95th percentile","dot-range","ms","Time per call","p50-p95",[{"name":"Model API time","points":[{"label":"Claude Sonnet 5.5 (effort low, via Claude Code)","value":1596,"lo":1596,"hi":2583,"n":82},{"label":"Claude Haiku 4.5 (thinking on, via Claude Code)","value":10508,"lo":10508,"hi":32132,"n":82}]},{"name":"CLI and harness time","points":[{"label":"Claude Sonnet 5.5 (effort low, via Claude Code)","value":973,"lo":973,"hi":1277,"n":82},{"label":"Claude Haiku 4.5 (thinking on, via Claude Code)","value":1698,"lo":1698,"hi":2677,"n":82}]}],"Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.",["agent-routing-overhead","agent-routing"]],["router-overhead-completed","Routing calls that returned a decision","\u0001","dot-range","rate","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","agent-routing","agent-jev-live"]],["router-overhead-cost-reported","Cost per 1,000 routing decisions: no model call vs provider-reported","\u0001","bar","usd","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","agent-routing","price-jev"]],["router-overhead-cost-list-price","Cost per 1,000 routing decisions for the model routers (calculation)","\u0001","bar","usd","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]],["router-overhead-cost-per-1000-tasks","Added routing cost per 1,000 tasks (calculation)","\u0001","grouped-bar","usd","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]],["router-overhead-delay-per-task","Added routing delay per task (calculation)","\u0001","grouped-bar","seconds","\u0001","\u0001",[],"\u0001",["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]],["cli-startup-tax","CLI start-up tax on a one-word answer","Median of 5 runs; whiskers = fastest and slowest run","dot-range","ms","Time to output","minmax",{"$k":["name","points"],"$r":[["First output event",[{"label":"Claude Code · Claude Haiku 4.5","value":563,"lo":519,"hi":726,"n":5},{"label":"Codex CLI (default model)","value":489,"lo":354,"hi":1304,"n":5}]],["First model output",[{"label":"Claude Code · Claude Haiku 4.5","value":1461,"lo":1206,"hi":2308,"n":5},{"label":"Codex CLI (default model)","value":5059,"lo":4391,"hi":5478,"n":5}]],["Total wall time",[{"label":"Claude Code · Claude Haiku 4.5","value":2529,"lo":2273,"hi":3382,"n":5},{"label":"Codex CLI (default model)","value":5999,"lo":5367,"hi":6506,"n":5}]]]},"Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.",["agent-routing-overhead"]],["cli-startup-input-tokens","Input tokens a CLI sends for a one-word answer","\u0001","bar","tokens","\u0001","\u0001",[],"\u0001",["agent-routing-overhead"]]]},[],[],"\u0001"],["inference-provider-index","Inference provider index: 27 models, 52 providers","Inference provider prices: OpenRouter vs direct","","","","2026-10-06","2026-10-06",[],[],[],["openrouter-api-snapshot","openrouter-fees","price-anthropic","price-openai","price-google"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["provider-index-spread","How much more the priciest provider charges than the cheapest","bar","ratio",["openrouter-api-snapshot"],[]],["gateway-markup-vs-first-party","OpenRouter markup over the first-party list price","grouped-bar","percent",["openrouter-api-snapshot","openrouter-fees","price-anthropic","price-openai","price-google"],[]],["gateway-vs-direct-claude-haiku-4-5","Claude Haiku 4.5: OpenRouter vs Anthropic list price","grouped-bar","usd",["openrouter-api-snapshot","price-anthropic"],[]],["gateway-vs-direct-claude-sonnet-5","Claude Sonnet 5: OpenRouter vs Anthropic list price","grouped-bar","usd",["openrouter-api-snapshot","price-anthropic"],[]],["gateway-vs-direct-claude-sonnet-5-5","Claude Sonnet 5.5: OpenRouter vs Anthropic list price","grouped-bar","usd",["openrouter-api-snapshot","price-anthropic"],[]],["gateway-vs-direct-claude-opus-4-8","Claude Opus 4.8: OpenRouter vs Anthropic list price","grouped-bar","usd",["openrouter-api-snapshot","price-anthropic"],[]],["gateway-vs-direct-claude-opus-5","Claude Opus 5: OpenRouter vs Anthropic list price","grouped-bar","usd",["openrouter-api-snapshot","price-anthropic"],[]],["gateway-vs-direct-claude-opus-5-5","Claude Opus 5.5: OpenRouter vs Anthropic list price","grouped-bar","usd",["openrouter-api-snapshot","price-anthropic"],[]],["gateway-vs-direct-claude-fable-5-1","Claude Fable 5.1: OpenRouter vs Anthropic list price","grouped-bar","usd",["openrouter-api-snapshot","price-anthropic"],[]],["gateway-vs-direct-gpt-6-luna","GPT-6 Luna: OpenRouter vs OpenAI list price","grouped-bar","usd",["openrouter-api-snapshot","price-openai"],[]],["gateway-vs-direct-gemini-3-8-flash","Gemini 3.8 Flash: OpenRouter vs Google list price","grouped-bar","usd",["openrouter-api-snapshot","price-google"],[]],["gateway-vs-direct-gemini-3-5-flash","Gemini 3.5 Flash: OpenRouter vs Google list price","grouped-bar","usd",["openrouter-api-snapshot","price-google"],[]],["provider-prices-claude-haiku-4-5","Claude Haiku 4.5: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-claude-sonnet-5","Claude Sonnet 5: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-claude-opus-4-8","Claude Opus 4.8: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-claude-opus-5","Claude Opus 5: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-claude-opus-5-5","Claude Opus 5.5: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-claude-fable-5-1","Claude Fable 5.1: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-gpt-6-sol","GPT-6 Sol: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-gpt-6-luna","GPT-6 Luna: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-gpt-6-astra","GPT-6 Astra: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-gpt-5-5","GPT-5.5: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-gpt-oss-120b","gpt-oss-120b: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-gemini-3-8-flash","Gemini 3.8 Flash: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-gemini-3-5-flash","Gemini 3.5 Flash: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-gemini-3-5-flash-lite","Gemini 3.5 Flash Lite: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-gemini-3-1-pro-preview","Gemini 3.1 Pro Preview: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-llama-4-maverick","Llama 4 Maverick: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-deepseek-v4-pro","DeepSeek V4 Pro 0423: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-deepseek-v4-flash","DeepSeek V4 Flash 0423: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-kimi-k3","Kimi K3: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]],["provider-prices-glm-5-3","GLM 5.3: price per million tokens by provider","grouped-bar","usd",["openrouter-api-snapshot"],[]]]},[],[],"\u0001"],["cost-thought-experiments","What if every call ran on Opus? Repricing real agent tokens","Agent token costs repriced: Haiku vs Sonnet vs Opus vs Fable","","","","2026-10-05","2026-10-05",[],[],[],["calc-repricing","agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","price-anthropic","price-google","price-openai","price-jev"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["repriced-cost-per-resolved","Thought experiment: the same tokens at other list prices","bar","usd",["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic","price-google","price-openai","price-jev"],[]],["cost-per-resolved-agent-vs-panel","Recorded cost per resolved instance: Agent vs the public panel","bar","usd",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"],[]],["prompt-cache-savings","Thought experiment: what prompt caching saved","grouped-bar","usd",["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic"],[]],["token-cost-mix","Where the token dollars go","bar","usd",["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic"],[]],["router-overhead-jev","Thought experiment: the price of a routing decision on every call","bar","usd",["calc-repricing","price-jev","agent-swebench-c1","agent-swebench-c2"],[]]]},[],[],"\u0001"],["cli-model-latency-tokens","Claude Code CLI vs Codex CLI vs the API: latency and tokens","Claude Code vs Codex CLI vs API: latency and token overhead","","","","2026-10-03","2026-10-05",[],[],[],["agent-provider-explorer"],[],{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["cli-vs-api-exact-reply-latency","CLI vs API: time for a one-line answer","Matched cohort, fixed exact reply, 5 runs per configuration","dot-range","seconds","Seconds",[{"name":"Total time","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.97,0.65,1.5,5],["OpenAI API · GPT-6.1 Sol · low",1.02,0.96,1.87,5],["OpenAI API · GPT-6.1 Sol · high",1.52,1.35,2.23,5],["Codex CLI · GPT-6 Luna · none",3.19,2.88,3.83,5],["Codex CLI · GPT-6.1 Sol · low",4.18,3.86,4.53,5],["Codex CLI · GPT-6.1 Sol · high",4.19,3.81,4.69,5]]}},{"name":"First useful output","points":{"$k":["label","value","lo","hi","n"],"$r":[["OpenAI API · GPT-6 Luna · none",0.82,0.51,1.37,5],["OpenAI API · GPT-6.1 Sol · low",0.87,0.84,1.74,5],["OpenAI API · GPT-6.1 Sol · high",1.34,1.26,2.12,5],["Codex CLI · GPT-6 Luna · none",2.79,2.46,3.42,5],["Codex CLI · GPT-6.1 Sol · low",3.75,3.44,4.1,5],["Codex CLI · GPT-6.1 Sol · high",3.79,3.37,4.3,5]]}}],"Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.",["agent-provider-explorer"]],["cli-vs-api-small-coding-latency","CLI vs API: time for a small coding task","\u0001","dot-range","seconds","\u0001",[],"\u0001",["agent-provider-explorer"]],["cli-vs-api-prompt-overhead","Hidden prompt: input tokens for the same one-line request","\u0001","bar","tokens","\u0001",[],"\u0001",["agent-provider-explorer"]],["scheduler-repair-claude-vs-codex","Repairing a scheduler: Claude Code vs Codex vs API","\u0001","dot-range","seconds","\u0001",[],"\u0001",["agent-provider-explorer"]],["scheduler-repair-output-tokens","Output tokens to repair the scheduler","\u0001","grouped-bar","tokens","\u0001",[],"\u0001",["agent-provider-explorer"]]]},[],[],"\u0001"],["coding-calibration","Coding calibration: what broke on three real pull requests","AI coding agent calibration on fastify, h3 and uvicorn tasks","","","","2026-10-02","2026-10-05",[],[],[],["agent-coding-calibration"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["calibration-outcomes-by-slice","Three real tasks, four platform builds","grouped-bar","count",["agent-coding-calibration"],[]],["calibration-cost-by-task","Notional model cost per task, by slice","grouped-bar","usd",["agent-coding-calibration"],[]],["calibration-minutes-by-task","Wall time per task, by slice","grouped-bar","minutes",["agent-coding-calibration"],[]],["calibration-guardrail-refusals","Guardrail refusals per slice","bar","count",["agent-coding-calibration"],[]]]},[],[],"\u0001"],["single-call-vs-agent-loop","Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks","Agent loop vs single call: does tool use improve accuracy?","","","","2026-10-06","2026-10-06",[],[],[],["agent-agent-loop","agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","sourceIds"],"$r":[["agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks","Same tasks and validators. Whiskers are 95% Wilson intervals","dot-range","rate","higher","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",0.4583,0.2789,0.6493,24],["Claude Haiku 4.5 (agent loop) · Claude Code",0.5417,0.3507,0.7211,24],["Claude Sonnet 5.5 (single call) · Claude Code",1,0.862,1,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",1,0.8064,1,16],["GPT-6 Luna (single call) · Codex CLI",0.625,0.3864,0.8152,16],["GPT-6 Luna (agent loop) · Codex CLI",0.8571,0.6006,0.9599,14]]}}],"Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.","ci95",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-by-task","Strict passes per task: single call vs agent loop","\u0001","grouped-bar","rate","higher","\u0001",[],"\u0001","\u0001",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-total-time","Total time per attempt: single call vs agent loop","Median per configuration; whiskers = fastest and slowest attempt","dot-range","seconds","\u0001","Seconds",[{"name":"Total time per attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (single call) · Claude Code",39.01,15.27,75.13,24],["Claude Haiku 4.5 (agent loop) · Claude Code",56.77,24.53,223.7,24],["Claude Sonnet 5.5 (single call) · Claude Code",7.75,2.26,34.79,24],["Claude Sonnet 5.5 (agent loop) · Claude Code",7.41,2.75,24.19,16],["GPT-6 Luna (single call) · Codex CLI",5.16,3.59,11.32,16],["GPT-6 Luna (agent loop) · Codex CLI",9.32,3.78,15.89,14]]}}],"Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.","minmax",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-tokens","Tokens per attempt: single call vs agent loop","\u0001","grouped-bar","tokens","\u0001","\u0001",[],"\u0001","\u0001",["agent-agent-loop","agent-provider-h2h-hard"]],["agent-loop-tool-calls","Tool calls per agent-loop attempt","\u0001","dot-range","count","\u0001","\u0001",[],"\u0001","\u0001",["agent-agent-loop"]],["agent-loop-cost-per-pass","List-price cost per strict pass: single call vs agent loop (calculation)","\u0001","bar","usd","\u0001","\u0001",[],"\u0001","\u0001",["agent-agent-loop","agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]]]},[],[],"\u0001"],["haiku-thinking-on-off","Does thinking pay for Claude Haiku 4.5? Thinking on vs off","Claude Haiku 4.5 thinking on vs off: benchmark results","","","","2026-10-06","2026-10-06",[],[],[],["agent-haiku-thinking","agent-routing","calc-repricing","price-anthropic","agent-provider-h2h-hard"],[],{"$k":["id","title","kind","unit","polarity","sourceIds","series","subtitle","yLabel","note","whisker","factContext"],"$r":[["haiku-thinking-router-exact","Haiku thinking study: typed routing decisions answered exactly right","dot-range","rate","higher",["agent-haiku-thinking","agent-routing"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"],["haiku-thinking-router-latency","Haiku thinking study: time per routing decision","dot-range","seconds","\u0001",["agent-haiku-thinking","agent-routing"],[{"name":"Wall time (CLI)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",4.66,4.66,8.18,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",12.54,12.54,34.48,82,false],["Claude Sonnet 5.5 (low) · Claude Code",2.6,2.6,4.3,82,false]]}},{"name":"Model time (API)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",3.79,3.79,7.43,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",10.51,10.51,32.13,82,false],["Claude Sonnet 5.5 (low) · Claude Code",1.6,1.6,2.58,82,false]]}}],"Median wall time and model (API) time; whisker to the 95th percentile","Seconds","Whiskers run from p50 to p95, not a confidence interval. Percentiles use the nearest-rank rule of the routing-overhead study, so the thinking-on and Sonnet medians match that study (the Jev-vs-LLM page interpolates between ranks and shows slightly different values). One call at a time through the Claude Code CLI; the thinking-on and Sonnet arms ran on another day.","p50-p95","typed routing decisions, thinking on vs off"],["haiku-thinking-router-tokens","Haiku thinking study: thinking and visible output tokens per routing decision","grouped-bar","tokens","\u0001",["agent-haiku-thinking","agent-routing"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"],["haiku-thinking-router-cost","Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)","bar","usd","\u0001",["agent-haiku-thinking","agent-routing","calc-repricing","price-anthropic"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"],["haiku-thinking-hard-pass","Haiku thinking study: pass rate on eight hard tasks","dot-range","rate","higher",["agent-haiku-thinking","agent-provider-h2h-hard"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"],["haiku-thinking-hard-time","Haiku thinking study: total time per call on hard tasks","dot-range","seconds","\u0001",["agent-haiku-thinking","agent-provider-h2h-hard"],[],"\u0001","\u0001","\u0001","\u0001","\u0001"]]},[],[],{"statIds":["haiku-thinking-exact-off","haiku-thinking-exact-on"],"testStatId":"haiku-thinking-mcnemar"}],["json-schema-vs-instructions","Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI","LLM structured output: JSON schema vs instructions","","","","2026-10-07","2026-10-07",[],[],[],["agent-structured-output"],[],{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","sourceIds"],"$r":[["structured-output-pass-rate","Does a JSON schema raise the pass rate? Instructions vs schema mode","Three extraction prompts pooled; whiskers are 95% Wilson intervals","dot-range","rate","higher","Passed",[{"name":"Strict pass: the whole reply is the right JSON","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0,0,0.138,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}},{"name":"Right answer in any format (strict pass or format miss)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0.7083,0.5083,0.8509,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}}],"Whiskers are 95% Wilson intervals (a calculation) over 24 calls and 12 calls per configuration; every error counts as a fail. Strict: the whole reply parses as JSON and matches the expected answer exactly. A format miss is a right answer inside a code fence or prose, so it is never a strict pass.","ci95",["agent-structured-output"]],["structured-output-outcomes","What each call produced: strict pass, format miss, wrong values or error","\u0001","stacked-bar","count","\u0001","\u0001",[],"\u0001","\u0001",["agent-structured-output"]],["structured-output-time","Time per call, instructions vs schema mode","Median; whiskers = fastest and slowest completed call","dot-range","seconds","\u0001","Seconds",[{"name":"Median time per call (the three prompts pooled)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",9.52,5.67,17,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",8.46,5.9,12.23,24],["Claude Sonnet 5.5 (instructions) · Claude Code",3.52,2.67,4.12,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",4.2,2.95,6.14,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",6.21,4.2,12.27,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",5.96,4.62,20.97,12]]}}],"Whiskers are a range (fastest and slowest call), not a confidence interval. Wall time from process start to exit, so it includes CLI start-up; Codex CLI timings include its larger system prompt. The three prompts differ in length, which widens every range.","minmax",["agent-structured-output"]],["structured-output-tokens","Output and reasoning tokens per call, instructions vs schema mode","\u0001","grouped-bar","tokens","\u0001","\u0001",[],"\u0001","\u0001",["agent-structured-output"]]]},[],[],"\u0001"],["prompt-cache-across-sessions","Does a new Claude Code session reuse the prompt cache of an earlier one?","Claude Code prompt caching across sessions, tested","","","","2026-10-07","2026-10-07",[],[],[],["agent-cache-sessions","calc-cache-pricing","price-anthropic","agent-caching-consistency"],[],{"$k":["id","title","kind","unit","polarity","sourceIds","series"],"$r":[["cache-sessions-turn1-read-share","Claude Sonnet 5.5 · Claude Code: share of turn-1 input read from the cache, by setup and session","grouped-bar","rate","none",["agent-cache-sessions"],[]],["cache-sessions-turn1-cost","Claude Sonnet 5.5 · Claude Code: list-price cost of turn 1, by setup and session (calculation)","grouped-bar","usd","\u0001",["agent-cache-sessions","calc-cache-pricing","price-anthropic"],[]],["cache-sessions-turn-time","Claude Sonnet 5.5 · Claude Code: time per turn, by setup","dot-range","seconds","\u0001",["agent-cache-sessions"],[]],["cache-sessions-codex-turn1-cached","GPT-6.1 Sol (medium) · Codex CLI: cached input tokens on turn 1, by setup and session","grouped-bar","tokens","\u0001",["agent-cache-sessions"],[]]]},[],[],{"statIds":["cache-sessions-later-reuse-a","cache-sessions-later-reuse-b"]}],["prompt-cache-break-even","Prompt cache break-even: after how many reuses does a cached prefix cost less?","Prompt cache break-even by model and session length","","","","2026-10-06","2026-10-06",[],[],[],["calc-cache-pricing","agent-caching-consistency","price-anthropic","price-openai"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["cache-break-even-reads","Reuses before a cached prefix costs less, by model and write type (calculation)","grouped-bar","score",["calc-cache-pricing","price-anthropic","price-openai"],[]],["cache-break-even-cost-curve","Cost of a reused prefix with and without the cache, by session length (calculation)","line","usd",["calc-cache-pricing","agent-caching-consistency","price-anthropic","price-openai"],[]],["cache-break-even-session-split","One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation)","grouped-bar","usd",["calc-cache-pricing","agent-caching-consistency","price-anthropic","price-openai"],[]]]},[],[],"\u0001"],["routing-holdout","Jev vs Claude routers on unseen decisions: a blind holdout","Jev vs Claude router accuracy on unseen decisions","","","","2026-10-06","2026-10-06",[],[],[],["agent-routing-holdout","agent-routing","calc-repricing","price-jev","price-anthropic"],[],{"$k":["id","title","kind","unit","polarity","sourceIds","series","subtitle","yLabel","note","whisker"],"$r":[["routing-holdout-exact","Unseen routing decisions answered exactly right","dot-range","rate","higher",["agent-routing-holdout"],[],"\u0001","\u0001","\u0001","\u0001"],["routing-holdout-key-accuracy","Per-question accuracy on unseen decisions","dot-range","rate","higher",["agent-routing-holdout"],[],"\u0001","\u0001","\u0001","\u0001"],["routing-holdout-by-purpose","Exact rate on unseen decisions, by decision type","grouped-bar","rate","higher",["agent-routing-holdout"],[],"\u0001","\u0001","\u0001","\u0001"],["routing-holdout-tuned-vs-unseen","Tuned case set vs unseen holdout: exact rate per router","grouped-bar","rate","higher",["agent-routing-holdout","agent-routing"],[],"\u0001","\u0001","\u0001","\u0001"],["routing-holdout-latency","Time per routing decision, by route","dot-range","seconds","\u0001",["agent-routing-holdout"],[{"name":"Wall time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.139,0.139,0.192,168,true],["Claude Haiku 4.5 · Claude Code",9.444,9.444,25.413,56,false],["Claude Sonnet 5.5 (low) · Claude Code",2.359,2.359,3.657,56,false]]}},{"name":"Model time (API, CLI-reported)","points":[{"label":"Claude Haiku 4.5 · Claude Code","value":7.522,"lo":7.522,"hi":23.913,"n":56},{"label":"Claude Sonnet 5.5 (low) · Claude Code","value":1.485,"lo":1.485,"hi":2.377,"n":56}]}],"Median, whisker to the 95th percentile","Seconds","The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.","p50-p95"],["routing-holdout-cost-per-1000","Cost per 1,000 unseen routing decisions","bar","usd","\u0001",["agent-routing-holdout","calc-repricing","price-jev","price-anthropic"],[],"\u0001","\u0001","\u0001","\u0001"]]},[],[],"\u0001"],["thinking-token-bill","How much of an AI bill is thinking? Reasoning tokens by model and effort","Thinking tokens by model and effort: share and cost","","","","2026-10-06","2026-10-06",[],[],[],["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],[],{"$k":["id","title","kind","unit","polarity","sourceIds","series"],"$r":[["thinking-bill-share","Reasoning share of output tokens per call on hard tasks (calculation)","bar","percent","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],[]],["thinking-bill-cost-per-call","List-price cost per call: reasoning, remaining output and input (calculation)","stacked-bar","usd","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],[]],["thinking-bill-by-effort","Reasoning cost per strict pass by effort, with the total (calculation)","grouped-bar","usd","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],[]],["thinking-bill-vs-time","Reasoning tokens vs total time per call (calculation)","scatter","seconds","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],[]],["thinking-bill-short-vs-hard","Reasoning share on short tasks vs hard tasks (calculation)","grouped-bar","percent","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],[]]]},[],[],{"statIds":["thinking-bill-share-highest","thinking-bill-share-lowest"]}],["llm-speed-anatomy","Where the seconds go: first text, output speed and prompt size for 6 LLMs","LLM speed: first text and tokens per second through CLIs","","","","2026-10-07","2026-10-07",[],[],[],["agent-speed-anatomy"],[],{"$k":["id","title","kind","unit","sourceIds","series","polarity"],"$r":[["speed-anatomy-first-text","Time to first text: a 250-line answer, six models","dot-range","seconds",["agent-speed-anatomy"],[],"\u0001"],["speed-anatomy-output-speed","Output speed after the first text: visible tokens per second (calculation)","dot-range","tokens",["agent-speed-anatomy"],[],"\u0001"],["speed-anatomy-chars-per-second","Output speed in characters per second after the first text (calculation)","dot-range","count",["agent-speed-anatomy"],[],"\u0001"],["speed-anatomy-prompt-size","Time to first text as the prompt grows","line","seconds",["agent-speed-anatomy"],[],"\u0001"],["speed-anatomy-total-by-size","Total time per call by prompt size","grouped-bar","seconds",["agent-speed-anatomy"],[],"\u0001"],["speed-anatomy-lookup-correct","Exact lookup answers at the 1k, 16k and 64k prompt-size targets","dot-range","rate",["agent-speed-anatomy"],[],"higher"]]},[],[],{"statIds":["speed-anatomy-token-count-ratio"]}],["haiku-retry-or-escalate","Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts","Haiku retry and escalate: cost per correct answer","","","","2026-10-07","2026-10-07",[],[],[],["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],[],{"$k":["id","title","kind","unit","sourceIds","series","polarity"],"$r":[["retry-escalate-cost-per-correct","Expected list-price cost per correct answer, by retry policy (calculation)","bar","usd",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],[],"\u0001"],["retry-escalate-time-per-correct","Expected wall time per correct answer, by retry policy (calculation)","bar","seconds",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],[],"\u0001"],["retry-escalate-success","Expected success rate by retry policy (calculation)","bar","rate",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],[],"higher"],["retry-escalate-haiku-by-task","Strict pass rate of Haiku 4.5 on each hard task","bar","rate",["agent-provider-h2h-hard"],[],"higher"],["retry-escalate-call-cost-by-task","List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation)","grouped-bar","usd",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],[],"\u0001"],["retry-escalate-sensitivity","Sensitivity: cost per correct answer if Haiku’s pass rate is higher or lower (calculation)","dot-range","usd",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],[],"\u0001"]]},[],[],"\u0001"],["harder-tasks-head-to-head","GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks","GPT-6.1 Sol vs Claude Opus 5.5 on harder tasks","","","","2026-10-07","2026-10-07",[],[],[],["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","whisker","series","note","sourceIds"],"$r":[["harder-h2h-pass-rate","Pass rate on 4 harder tasks","Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","dot-range","rate","higher","Passed","ci95",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.5,0.2538,0.7462,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}}],"Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row.\n\nCounted calls are new calls.",["agent-harder-tasks"]],["harder-h2h-outcomes","What happened on every call","\u0001","stacked-bar","count","\u0001","\u0001","\u0001",[],"\u0001",["agent-harder-tasks"]],["harder-h2h-tool-attempts","Calls that tried a tool although tools were off","\u0001","dot-range","rate","none","\u0001","\u0001",[],"\u0001",["agent-harder-tasks"]],["harder-h2h-pass-by-task","Strict pass rate by task","One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide","grouped-bar","rate","higher","Strict pass rate","ci95",{"$k":["name","points"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",0.75,0.3006,0.9544,4],["Sudoku, 22 givens",0.25,0.0456,0.6994,4],["6x6 Skyscrapers",1,0.5101,1,4],["Seeded shuffle output",0.75,0.3006,0.9544,4]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",1,0.4385,1,3],["Sudoku, 22 givens",0,0,0.5615,3],["6x6 Skyscrapers",0.3333,0.0615,0.7923,3],["Seeded shuffle output",0.3333,0.0615,0.7923,3]]}],["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",1,0.5101,1,4],["Sudoku, 22 givens",0,0,0.4899,4],["6x6 Skyscrapers",0,0,0.4899,4],["Seeded shuffle output",0.5,0.15,0.85,4]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",0,0,0.5615,3],["Sudoku, 22 givens",0,0,0.5615,3],["6x6 Skyscrapers",0,0,0.5615,3],["Seeded shuffle output",0,0,0.5615,3]]}]]},"Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.",["agent-harder-tasks"]],["harder-h2h-total-latency","Total time per call on harder tasks","\u0001","dot-range","seconds","\u0001","\u0001","\u0001",[],"\u0001",["agent-harder-tasks"]],["harder-h2h-output-tokens","Output tokens per call on harder tasks","\u0001","grouped-bar","tokens","none","\u0001","\u0001",[],"\u0001",["agent-harder-tasks"]],["harder-h2h-cost-per-pass","List-price cost per strict pass on harder tasks (calculation)","\u0001","bar","usd","\u0001","\u0001","\u0001",[],"\u0001",["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"]],["harder-h2h-frontier","Observed quality vs cost frontier (calculation)","\u0001","scatter","rate","higher","\u0001","\u0001",[],"\u0001",["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"]]]},[],[],"\u0001"],["voice-agent-latency-budget","Voice agent latency budget: component calculations, not a measured turn","Voice agent latency budget: calculated component times","","","","2026-10-06","2026-10-06",[],[],[],["calc-latency-budget","agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["latency-budget-fast-steps","Steps that take under a second, against three assumed budgets (median to p95)","dot-range","ms",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"],[]],["latency-budget-fast-steps-ranges","Steps that take under a second, against three assumed budgets (observed ranges)","dot-range","ms",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"],[]],["latency-budget-slow-steps","Steps that take a second or more, against three assumed budgets (median to p95)","dot-range","ms",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"],[]],["latency-budget-slow-steps-ranges","Steps that take a second or more, against three assumed budgets (observed ranges)","dot-range","ms",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"],[]],["latency-budget-fit","How many of the steps fit each assumed budget (calculation)","grouped-bar","count",["calc-latency-budget","agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"],[]]]},[],[],"\u0001"],["routing-at-scale","What does routing a million AI requests a day cost? A calculation from measured runs","LLM router cost at scale: 1 million decisions a day","","","","2026-10-06","2026-10-06",[],[],[],["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"],[],{"$k":["id","title","kind","unit","sourceIds","series"],"$r":[["routing-at-scale-daily-cost","Daily cost of routing at 10,000 to 10 million decisions a day (calculation)","grouped-bar","usd",["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev"],[]],["routing-at-scale-in-flight","In-flight scenarios at 10,000 to 10 million a day (calculation)","grouped-bar","calls",["calc-routing-at-scale","agent-routing-overhead","agent-routing","agent-jev-live"],[]],["routing-at-scale-waiting-hours","Waiting scenarios per day at 1 million decisions a day (calculation)","bar","count",["calc-routing-at-scale","agent-routing-overhead","agent-routing","agent-jev-live"],[]],["routing-at-scale-scenarios","Daily cost at 1 million model calls a day: route every call or only System One decisions (calculation)","grouped-bar","usd",["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev"],[]]]},[],[],"\u0001"]]},"entities":{"$k":["slug","name","vendor","kind","description","aliases","facts"],"$r":[["claude-sonnet-5-5","Claude Sonnet 5.5","Anthropic","model","Anthropic’s mid-tier Claude model. Agent’s default coding model; measured here through Claude Code at low, medium, high and default effort, and as an LLM router.",["Claude Sonnet 5.5","Sonnet 5.5"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","polarity","calculation"],"$r":[["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Sonnet 5.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Sonnet 5.5 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001"],["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","Claude Sonnet 5.5 · Claude Code","coding-agents-wall-time","Time per coding session",23.1,"seconds","23.1 s",12,"\u0001","minmax","Claude Code · six small repository tasks with hidden tests",[18.7,44.5],"\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Exact number","Claude Sonnet 5.5 · Claude Code","consistency-pass-rate/Exact number","Same prompt, 10 times: strict pass rate (Exact number)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","JSON object","Claude Sonnet 5.5 · Claude Code","consistency-pass-rate/JSON object","Same prompt, 10 times: strict pass rate (JSON object)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",6.89,"seconds","6.89 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[5.81,7.81],"\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Code fix","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",2.67,"seconds","2.67 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[2.32,4.34],"\u0001","\u0001"],["agent-memory","memory-full-pass","Claude Sonnet 5.5","Handbook, 210 lines","memory-full-pass/Handbook, 210 lines","Full pass rate by kind of memory: Handbook, 210 lines",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","","\u0001","\u0001","\u0001"],["agent-memory","memory-team-knowledge-by-model","Claude Sonnet 5.5","/init CLAUDE.md","memory-team-knowledge-by-model//init CLAUDE.md","Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md",0.6667,"rate","67% (10/15)",15,[0.4171,0.8482],"ci95","","\u0001","\u0001","\u0001"],["agent-memory","memory-team-knowledge-by-model","Claude Sonnet 5.5","Handbook, 210 lines","memory-team-knowledge-by-model/Handbook, 210 lines","Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","","\u0001","\u0001","\u0001"],["agent-memory","memory-broken-test-command","Claude Sonnet 5.5","Raw notes, 60 lines","memory-broken-test-command/Raw notes, 60 lines","A stale README command: who still ran it?: Raw notes, 60 lines",0,"rate","0% (0/15)",15,[0,0.2039],"ci95","","\u0001","lower","\u0001"],["routing-jev-vs-llm","routing-decision-latency","Wall time (CLI)","Claude Sonnet 5.5","routing-decision-latency/Wall time (CLI)","Time per routing decision (Wall time (CLI))",2598,"ms","2,598 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code",[2598,4298],"\u0001","\u0001"],["routing-jev-vs-llm","routing-decision-latency","Model time (API)","Claude Sonnet 5.5","routing-decision-latency/Model time (API)","Time per routing decision (Model time (API))",1599,"ms","1,599 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code",[1599,2574],"\u0001","\u0001"],["routing-overhead","router-overhead-decision-latency","Decision time","Claude Sonnet 5.5 (effort low, via Claude Code)","router-overhead-decision-latency","Time to make one routing decision",2597,"ms","2,597 ms",82,"\u0001","p50-p95","effort low · via Claude Code · routing overhead per decision",[2597,4298],"\u0001","\u0001"],["routing-overhead","router-overhead-cli-vs-model-time","Model API time","Claude Sonnet 5.5 (effort low, via Claude Code)","router-overhead-cli-vs-model-time/Model API time","Where an LLM router’s time goes: model vs CLI (Model API time)",1596,"ms","1,596 ms",82,"\u0001","p50-p95","effort low · via Claude Code · routing overhead per decision",[1596,2583],"\u0001","\u0001"],["routing-overhead","router-overhead-cli-vs-model-time","CLI and harness time","Claude Sonnet 5.5 (effort low, via Claude Code)","router-overhead-cli-vs-model-time/CLI and harness time","Where an LLM router’s time goes: model vs CLI (CLI and harness time)",973,"ms","973 ms",82,"\u0001","p50-p95","effort low · via Claude Code · routing overhead per decision",[973,1277],"\u0001","\u0001"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","Claude Sonnet 5.5 (single call) · Claude Code","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · single call","\u0001","higher","\u0001"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","Claude Sonnet 5.5 (agent loop) · Claude Code","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · agent loop","\u0001","higher","\u0001"],["haiku-thinking-on-off","haiku-thinking-router-latency","Wall time (CLI)","Claude Sonnet 5.5 (low) · Claude Code","haiku-thinking-router-latency/Wall time (CLI)","Haiku thinking study: time per routing decision (Wall time (CLI))",2.6,"seconds","2.60 s",82,"\u0001","p50-p95","Claude Code · effort low · typed routing decisions, thinking on vs off",[2.6,4.3],"\u0001","\u0001"],["haiku-thinking-on-off","haiku-thinking-router-latency","Model time (API)","Claude Sonnet 5.5 (low) · Claude Code","haiku-thinking-router-latency/Model time (API)","Haiku thinking study: time per routing decision (Model time (API))",1.6,"seconds","1.60 s",82,"\u0001","p50-p95","Claude Code · effort low · typed routing decisions, thinking on vs off",[1.6,2.58],"\u0001","\u0001"],["json-schema-vs-instructions","structured-output-pass-rate","Strict pass: the whole reply is the right JSON","Claude Sonnet 5.5 (instructions) · Claude Code","structured-output-pass-rate/Strict pass: the whole reply is the right JSON","Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","Claude Code · instructions","\u0001","higher",true],["json-schema-vs-instructions","structured-output-pass-rate","Strict pass: the whole reply is the right JSON","Claude Sonnet 5.5 (JSON schema) · Claude Code","structured-output-pass-rate/Strict pass: the whole reply is the right JSON","Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","Claude Code · JSON schema","\u0001","higher",true],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","Claude Sonnet 5.5 (instructions) · Claude Code","structured-output-time","Time per call, instructions vs schema mode",3.52,"seconds","3.52 s",12,"\u0001","minmax","Claude Code · instructions",[2.67,4.12],"\u0001","\u0001"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","Claude Sonnet 5.5 (JSON schema) · Claude Code","structured-output-time","Time per call, instructions vs schema mode",4.2,"seconds","4.20 s",12,"\u0001","minmax","Claude Code · JSON schema",[2.95,6.14],"\u0001","\u0001"],["routing-holdout","routing-holdout-latency","Wall time","Claude Sonnet 5.5 (low) · Claude Code","routing-holdout-latency/Wall time","Time per routing decision, by route (Wall time)",2.359,"seconds","2.36 s",56,"\u0001","p50-p95","Claude Code · effort low",[2.359,3.657],"\u0001","\u0001"],["routing-holdout","routing-holdout-latency","Model time (API, CLI-reported)","Claude Sonnet 5.5 (low) · Claude Code","routing-holdout-latency/Model time (API, CLI-reported)","Time per routing decision, by route (Model time (API, CLI-reported))",1.485,"seconds","1.49 s",56,"\u0001","p50-p95","Claude Code · effort low",[1.485,2.377],"\u0001","\u0001"],["harder-tasks-head-to-head","harder-h2h-pass-by-task","Claude Sonnet 5.5 · Claude Code","6x6 Skyscrapers","harder-h2h-pass-by-task/6x6 Skyscrapers","Strict pass rate by task: 6x6 Skyscrapers",0,"rate","0% (0/4)",4,[0,0.4899],"ci95","Claude Code","\u0001","higher","\u0001"]]}],["claude-opus-5-5","Claude Opus 5.5","Anthropic","model","Anthropic’s large Claude model, measured through Claude Code at low, medium, high and default effort.",["Claude Opus 5.5","Opus 5.5"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","polarity"],"$r":[["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Opus 5.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Opus 5.5 (high) · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · effort high · eight hard validated tasks","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Opus 5.5 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Opus 5.5 (high) · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · effort high · eight hard validated tasks","\u0001"],["harder-tasks-head-to-head","harder-h2h-pass-rate","Lenient (format misses counted)","Claude Opus 5.5 · Claude Code","harder-h2h-pass-rate/Lenient (format misses counted)","Pass rate on 4 harder tasks (Lenient (format misses counted))",0.5,"rate","50% (6/12)",12,[0.2538,0.7462],"ci95","Claude Code","higher"]]}],["claude-haiku-4-5","Claude Haiku 4.5","Anthropic","model","Anthropic’s small, low-price Claude model, measured through Claude Code, as an LLM router and in the public SWE-bench panel.",["Claude Haiku 4.5","Haiku 4.5","Claude 4.5 Haiku"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","polarity","calculation"],"$r":[["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Haiku 4.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",0.4583,"rate","46% (11/24)",24,[0.2789,0.6493],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Haiku 4.5 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,"rate","67% (16/24)",24,[0.4671,0.8203],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Exact number","Claude Haiku 4.5 · Claude Code","consistency-pass-rate/Exact number","Same prompt, 10 times: strict pass rate (Exact number)",0,"rate","0% (0/10)",10,[0,0.2775],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","JSON object","Claude Haiku 4.5 · Claude Code","consistency-pass-rate/JSON object","Same prompt, 10 times: strict pass rate (JSON object)",0.1,"rate","10% (1/10)",10,[0.0179,0.4042],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",5.06,"seconds","5.06 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[4.42,6.2],"\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Code fix","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",5.95,"seconds","5.95 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[4.89,7.33],"\u0001","\u0001"],["agent-memory","memory-full-pass","Claude Haiku 4.5","Handbook, 210 lines","memory-full-pass/Handbook, 210 lines","Full pass rate by kind of memory: Handbook, 210 lines",0.3,"rate","30% (3/10)",10,[0.1078,0.6032],"ci95","","\u0001","\u0001","\u0001"],["agent-memory","memory-team-knowledge-by-model","Claude Haiku 4.5","/init CLAUDE.md","memory-team-knowledge-by-model//init CLAUDE.md","Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md",0.1,"rate","10% (1/10)",10,[0.0179,0.4042],"ci95","","\u0001","\u0001","\u0001"],["agent-memory","memory-team-knowledge-by-model","Claude Haiku 4.5","Handbook, 210 lines","memory-team-knowledge-by-model/Handbook, 210 lines","Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines",0.3,"rate","30% (3/10)",10,[0.1078,0.6032],"ci95","","\u0001","\u0001","\u0001"],["agent-memory","memory-broken-test-command","Claude Haiku 4.5","Raw notes, 60 lines","memory-broken-test-command/Raw notes, 60 lines","A stale README command: who still ran it?: Raw notes, 60 lines",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","","\u0001","lower","\u0001"],["routing-jev-vs-llm","routing-decision-latency","Wall time (CLI)","Claude Haiku 4.5","routing-decision-latency/Wall time (CLI)","Time per routing decision (Wall time (CLI))",12674,"ms","12,674 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code",[12674,34413],"\u0001","\u0001"],["routing-jev-vs-llm","routing-decision-latency","Model time (API)","Claude Haiku 4.5","routing-decision-latency/Model time (API)","Time per routing decision (Model time (API))",10734,"ms","10,734 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code",[10734,32072],"\u0001","\u0001"],["routing-overhead","router-overhead-decision-latency","Decision time","Claude Haiku 4.5 (thinking on, via Claude Code)","router-overhead-decision-latency","Time to make one routing decision",12543,"ms","12,543 ms",82,"\u0001","p50-p95","thinking on · via Claude Code · routing overhead per decision",[12543,34481],"\u0001","\u0001"],["routing-overhead","router-overhead-cli-vs-model-time","Model API time","Claude Haiku 4.5 (thinking on, via Claude Code)","router-overhead-cli-vs-model-time/Model API time","Where an LLM router’s time goes: model vs CLI (Model API time)",10508,"ms","10,508 ms",82,"\u0001","p50-p95","thinking on · via Claude Code · routing overhead per decision",[10508,32132],"\u0001","\u0001"],["routing-overhead","router-overhead-cli-vs-model-time","CLI and harness time","Claude Haiku 4.5 (thinking on, via Claude Code)","router-overhead-cli-vs-model-time/CLI and harness time","Where an LLM router’s time goes: model vs CLI (CLI and harness time)",1698,"ms","1,698 ms",82,"\u0001","p50-p95","thinking on · via Claude Code · routing overhead per decision",[1698,2677],"\u0001","\u0001"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","Claude Haiku 4.5 (single call) · Claude Code","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.4583,"rate","46% (11/24)",24,[0.2789,0.6493],"ci95","Claude Code · single call","\u0001","higher","\u0001"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","Claude Haiku 4.5 (agent loop) · Claude Code","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.5417,"rate","54% (13/24)",24,[0.3507,0.7211],"ci95","Claude Code · agent loop","\u0001","higher","\u0001"],["single-call-vs-agent-loop","agent-loop-total-time","Total time per attempt","Claude Haiku 4.5 (single call) · Claude Code","agent-loop-total-time","Total time per attempt: single call vs agent loop",39.01,"seconds","39.0 s",24,"\u0001","minmax","Claude Code · single call",[15.27,75.13],"\u0001","\u0001"],["single-call-vs-agent-loop","agent-loop-total-time","Total time per attempt","Claude Haiku 4.5 (agent loop) · Claude Code","agent-loop-total-time","Total time per attempt: single call vs agent loop",56.77,"seconds","56.8 s",24,"\u0001","minmax","Claude Code · agent loop",[24.53,223.7],"\u0001","\u0001"],["haiku-thinking-on-off","haiku-thinking-router-latency","Wall time (CLI)","Claude Haiku 4.5 (thinking off) · Claude Code","haiku-thinking-router-latency/Wall time (CLI)","Haiku thinking study: time per routing decision (Wall time (CLI))",4.66,"seconds","4.66 s",82,"\u0001","p50-p95","Claude Code · thinking off · typed routing decisions, thinking on vs off",[4.66,8.18],"\u0001","\u0001"],["haiku-thinking-on-off","haiku-thinking-router-latency","Wall time (CLI)","Claude Haiku 4.5 (thinking on) · Claude Code","haiku-thinking-router-latency/Wall time (CLI)","Haiku thinking study: time per routing decision (Wall time (CLI))",12.54,"seconds","12.5 s",82,"\u0001","p50-p95","Claude Code · thinking on · typed routing decisions, thinking on vs off",[12.54,34.48],"\u0001","\u0001"],["haiku-thinking-on-off","haiku-thinking-router-latency","Model time (API)","Claude Haiku 4.5 (thinking off) · Claude Code","haiku-thinking-router-latency/Model time (API)","Haiku thinking study: time per routing decision (Model time (API))",3.79,"seconds","3.79 s",82,"\u0001","p50-p95","Claude Code · thinking off · typed routing decisions, thinking on vs off",[3.79,7.43],"\u0001","\u0001"],["haiku-thinking-on-off","haiku-thinking-router-latency","Model time (API)","Claude Haiku 4.5 (thinking on) · Claude Code","haiku-thinking-router-latency/Model time (API)","Haiku thinking study: time per routing decision (Model time (API))",10.51,"seconds","10.5 s",82,"\u0001","p50-p95","Claude Code · thinking on · typed routing decisions, thinking on vs off",[10.51,32.13],"\u0001","\u0001"],["json-schema-vs-instructions","structured-output-pass-rate","Strict pass: the whole reply is the right JSON","Claude Haiku 4.5 (instructions) · Claude Code","structured-output-pass-rate/Strict pass: the whole reply is the right JSON","Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",0,"rate","0% (0/24)",24,[0,0.138],"ci95","Claude Code · instructions","\u0001","higher",true],["json-schema-vs-instructions","structured-output-pass-rate","Strict pass: the whole reply is the right JSON","Claude Haiku 4.5 (JSON schema) · Claude Code","structured-output-pass-rate/Strict pass: the whole reply is the right JSON","Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",0.75,"rate","75% (18/24)",24,[0.551,0.88],"ci95","Claude Code · JSON schema","\u0001","higher",true],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","Claude Haiku 4.5 (instructions) · Claude Code","structured-output-time","Time per call, instructions vs schema mode",9.52,"seconds","9.52 s",24,"\u0001","minmax","Claude Code · instructions",[5.67,17],"\u0001","\u0001"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","Claude Haiku 4.5 (JSON schema) · Claude Code","structured-output-time","Time per call, instructions vs schema mode",8.46,"seconds","8.46 s",24,"\u0001","minmax","Claude Code · JSON schema",[5.9,12.23],"\u0001","\u0001"],["routing-holdout","routing-holdout-latency","Wall time","Claude Haiku 4.5 · Claude Code","routing-holdout-latency/Wall time","Time per routing decision, by route (Wall time)",9.444,"seconds","9.44 s",56,"\u0001","p50-p95","Claude Code",[9.444,25.413],"\u0001","\u0001"],["routing-holdout","routing-holdout-latency","Model time (API, CLI-reported)","Claude Haiku 4.5 · Claude Code","routing-holdout-latency/Model time (API, CLI-reported)","Time per routing decision, by route (Model time (API, CLI-reported))",7.522,"seconds","7.52 s",56,"\u0001","p50-p95","Claude Code",[7.522,23.913],"\u0001","\u0001"],["harder-tasks-head-to-head","harder-h2h-pass-rate","Strict pass","Claude Haiku 4.5 · Claude Code","harder-h2h-pass-rate/Strict pass","Pass rate on 4 harder tasks (Strict pass)",0,"rate","0% (0/12)",12,[0,0.2425],"ci95","Claude Code","\u0001","higher","\u0001"],["harder-tasks-head-to-head","harder-h2h-pass-rate","Lenient (format misses counted)","Claude Haiku 4.5 · Claude Code","harder-h2h-pass-rate/Lenient (format misses counted)","Pass rate on 4 harder tasks (Lenient (format misses counted))",0,"rate","0% (0/12)",12,[0,0.2425],"ci95","Claude Code","\u0001","higher","\u0001"]]}],["claude-fable-5-1","Claude Fable 5.1","Anthropic","model","Anthropic’s highest-priced Claude model in these studies, measured through Claude Code.",["Claude Fable 5.1","Fable 5.1"],[{"studySlug":"hard-model-head-to-head","chartId":"hard-h2h-pass-rate","series":"Strict pass","point":"Claude Fable 5.1 · Claude Code","metric":"hard-h2h-pass-rate/Strict pass","label":"Pass rate on eight hard tasks (Strict pass)","value":1,"unit":"rate","display":"100% (24/24)","n":24,"ci":[0.862,1],"spanKind":"ci95","context":"Claude Code · eight hard validated tasks"},{"studySlug":"hard-model-head-to-head","chartId":"hard-h2h-pass-rate","series":"Lenient (format misses counted)","point":"Claude Fable 5.1 · Claude Code","metric":"hard-h2h-pass-rate/Lenient (format misses counted)","label":"Pass rate on eight hard tasks (Lenient (format misses counted))","value":1,"unit":"rate","display":"100% (24/24)","n":24,"ci":[0.862,1],"spanKind":"ci95","context":"Claude Code · eight hard validated tasks"}]],["gpt-6-1-sol-codex-cli","GPT-6.1 Sol (Codex CLI)","OpenAI","model","OpenAI’s GPT-6.1 Sol model run through the Codex CLI at low, medium and high effort.",["GPT-6.1 Sol · Codex CLI"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation","polarity"],"$r":[["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","GPT-6.1 Sol (high) · Codex CLI","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001"],["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI","coding-agents-wall-time","Time per coding session",113.4,"seconds","113.4 s",12,"\u0001","minmax","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests",[78.5,221.9],"\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Exact number","GPT-6.1 Sol (medium) · Codex CLI","consistency-pass-rate/Exact number","Same prompt, 10 times: strict pass rate (Exact number)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","JSON object","GPT-6.1 Sol (medium) · Codex CLI","consistency-pass-rate/JSON object","Same prompt, 10 times: strict pass rate (JSON object)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",13.38,"seconds","13.4 s",10,"\u0001","minmax","Codex CLI · effort medium · same prompt repeated 10 times",[12.29,17.97],"\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Code fix","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",11.29,"seconds","11.3 s",10,"\u0001","minmax","Codex CLI · effort medium · same prompt repeated 10 times",[9.08,14.85],"\u0001","\u0001"],["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","Total time","Codex CLI · GPT-6.1 Sol · low","cli-vs-api-exact-reply-latency/Total time","CLI vs API: time for a one-line answer (Total time)",4.18,"seconds","4.18 s",5,"\u0001","minmax","Codex CLI · effort low · fixed exact reply, 5 runs",[3.86,4.53],"\u0001","\u0001"],["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","Total time","Codex CLI · GPT-6.1 Sol · high","cli-vs-api-exact-reply-latency/Total time","CLI vs API: time for a one-line answer (Total time)",4.19,"seconds","4.19 s",5,"\u0001","minmax","Codex CLI · effort high · fixed exact reply, 5 runs",[3.81,4.69],"\u0001","\u0001"],["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","First useful output","Codex CLI · GPT-6.1 Sol · low","cli-vs-api-exact-reply-latency/First useful output","CLI vs API: time for a one-line answer (First useful output)",3.75,"seconds","3.75 s",5,"\u0001","minmax","Codex CLI · effort low · fixed exact reply, 5 runs",[3.44,4.1],"\u0001","\u0001"],["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","First useful output","Codex CLI · GPT-6.1 Sol · high","cli-vs-api-exact-reply-latency/First useful output","CLI vs API: time for a one-line answer (First useful output)",3.79,"seconds","3.79 s",5,"\u0001","minmax","Codex CLI · effort high · fixed exact reply, 5 runs",[3.37,4.3],"\u0001","\u0001"],["json-schema-vs-instructions","structured-output-pass-rate","Strict pass: the whole reply is the right JSON","GPT-6.1 Sol (low, instructions) · Codex CLI","structured-output-pass-rate/Strict pass: the whole reply is the right JSON","Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","Codex CLI · effort low · instructions","\u0001",true,"higher"],["json-schema-vs-instructions","structured-output-pass-rate","Strict pass: the whole reply is the right JSON","GPT-6.1 Sol (low, JSON schema) · Codex CLI","structured-output-pass-rate/Strict pass: the whole reply is the right JSON","Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","Codex CLI · effort low · JSON schema","\u0001",true,"higher"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","GPT-6.1 Sol (low, instructions) · Codex CLI","structured-output-time","Time per call, instructions vs schema mode",6.21,"seconds","6.21 s",12,"\u0001","minmax","Codex CLI · effort low · instructions",[4.2,12.27],"\u0001","\u0001"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","GPT-6.1 Sol (low, JSON schema) · Codex CLI","structured-output-time","Time per call, instructions vs schema mode",5.96,"seconds","5.96 s",12,"\u0001","minmax","Codex CLI · effort low · JSON schema",[4.62,20.97],"\u0001","\u0001"],["harder-tasks-head-to-head","harder-h2h-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","harder-h2h-pass-rate/Strict pass","Pass rate on 4 harder tasks (Strict pass)",0.6875,"rate","69% (11/16)",16,[0.444,0.8584],"ci95","Codex CLI · effort medium","\u0001","\u0001","higher"],["harder-tasks-head-to-head","harder-h2h-pass-rate","Lenient (format misses counted)","GPT-6.1 Sol (medium) · Codex CLI","harder-h2h-pass-rate/Lenient (format misses counted)","Pass rate on 4 harder tasks (Lenient (format misses counted))",0.6875,"rate","69% (11/16)",16,[0.444,0.8584],"ci95","Codex CLI · effort medium","\u0001","\u0001","higher"],["harder-tasks-head-to-head","harder-h2h-pass-by-task","GPT-6.1 Sol (medium) · Codex CLI","6x6 Skyscrapers","harder-h2h-pass-by-task/6x6 Skyscrapers","Strict pass rate by task: 6x6 Skyscrapers",1,"rate","100% (4/4)",4,[0.5101,1],"ci95","Codex CLI · effort medium","\u0001","\u0001","higher"]]}],["claude-code-cli","Claude Code","Anthropic","cli","Anthropic’s coding CLI. Each measurement pairs it with one Claude model; the context names the model.",["Claude Code","Claude Code CLI"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","range","spanKind","context","ci","polarity"],"$r":[["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","Claude Sonnet 5.5 · Claude Code","coding-agents-wall-time","Time per coding session",23.1,"seconds","23.1 s",12,[18.7,44.5],"minmax","Claude Sonnet 5.5 · six small repository tasks with hidden tests","\u0001","\u0001"],["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","Claude Opus 5.5 · Claude Code","coding-agents-wall-time","Time per coding session",56.9,"seconds","56.9 s",12,[29.8,185.8],"minmax","Claude Opus 5.5 · six small repository tasks with hidden tests","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",5.06,"seconds","5.06 s",10,[4.42,6.2],"minmax","Claude Haiku 4.5 · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",6.89,"seconds","6.89 s",10,[5.81,7.81],"minmax","Claude Sonnet 5.5 · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Code fix","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",5.95,"seconds","5.95 s",10,[4.89,7.33],"minmax","Claude Haiku 4.5 · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Code fix","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",2.67,"seconds","2.67 s",10,[2.32,4.34],"minmax","Claude Sonnet 5.5 · same prompt repeated 10 times","\u0001","\u0001"],["routing-overhead","cli-startup-tax","First model output","Claude Code · Claude Haiku 4.5","cli-startup-tax/First model output","CLI start-up tax on a one-word answer (First model output)",1461,"ms","1,461 ms",5,[1206,2308],"minmax","Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","\u0001","\u0001"],["routing-overhead","cli-startup-tax","Total wall time","Claude Code · Claude Haiku 4.5","cli-startup-tax/Total wall time","CLI start-up tax on a one-word answer (Total wall time)",2529,"ms","2,529 ms",5,[2273,3382],"minmax","Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","\u0001","\u0001"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","Claude Haiku 4.5 (single call) · Claude Code","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.4583,"rate","46% (11/24)",24,"\u0001","ci95","Claude Haiku 4.5 · single call",[0.2789,0.6493],"higher"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","Claude Haiku 4.5 (agent loop) · Claude Code","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.5417,"rate","54% (13/24)",24,"\u0001","ci95","Claude Haiku 4.5 · agent loop",[0.3507,0.7211],"higher"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","Claude Sonnet 5.5 (single call) · Claude Code","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",1,"rate","100% (24/24)",24,"\u0001","ci95","Claude Sonnet 5.5 · single call",[0.862,1],"higher"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","Claude Sonnet 5.5 (agent loop) · Claude Code","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",1,"rate","100% (16/16)",16,"\u0001","ci95","Claude Sonnet 5.5 · agent loop",[0.8064,1],"higher"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","Claude Haiku 4.5 (instructions) · Claude Code","structured-output-time","Time per call, instructions vs schema mode",9.52,"seconds","9.52 s",24,[5.67,17],"minmax","Claude Haiku 4.5 · instructions","\u0001","\u0001"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","Claude Haiku 4.5 (JSON schema) · Claude Code","structured-output-time","Time per call, instructions vs schema mode",8.46,"seconds","8.46 s",24,[5.9,12.23],"minmax","Claude Haiku 4.5 · JSON schema","\u0001","\u0001"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","Claude Sonnet 5.5 (instructions) · Claude Code","structured-output-time","Time per call, instructions vs schema mode",3.52,"seconds","3.52 s",12,[2.67,4.12],"minmax","Claude Sonnet 5.5 · instructions","\u0001","\u0001"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","Claude Sonnet 5.5 (JSON schema) · Claude Code","structured-output-time","Time per call, instructions vs schema mode",4.2,"seconds","4.20 s",12,[2.95,6.14],"minmax","Claude Sonnet 5.5 · JSON schema","\u0001","\u0001"],["harder-tasks-head-to-head","harder-h2h-pass-by-task","Claude Opus 5.5 · Claude Code","6x6 Skyscrapers","harder-h2h-pass-by-task/6x6 Skyscrapers","Strict pass rate by task: 6x6 Skyscrapers",0.3333,"rate","33% (1/3)",3,"\u0001","ci95","Claude Opus 5.5",[0.0615,0.7923],"higher"],["harder-tasks-head-to-head","harder-h2h-pass-by-task","Claude Sonnet 5.5 · Claude Code","6x6 Skyscrapers","harder-h2h-pass-by-task/6x6 Skyscrapers","Strict pass rate by task: 6x6 Skyscrapers",0,"rate","0% (0/4)",4,"\u0001","ci95","Claude Sonnet 5.5",[0,0.4899],"higher"],["harder-tasks-head-to-head","harder-h2h-pass-by-task","Claude Haiku 4.5 · Claude Code","6x6 Skyscrapers","harder-h2h-pass-by-task/6x6 Skyscrapers","Strict pass rate by task: 6x6 Skyscrapers",0,"rate","0% (0/3)",3,"\u0001","ci95","Claude Haiku 4.5",[0,0.5615],"higher"]]}],["codex-cli","Codex CLI","OpenAI","cli","OpenAI’s coding CLI. Each measurement pairs it with one GPT model and effort; the context names them.",["Codex CLI"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","range","spanKind","context","ci","polarity"],"$r":[["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI","coding-agents-wall-time","Time per coding session",113.4,"seconds","113.4 s",12,[78.5,221.9],"minmax","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",13.38,"seconds","13.4 s",10,[12.29,17.97],"minmax","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Code fix","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",11.29,"seconds","11.3 s",10,[9.08,14.85],"minmax","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001"],["routing-overhead","cli-startup-tax","First model output","Codex CLI (default model)","cli-startup-tax/First model output","CLI start-up tax on a one-word answer (First model output)",5059,"ms","5,059 ms",5,[4391,5478],"minmax","default model · CLI start-up, one-word prompt, 5 runs","\u0001","\u0001"],["routing-overhead","cli-startup-tax","Total wall time","Codex CLI (default model)","cli-startup-tax/Total wall time","CLI start-up tax on a one-word answer (Total wall time)",5999,"ms","5,999 ms",5,[5367,6506],"minmax","default model · CLI start-up, one-word prompt, 5 runs","\u0001","\u0001"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","GPT-6 Luna (single call) · Codex CLI","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.625,"rate","63% (10/16)",16,"\u0001","ci95","GPT-6 Luna · single call",[0.3864,0.8152],"higher"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","GPT-6 Luna (agent loop) · Codex CLI","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.8571,"rate","86% (12/14)",14,"\u0001","ci95","GPT-6 Luna · agent loop",[0.6006,0.9599],"higher"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","GPT-6.1 Sol (low, instructions) · Codex CLI","structured-output-time","Time per call, instructions vs schema mode",6.21,"seconds","6.21 s",12,[4.2,12.27],"minmax","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001"],["json-schema-vs-instructions","structured-output-time","Median time per call (the three prompts pooled)","GPT-6.1 Sol (low, JSON schema) · Codex CLI","structured-output-time","Time per call, instructions vs schema mode",5.96,"seconds","5.96 s",12,[4.62,20.97],"minmax","GPT-6.1 Sol · effort low · JSON schema","\u0001","\u0001"],["harder-tasks-head-to-head","harder-h2h-pass-by-task","GPT-6.1 Sol (medium) · Codex CLI","6x6 Skyscrapers","harder-h2h-pass-by-task/6x6 Skyscrapers","Strict pass rate by task: 6x6 Skyscrapers",1,"rate","100% (4/4)",4,"\u0001","ci95","GPT-6.1 Sol · effort medium",[0.5101,1],"higher"]]}],["jev-1-13","Jev 1.13","TypeSafe","router","A small routing model from TypeSafe that answers typed routing decisions over an API. Priced on input tokens only. Measured here as a router (accuracy, live time per call and cost per 1,000 decisions) and in the System One arena.",["Jev 1.13"],[{"studySlug":"routing-overhead","chartId":"router-overhead-decision-latency","series":"Decision time","point":"Jev 1.13 (TypeSafe)","metric":"router-overhead-decision-latency","label":"Time to make one routing decision","value":136.5,"unit":"ms","display":"137 ms","n":246,"range":[136.5,195.7],"spanKind":"p50-p95","context":"routing overhead per decision · TypeSafe API"},{"studySlug":"routing-holdout","chartId":"routing-holdout-latency","series":"Wall time","point":"Jev 1.13 (TypeSafe)","metric":"routing-holdout-latency/Wall time","label":"Time per routing decision, by route (Wall time)","value":0.139,"unit":"seconds","display":"0.14 s","n":168,"range":[0.139,0.192],"spanKind":"p50-p95","context":""}]],["agent-harness","Agent","Agent","harness","The Agent coding pipeline: onboarding, research, plan, act, verify and review, on Claude Sonnet 5.5 through a Claude subscription.",["Agent"],[]],["gpt-6-1-sol-openai-api","GPT-6.1 Sol (OpenAI API)","OpenAI","model","OpenAI’s GPT-6.1 Sol model called directly through the OpenAI API, without a CLI.",["GPT-6.1 Sol · OpenAI API"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","range","spanKind","context"],"$r":[["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","Total time","OpenAI API · GPT-6.1 Sol · low","cli-vs-api-exact-reply-latency/Total time","CLI vs API: time for a one-line answer (Total time)",1.02,"seconds","1.02 s",5,[0.96,1.87],"minmax","OpenAI API · effort low · fixed exact reply, 5 runs"],["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","Total time","OpenAI API · GPT-6.1 Sol · high","cli-vs-api-exact-reply-latency/Total time","CLI vs API: time for a one-line answer (Total time)",1.52,"seconds","1.52 s",5,[1.35,2.23],"minmax","OpenAI API · effort high · fixed exact reply, 5 runs"],["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","First useful output","OpenAI API · GPT-6.1 Sol · low","cli-vs-api-exact-reply-latency/First useful output","CLI vs API: time for a one-line answer (First useful output)",0.87,"seconds","0.87 s",5,[0.84,1.74],"minmax","OpenAI API · effort low · fixed exact reply, 5 runs"],["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","First useful output","OpenAI API · GPT-6.1 Sol · high","cli-vs-api-exact-reply-latency/First useful output","CLI vs API: time for a one-line answer (First useful output)",1.34,"seconds","1.34 s",5,[1.26,2.12],"minmax","OpenAI API · effort high · fixed exact reply, 5 runs"]]}],["gpt-6-luna-codex-cli","GPT-6 Luna (Codex CLI)","OpenAI","model","OpenAI’s GPT-6 Luna model run through the Codex CLI.",["GPT-6 Luna · Codex CLI"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","range","spanKind","context","ci","polarity"],"$r":[["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","Total time","Codex CLI · GPT-6 Luna · none","cli-vs-api-exact-reply-latency/Total time","CLI vs API: time for a one-line answer (Total time)",3.19,"seconds","3.19 s",5,[2.88,3.83],"minmax","Codex CLI · effort none · fixed exact reply, 5 runs","\u0001","\u0001"],["cli-model-latency-tokens","cli-vs-api-exact-reply-latency","First useful output","Codex CLI · GPT-6 Luna · none","cli-vs-api-exact-reply-latency/First useful output","CLI vs API: time for a one-line answer (First useful output)",2.79,"seconds","2.79 s",5,[2.46,3.42],"minmax","Codex CLI · effort none · fixed exact reply, 5 runs","\u0001","\u0001"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","GPT-6 Luna (single call) · Codex CLI","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.625,"rate","63% (10/16)",16,"\u0001","ci95","Codex CLI · single call",[0.3864,0.8152],"higher"],["single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","GPT-6 Luna (agent loop) · Codex CLI","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.8571,"rate","86% (12/14)",14,"\u0001","ci95","Codex CLI · agent loop",[0.6006,0.9599],"higher"],["single-call-vs-agent-loop","agent-loop-total-time","Total time per attempt","GPT-6 Luna (single call) · Codex CLI","agent-loop-total-time","Total time per attempt: single call vs agent loop",5.16,"seconds","5.16 s",16,[3.59,11.32],"minmax","Codex CLI · single call","\u0001","\u0001"],["single-call-vs-agent-loop","agent-loop-total-time","Total time per attempt","GPT-6 Luna (agent loop) · Codex CLI","agent-loop-total-time","Total time per attempt: single call vs agent loop",9.32,"seconds","9.32 s",14,[3.78,15.89],"minmax","Codex CLI · agent loop","\u0001","\u0001"]]}],["gpt-6-luna-openai-api","GPT-6 Luna (OpenAI API)","OpenAI","model","OpenAI’s GPT-6 Luna model called directly through the OpenAI API, without a CLI.",["GPT-6 Luna · OpenAI API"],[{"studySlug":"cli-model-latency-tokens","chartId":"cli-vs-api-exact-reply-latency","series":"Total time","point":"OpenAI API · GPT-6 Luna · none","metric":"cli-vs-api-exact-reply-latency/Total time","label":"CLI vs API: time for a one-line answer (Total time)","value":0.97,"unit":"seconds","display":"0.97 s","n":5,"range":[0.65,1.5],"spanKind":"minmax","context":"OpenAI API · effort none · fixed exact reply, 5 runs"},{"studySlug":"cli-model-latency-tokens","chartId":"cli-vs-api-exact-reply-latency","series":"First useful output","point":"OpenAI API · GPT-6 Luna · none","metric":"cli-vs-api-exact-reply-latency/First useful output","label":"CLI vs API: time for a one-line answer (First useful output)","value":0.82,"unit":"seconds","display":"0.82 s","n":5,"range":[0.51,1.37],"spanKind":"minmax","context":"OpenAI API · effort none · fixed exact reply, 5 runs"}]],["gpt-5-2","GPT 5.2","OpenAI","model","OpenAI model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).",["GPT 5.2"],[]],["gemini-3-flash","Gemini 3 Flash","Google","model","Google model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).",["Gemini 3 Flash"],[]],["glm-5","GLM 5","Z.ai","model","Z.ai model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).",["GLM 5"],[]],["claude-sonnet-4-5","Claude Sonnet 4.5","Anthropic","model","Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).",["Claude Sonnet 4.5","Claude 4.5 Sonnet"],[]],["claude-opus-4-5","Claude Opus 4.5","Anthropic","model","Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).",["Claude Opus 4.5","Claude 4.5 Opus"],[]],["claude-opus-4-6","Claude Opus 4.6","Anthropic","model","Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2).",["Claude Opus 4.6","Claude 4.6 Opus"],[]],["deepseek-v3-2","DeepSeek V3.2","DeepSeek","model","DeepSeek model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).",["DeepSeek V3.2"],[]],["minimax-m2-5","MiniMax M2.5","MiniMax","model","MiniMax model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).",["MiniMax M2.5"],[]],["kimi-k2-5","Kimi K2.5","Moonshot AI","model","Moonshot AI model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).",["Kimi K2.5"],[]],["gpt-5-mini","GPT 5 mini","OpenAI","model","OpenAI model in the public SWE-bench Verified panel (mini-SWE-agent v2).",["GPT 5 mini"],[]],["claude-opus-5","Claude Opus 5","Anthropic","model","Anthropic model, measured as a blind code-review critic and in a list-price calculation.",["Claude Opus 5"],[]],["claude-fable-5","Claude Fable 5","Anthropic","model","Anthropic model, measured as a blind code-review critic.",["Claude Fable 5"],[]],["gpt-5-5","GPT 5.5","OpenAI","model","OpenAI model, measured as a blind code-review critic on a small number of pairs.",["GPT 5.5"],[]],["deterministic-routing-policy","Deterministic routing policy","Agent","router","Agent’s rule-based routing: an in-process policy picks the model and effort for each call from the task stage and signals. No model call, so no token cost.",["Deterministic routing policy"],[{"studySlug":"routing-overhead","chartId":"router-overhead-decision-latency","series":"Decision time","point":"Deterministic routing policy (Agent, in process)","metric":"router-overhead-decision-latency","label":"Time to make one routing decision","value":0.00142,"unit":"ms","display":"1.42 µs","n":20000,"range":[0.00142,0.00233],"spanKind":"p50-p95","context":"Agent · in process · routing overhead per decision"}]],["openrouter","OpenRouter","OpenRouter","provider","A gateway that routes one API to many inference providers. Its per-token price is compared with first-party list prices; its fee is charged when credits are bought.",["OpenRouter"],[]],["anthropic","Anthropic","Anthropic","provider","Anthropic’s own API: the first-party list price of the Claude models, and Anthropic’s endpoint as listed on OpenRouter.",["Anthropic"],[]],["openai","OpenAI","OpenAI","provider","OpenAI’s own API: the first-party list price of GPT models, and OpenAI’s endpoints as listed on OpenRouter.",["OpenAI"],[]],["google-ai-studio","Google AI Studio","Google","provider","Google’s Gemini API (AI Studio): the first-party list price of Gemini models, and its endpoints as listed on OpenRouter.",["Google AI Studio"],[]],["google-vertex","Google Vertex AI","Google","provider","Google Cloud’s model platform. An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Google Vertex"],[]],["amazon-bedrock","Amazon Bedrock","Amazon","provider","Amazon Web Services’ model platform. An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Amazon Bedrock"],[]],["azure","Azure","Microsoft","provider","Microsoft’s cloud model platform. An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Azure"],[]],["claude-platform-on-aws","Claude Platform on AWS","Anthropic","provider","Anthropic’s Claude platform hosted on AWS. An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Claude Platform on AWS"],[]],["groq","Groq","Groq","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Groq"],[]],["together","Together AI","Together AI","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Together"],[]],["fireworks","Fireworks AI","Fireworks AI","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Fireworks"],[]],["deepinfra","DeepInfra","DeepInfra","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["DeepInfra"],[]],["cerebras","Cerebras","Cerebras","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Cerebras"],[]],["sambanova","SambaNova","SambaNova","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["SambaNova"],[]],["nebius","Nebius","Nebius","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Nebius"],[]],["parasail","Parasail","Parasail","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Parasail"],[]],["novita","Novita AI","Novita AI","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Novita"],[]],["baseten","Baseten","Baseten","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["BaseTen"],[]],["cloudflare","Cloudflare Workers AI","Cloudflare","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["Cloudflare"],[]],["siliconflow","SiliconFlow","SiliconFlow","provider","An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.",["SiliconFlow"],[]]]},"comparisons":{"$k":["slug","a","b","title","seoTitle","description","verdict","rows"],"$r":[["claude-sonnet-5-5-vs-claude-opus-5-5","claude-sonnet-5-5","claude-opus-5-5","Claude Sonnet 5.5 vs Claude Opus 5.5","Claude Sonnet 5.5 vs Claude Opus 5.5: measured benchmarks","Claude Sonnet 5.5 vs Claude Opus 5.5: 35 measured metrics from 8 studies, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 and Claude Opus 5.5 share 35 measured metrics and 31 list-price calculations from 10 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 16 ties and 50 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved on the same 3 SWE-bench Verified instances (interim)",0.3333,0.6667,"rate","33% (1/3)","67% (2/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 6% to 79%; Claude Opus 5.5 21% to 94%), so this sample cannot separate them.","swe-bench-opus-vs-sonnet",3,"swebench-opus-sonnet-resolved",3,3,"Agent · older builds · SWE-bench Verified, interim paired probe","Agent · new build · SWE-bench Verified, interim paired probe","ci95","ci95",[0.0615,0.7923],[0.2077,0.9385],"\u0001"],["List-price cost per attempt (calculation)",2.88,7.59,"usd","$2.88","$7.59","unclear","No interval or range was recorded for either side, so the gap ($2.88 vs $7.59, 2.6x) is not tested against run-to-run variation.","swe-bench-opus-vs-sonnet",3,"swebench-opus-sonnet-cost-per-attempt",3,3,"Agent · older builds · SWE-bench Verified, interim paired probe","Agent · new build · SWE-bench Verified, interim paired probe","\u0001","\u0001","\u0001","\u0001",true],["Worker time per attempt",9.37,20.26,"minutes","9.4 min","20.3 min","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.7 min to 15.0 min; Claude Opus 5.5 9.8 min to 25.3 min); the medians alone do not show a reliable difference. A range is not a confidence interval.","swe-bench-opus-vs-sonnet",3,"swebench-opus-sonnet-minutes",3,3,"Agent · older builds · SWE-bench Verified, interim paired probe","Agent · new build · SWE-bench Verified, interim paired probe","range","minmax",[4.74,15],[9.76,25.29],"\u0001"],["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,2.75,"seconds","2.31 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.17,7.73],[2.47,8.91],"\u0001"],["Time to first useful output",1.56,1.92,"seconds","1.56 s","1.92 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; Claude Opus 5.5 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.99,6.39],[1.56,7.23],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1401,1401,"tokens","1,401","1,401","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",685,680,"tokens","685","680","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",107,64,"tokens","107","64","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.0036,0.00688,"usd","$0.0036","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 $0.0034 to $0.010; Claude Opus 5.5 $0.0059 to $0.022); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00342,0.01021],[0.00592,0.02226],true],["List-price cost per passing answer (calculation)",0.00624,0.01009,"usd","$0.0062","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Opus 5.5 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Opus 5.5 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,9.18,"seconds","7.75 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[2.26,34.79],[4.24,27.21],"\u0001"],["Time to first useful output on hard tasks",5.95,6.78,"seconds","5.95 s","6.78 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.86 s to 30.6 s; Claude Opus 5.5 2.39 s to 21.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[0.86,30.57],[2.39,21.77],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1050,945,"tokens","1,050","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.02824,"usd","$0.014","$0.028","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.028) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; Claude Opus 5.5 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",23.1,56.9,"seconds","23.1 s","56.9 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; Claude Opus 5.5 29.8 s to 185.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","range","minmax",[18.7,44.5],[29.8,185.8],"\u0001"],["Tool calls per coding session",7.5,7.5,"calls","7.5","7.5","tie","Same value. More or fewer is not better by itself for this metric.","coding-agents-head-to-head",12,"coding-agents-tool-calls",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","range","minmax",[3,14],[5,14],"\u0001"],["List-price cost per passing coding session (calculation)",0.085,0.2229,"usd","$0.085","$0.22","unclear","No interval or range was recorded for either side, so the gap ($0.085 vs $0.22, 2.6x) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 81% to 100%; Claude Opus 5.5 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.97,9.18,"seconds","7.97 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 21.6 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[2.26,21.61],[4.24,27.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",1054,945,"tokens","1,054","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01398,0.02893,"usd","$0.014","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.029, 2.1x) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded)",0.135003,0.255057,"usd","$0.14","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.14 vs $0.26) is not tested against run-to-run variation.","caching-consistency",15,"caching-cost-with-without",15,15,"Claude Code · calculation: 5-turn cached sessions over a fixed ledger","Claude Code · calculation: 5-turn cached sessions over a fixed ledger","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price)",0.269788,0.544228,"usd","$0.27","$0.54","unclear","No interval or range was recorded for either side, so the gap ($0.27 vs $0.54, 2.0x) is not tested against run-to-run variation.","caching-consistency",15,"caching-cost-with-without",15,15,"Claude Code · calculation: 5-turn cached sessions over a fixed ledger","Claude Code · calculation: 5-turn cached sessions over a fixed ledger","\u0001","\u0001","\u0001","\u0001",true],["Time per turn: first turn vs later turns in a cached session (Turn 1 (writes the ledger to the cache))",1.64,1.9,"seconds","1.64 s","1.90 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.58 s to 1.79 s; Claude Opus 5.5 1.78 s to 4.36 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",3,"caching-latency-first-vs-later",3,3,"Claude Code · 5-turn cached sessions over a fixed ledger","Claude Code · 5-turn cached sessions over a fixed ledger","range","minmax",[1.58,1.79],[1.78,4.36],"\u0001"],["Time per turn: first turn vs later turns in a cached session (Turns 2-5 (read the ledger from the cache))",1.61,2.4,"seconds","1.61 s","2.40 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.35 s to 5.63 s; Claude Opus 5.5 1.63 s to 12.7 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",12,"caching-latency-first-vs-later",12,12,"Claude Code · 5-turn cached sessions over a fixed ledger","Claude Code · 5-turn cached sessions over a fixed ledger","range","minmax",[1.35,5.63],[1.63,12.67],"\u0001"],["Cost of a reused prefix with and without the cache, by session length (calculation): 1 turn",15.66,31.31,"usd","$15.66","$31.31","unclear","No interval or range was recorded for either side, so the gap ($15.66 vs $31.31) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 2 turns",31.32,62.62,"usd","$31.32","$62.62","unclear","No interval or range was recorded for either side, so the gap ($31.32 vs $62.62) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 3 turns",46.99,93.94,"usd","$46.99","$93.94","unclear","No interval or range was recorded for either side, so the gap ($46.99 vs $93.94) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 5 turns",78.31,156.56,"usd","$78.31","$156.56","unclear","No interval or range was recorded for either side, so the gap ($78.31 vs $156.56) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 10 turns",156.62,313.12,"usd","$156.62","$313.12","unclear","No interval or range was recorded for either side, so the gap ($156.62 vs $313.12) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["Cost of a reused prefix with and without the cache, by session length (calculation): 20 turns",313.24,626.24,"usd","$313.24","$626.24","unclear","No interval or range was recorded for either side, so the gap ($313.24 vs $626.24) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-cost-curve","\u0001","\u0001","no cache","no cache","\u0001","\u0001","\u0001","\u0001",true],["One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (No cache (the same either way))",156.62,313.12,"usd","$156.62","$313.12","unclear","No interval or range was recorded for either side, so the gap ($156.62 vs $313.12) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-session-split","\u0001","\u0001","","cache read $0.2 per M","\u0001","\u0001","\u0001","\u0001",true],["One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (One 10-turn session, 1-hour cache)",45.42,76.71,"usd","$45.42","$76.71","unclear","No interval or range was recorded for either side, so the gap ($45.42 vs $76.71) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-session-split","\u0001","\u0001","","cache read $0.2 per M","\u0001","\u0001","\u0001","\u0001",true],["One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix))",313.24,626.24,"usd","$313.24","$626.24","unclear","No interval or range was recorded for either side, so the gap ($313.24 vs $626.24) is not tested against run-to-run variation.","prompt-cache-break-even","\u0001","cache-break-even-session-split","\u0001","\u0001","","cache read $0.2 per M","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",54.54,54.79,"percent","54.5%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[0,95.91],[29.92,95.6],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.006665,0.012528,"usd","$0.0067","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.003672,0.008003,"usd","$0.0037","$0.0080","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.004012,0.00771,"usd","$0.0040","$0.0077","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.006299,0.013104,"usd","$0.0063","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.013978,0.028925,"usd","$0.014","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.54,54.79,"percent","54.5%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[0,95.91],[29.92,95.6],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",0,0,"percent","0%","0%","tie","Same value. More or fewer is not better by itself for this metric.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[0,72.75],[0,93.33],true],["Time to first text: a 250-line answer, six models",1.96,1.97,"seconds","1.96 s","1.97 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; Claude Opus 5.5 1.70 s to 2.35 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[0.88,4.09],[1.7,2.35],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",231.7,155.5,"tokens","232","156","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[230.3,233],[154.6,156.4],true],["Output speed in characters per second after the first text (calculation)",517,347,"count","517","347","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Code","Claude Code","range","minmax",[513,519],[345,349],true],["Time to first text as the prompt grows: 1k",1.45,1.51,"seconds","1.45 s","1.51 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.23 s to 1.72 s; Claude Opus 5.5 1.46 s to 2.01 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.23,1.72],[1.46,2.01],true],["Time to first text as the prompt grows: 16k",1.78,1.74,"seconds","1.78 s","1.74 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.64 s to 2.11 s; Claude Opus 5.5 1.70 s to 2.97 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.64,2.11],[1.7,2.97],true],["Time to first text as the prompt grows: 64k",3.07,1.79,"seconds","3.07 s","1.79 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.38 s to 3.61 s; Claude Opus 5.5 1.72 s to 3.72 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.38,3.61],[1.72,3.72],true],["Total time per call by prompt size (1k prompt)",1.78,1.83,"seconds","1.78 s","1.83 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.57 s to 2.12 s; Claude Opus 5.5 1.82 s to 2.41 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[1.57,2.12],[1.82,2.41],"\u0001"],["Total time per call by prompt size (16k prompt)",2.1,2.36,"seconds","2.10 s","2.36 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.98 s to 2.48 s; Claude Opus 5.5 2.11 s to 3.40 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[1.98,2.48],[2.11,3.4],"\u0001"],["Total time per call by prompt size (64k prompt)",3.44,2.35,"seconds","3.44 s","2.35 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.74 s to 4.38 s; Claude Opus 5.5 2.26 s to 4.29 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[1.74,4.38],[2.26,4.29],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,0.5556,"rate","100% (9/9)","56% (5/9)","tie","The 95% intervals overlap (Claude Sonnet 5.5 70% to 100%; Claude Opus 5.5 27% to 81%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Claude Code","ci95","ci95",[0.7009,1],[0.2667,0.8112],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0.375,0.4167,"rate","38% (6/16)","42% (5/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; Claude Opus 5.5 19% to 68%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",16,12,"Claude Code","Claude Code","ci95","ci95",[0.1848,0.6136],[0.1933,0.6805],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0.375,0.5,"rate","38% (6/16)","50% (6/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; Claude Opus 5.5 25% to 75%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",16,12,"Claude Code","Claude Code","ci95","ci95",[0.1848,0.6136],[0.2538,0.7462],"\u0001"],["Calls that tried a tool although tools were off",0.3125,0.4167,"rate","31% (5/16)","42% (5/12)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-tool-attempts",16,12,"Claude Code","Claude Code","ci95","ci95",[0.1416,0.556],[0.1933,0.6805],"\u0001"],["Strict pass rate by task: 10x10 nonogram",1,1,"rate","100% (4/4)","100% (3/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 51% to 100%; Claude Opus 5.5 44% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",4,3,"Claude Code","Claude Code","ci95","ci95",[0.5101,1],[0.4385,1],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0,"rate","0% (0/4)","0% (0/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 0% to 49%; Claude Opus 5.5 0% to 56%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",4,3,"Claude Code","Claude Code","ci95","ci95",[0,0.4899],[0,0.5615],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,0.3333,"rate","0% (0/4)","33% (1/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 0% to 49%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",4,3,"Claude Code","Claude Code","ci95","ci95",[0,0.4899],[0.0615,0.7923],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0.5,0.3333,"rate","50% (2/4)","33% (1/3)","tie","The 95% intervals overlap (Claude Sonnet 5.5 15% to 85%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",4,3,"Claude Code","Claude Code","ci95","ci95",[0.15,0.85],[0.0615,0.7923],"\u0001"],["Total time per call on harder tasks",70.43,80.34,"seconds","70.4 s","80.3 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.32 s to 210.1 s; Claude Opus 5.5 3.82 s to 279.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",12,9,"Claude Code","Claude Code","range","minmax",[4.32,210.08],[3.82,279.5],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",9287,8420,"tokens","9,287","8,420","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",12,9,"Claude Code","Claude Code","range","minmax",[407,27921],[279,40044],"\u0001"],["List-price cost per strict pass on harder tasks (calculation)",0.23843,0.59333,"usd","$0.24","$0.59","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $0.59, 2.5x) is not tested against run-to-run variation.","harder-tasks-head-to-head","\u0001","harder-h2h-cost-per-pass",16,12,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-haiku-4-5-vs-claude-sonnet-5-5","claude-haiku-4-5","claude-sonnet-5-5","Claude Haiku 4.5 vs Claude Sonnet 5.5","Claude Haiku 4.5 vs Claude Sonnet 5.5: measured benchmarks","Claude Haiku 4.5 vs Claude Sonnet 5.5: 113 measured metrics from 12 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Haiku 4.5 and Claude Sonnet 5.5 share 113 measured metrics and 40 list-price calculations from 14 studies. Claude Sonnet 5.5 leads on 21 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); and 18 more. On those rows the 95% intervals, run ranges and p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 58 ties and 74 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 2 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,0.8,"rate","100% (15/15)","80% (12/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.5481,0.9295],"\u0001"],["Total time per call",4.43,2.31,"seconds","4.43 s","2.31 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Sonnet 5.5 2.17 s to 7.73 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[3.16,23.57],[2.17,7.73],"\u0001"],["Time to first useful output",3.63,1.56,"seconds","3.63 s","1.56 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; Claude Sonnet 5.5 0.99 s to 6.39 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.78,22.27],[0.99,6.39],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",0,1401,"tokens","0","1,401","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",3790,685,"tokens","3,790","685","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",367,107,"tokens","367","107","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00566,0.0036,"usd","$0.0057","$0.0036","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; Claude Sonnet 5.5 $0.0034 to $0.010); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00513,0.01804],[0.00342,0.01021],true],["List-price cost per passing answer (calculation)",0.00836,0.00624,"usd","$0.0084","$0.0062","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.0062) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Sonnet 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,7.75,"seconds","39.0 s","7.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Sonnet 5.5 2.26 s to 34.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[15.27,75.13],[2.26,34.79],"\u0001"],["Time to first useful output on hard tasks",35.54,5.95,"seconds","35.5 s","5.95 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; Claude Sonnet 5.5 0.86 s to 30.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[12.88,70.31],[0.86,30.57],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",5064,1050,"tokens","5,064","1,050","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.01435,"usd","$0.067","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.014, 4.7x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",0,1,"rate","0% (0/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","ci95","ci95",[0,0.2775],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",0.1,1,"rate","10% (1/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","ci95","ci95",[0.0179,0.4042],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: how many different answers (Exact number)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (JSON object)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (Code fix)",6,3,"count","6","3","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: time per call (Exact number)",5.06,6.89,"seconds","5.06 s","6.89 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 4.42 s to 6.20 s; Claude Sonnet 5.5 5.81 s to 7.81 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","range","minmax",[4.42,6.2],[5.81,7.81],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",7.03,2.89,"seconds","7.03 s","2.89 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.28 s to 12.3 s; Claude Sonnet 5.5 2.68 s to 5.30 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","range","minmax",[5.28,12.27],[2.68,5.3],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",5.95,2.67,"seconds","5.95 s","2.67 s","b","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; Claude Sonnet 5.5 2.32 s to 4.34 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","range","minmax",[4.89,7.33],[2.32,4.34],"\u0001"],["Full pass rate by kind of memory: No memory",0.2,0.6,"rate","20% (2/10)","60% (9/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 6% to 51%; Claude Sonnet 5.5 36% to 80%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.0567,0.5098],[0.3575,0.8018],"\u0001"],["Full pass rate by kind of memory: /init CLAUDE.md",0.2,0.6,"rate","20% (2/10)","60% (9/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 6% to 51%; Claude Sonnet 5.5 36% to 80%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.0567,0.5098],[0.3575,0.8018],"\u0001"],["Full pass rate by kind of memory: Curated, 11 lines",0.7,1,"rate","70% (7/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 40% to 89%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.3968,0.8922],[0.7961,1],"\u0001"],["Full pass rate by kind of memory: Raw notes, 60 lines",0.6,0.9333,"rate","60% (6/10)","93% (14/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 31% to 83%; Claude Sonnet 5.5 70% to 99%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.3127,0.8318],[0.7018,0.9881],"\u0001"],["Full pass rate by kind of memory: Dreamed notes",0.7,1,"rate","70% (7/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 40% to 89%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.3968,0.8922],[0.7961,1],"\u0001"],["Full pass rate by kind of memory: Handbook, 210 lines",0.3,1,"rate","30% (3/10)","100% (15/15)","b","The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%).","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.1078,0.6032],[0.7961,1],"\u0001"],["Full pass rate by kind of memory: Stop hook only",0.8,0.8,"rate","80% (8/10)","80% (12/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.4902,0.9433],[0.5481,0.9295],"\u0001"],["Full pass rate by kind of memory: Curated + hook",0.9,1,"rate","90% (9/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 60% to 98%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-full-pass",10,15,"","","ci95","ci95",[0.5958,0.9821],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: No memory",0,0.4,"rate","0% (0/10)","40% (6/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 20% to 64%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0,0.2775],[0.1982,0.6425],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md",0.1,0.6667,"rate","10% (1/10)","67% (10/15)","b","The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 42% to 85%).","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.0179,0.4042],[0.4171,0.8482],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Curated, 11 lines",0.8,1,"rate","80% (8/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.4902,0.9433],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Raw notes, 60 lines",0.6,1,"rate","60% (6/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 31% to 83%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.3127,0.8318],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Dreamed notes",0.8,1,"rate","80% (8/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.4902,0.9433],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines",0.3,1,"rate","30% (3/10)","100% (15/15)","b","The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%).","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.1078,0.6032],[0.7961,1],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Stop hook only",0.8,0.6667,"rate","80% (8/10)","67% (10/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 42% to 85%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.4902,0.9433],[0.4171,0.8482],"\u0001"],["Team knowledge followed, Sonnet vs Haiku: Curated + hook",1,1,"rate","100% (10/10)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them.","agent-memory","\u0001","memory-team-knowledge-by-model",10,15,"","","ci95","ci95",[0.7225,1],[0.7961,1],"\u0001"],["A stale README command: who still ran it?: No memory",1,0.8,"rate","100% (10/10)","80% (12/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.7225,1],[0.5481,0.9295],"\u0001"],["A stale README command: who still ran it?: /init CLAUDE.md",1,0.8667,"rate","100% (10/10)","87% (13/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 62% to 96%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.7225,1],[0.6212,0.9626],"\u0001"],["A stale README command: who still ran it?: Curated, 11 lines",0,0,"rate","0% (0/10)","0% (0/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0,0.2775],[0,0.2039],"\u0001"],["A stale README command: who still ran it?: Raw notes, 60 lines",1,0,"rate","100% (10/10)","0% (0/15)","b","The 95% intervals do not overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 0% to 20%).","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.7225,1],[0,0.2039],"\u0001"],["A stale README command: who still ran it?: Dreamed notes",0.1,0,"rate","10% (1/10)","0% (0/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.0179,0.4042],[0,0.2039],"\u0001"],["A stale README command: who still ran it?: Handbook, 210 lines",0,0,"rate","0% (0/10)","0% (0/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0,0.2775],[0,0.2039],"\u0001"],["A stale README command: who still ran it?: Stop hook only",1,0.6,"rate","100% (10/10)","60% (9/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 36% to 80%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0.7225,1],[0.3575,0.8018],"\u0001"],["A stale README command: who still ran it?: Curated + hook",0,0,"rate","0% (0/10)","0% (0/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them.","agent-memory","\u0001","memory-broken-test-command",10,15,"","","ci95","ci95",[0,0.2775],[0,0.2039],"\u0001"],["List-price cost per fully correct result (calculation): No memory",0.3786,0.1386,"usd","$0.38","$0.14","unclear","No interval or range was recorded for either side, so the gap ($0.38 vs $0.14, 2.7x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",2,9,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): /init CLAUDE.md",0.4317,0.1359,"usd","$0.43","$0.14","unclear","No interval or range was recorded for either side, so the gap ($0.43 vs $0.14, 3.2x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",2,9,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Curated, 11 lines",0.1095,0.0818,"usd","$0.11","$0.082","unclear","No interval or range was recorded for either side, so the gap ($0.11 vs $0.082) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",7,15,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Raw notes, 60 lines",0.1255,0.1009,"usd","$0.13","$0.10","unclear","No interval or range was recorded for either side, so the gap ($0.13 vs $0.10) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",6,14,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Dreamed notes",0.1153,0.0896,"usd","$0.12","$0.090","unclear","No interval or range was recorded for either side, so the gap ($0.12 vs $0.090) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",7,15,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Handbook, 210 lines",0.2615,0.1006,"usd","$0.26","$0.10","unclear","No interval or range was recorded for either side, so the gap ($0.26 vs $0.10, 2.6x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",3,15,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Stop hook only",0.1419,0.1278,"usd","$0.14","$0.13","unclear","No interval or range was recorded for either side, so the gap ($0.14 vs $0.13) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",8,12,"","","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per fully correct result (calculation): Curated + hook",0.096,0.0843,"usd","$0.096","$0.084","unclear","No interval or range was recorded for either side, so the gap ($0.096 vs $0.084) is not tested against run-to-run variation.","agent-memory","\u0001","memory-cost-per-full-pass",9,15,"","","\u0001","\u0001","\u0001","\u0001",true],["Time per session: No memory",54.2,18,"seconds","54.2 s","18.0 s","unclear","No interval or range was recorded for either side, so the gap (54.2 s vs 18.0 s, 3.0x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: /init CLAUDE.md",52.9,19,"seconds","52.9 s","19.0 s","unclear","No interval or range was recorded for either side, so the gap (52.9 s vs 19.0 s, 2.8x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Curated, 11 lines",51.7,21.9,"seconds","51.7 s","21.9 s","unclear","No interval or range was recorded for either side, so the gap (51.7 s vs 21.9 s, 2.4x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Raw notes, 60 lines",51,26.8,"seconds","51.0 s","26.8 s","unclear","No interval or range was recorded for either side, so the gap (51.0 s vs 26.8 s) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Dreamed notes",51.8,27.2,"seconds","51.8 s","27.2 s","unclear","No interval or range was recorded for either side, so the gap (51.8 s vs 27.2 s) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Handbook, 210 lines",49.9,23.6,"seconds","49.9 s","23.6 s","unclear","No interval or range was recorded for either side, so the gap (49.9 s vs 23.6 s, 2.1x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Stop hook only",68.5,27.3,"seconds","68.5 s","27.3 s","unclear","No interval or range was recorded for either side, so the gap (68.5 s vs 27.3 s, 2.5x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per session: Curated + hook",52.8,22,"seconds","52.8 s","22.0 s","unclear","No interval or range was recorded for either side, so the gap (52.8 s vs 22.0 s, 2.4x) is not tested against run-to-run variation.","agent-memory","\u0001","memory-wall-time",10,15,"","","\u0001","\u0001","\u0001","\u0001","\u0001"],["Typed routing decisions answered exactly right",0.8902,0.939,"rate","89% (73/82)","94% (77/82)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 94%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","routing-jev-vs-llm",82,"routing-exact-decisions",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.8044,0.9412],[0.8651,0.9737],"\u0001"],["Per-question accuracy",0.9433,0.9742,"rate","94% (183/194)","97% (189/194)","tie","The 95% intervals overlap (Claude Haiku 4.5 90% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them.","routing-jev-vs-llm",194,"routing-key-accuracy",194,194,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.9013,0.968],[0.9411,0.9889],"\u0001"],["Exact rate by decision type: Failure class",0.9444,1,"rate","94% (17/18)","100% (18/18)","tie","The 95% intervals overlap (Claude Haiku 4.5 74% to 99%; Claude Sonnet 5.5 82% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",18,"routing-exact-by-decision",18,18,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.7424,0.9901],[0.8241,1],"\u0001"],["Exact rate by decision type: Message intent",1,1,"rate","100% (20/20)","100% (20/20)","tie","The 95% intervals overlap (Claude Haiku 4.5 84% to 100%; Claude Sonnet 5.5 84% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",20,"routing-exact-by-decision",20,20,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.8389,1],[0.8389,1],"\u0001"],["Exact rate by decision type: Is it a rule?",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Haiku 4.5 76% to 100%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",12,"routing-exact-by-decision",12,12,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Exact rate by decision type: Context shape",0.75,0.8438,"rate","75% (24/32)","84% (27/32)","tie","The 95% intervals overlap (Claude Haiku 4.5 58% to 87%; Claude Sonnet 5.5 68% to 93%), so this sample cannot separate them.","routing-jev-vs-llm",32,"routing-exact-by-decision",32,32,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.5789,0.8675],[0.6825,0.9314],"\u0001"],["Cost per 1,000 routing decisions",8.924,4.996,"usd","$8.92","$5.00","unclear","No interval or range was recorded for either side, so the gap ($8.92 vs $5.00) is not tested against run-to-run variation.","routing-jev-vs-llm",82,"routing-cost-per-1000",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Time per routing decision (Wall time (CLI))",12674,2598,"ms","12,674 ms","2,598 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,674 ms to 34,413 ms; Claude Sonnet 5.5 2,598 ms to 4,298 ms); not a confidence interval.","routing-jev-vs-llm",82,"routing-decision-latency",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","range","p50-p95",[12674,34413],[2598,4298],"\u0001"],["Time per routing decision (Model time (API))",10734,1599,"ms","10,734 ms","1,599 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,734 ms to 32,072 ms; Claude Sonnet 5.5 1,599 ms to 2,574 ms); not a confidence interval.","routing-jev-vs-llm",82,"routing-decision-latency",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","range","p50-p95",[10734,32072],[1599,2574],"\u0001"],["Time to make one routing decision",12543,2597,"ms","12,543 ms","2,597 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,543 ms to 34,481 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval.","routing-overhead",82,"router-overhead-decision-latency",82,82,"thinking on · via Claude Code · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","range","p50-p95",[12543,34481],[2597,4298],"\u0001"],["Where an LLM router’s time goes: model vs CLI (Model API time)",10508,1596,"ms","10,508 ms","1,596 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,508 ms to 32,132 ms; Claude Sonnet 5.5 1,596 ms to 2,583 ms); not a confidence interval.","routing-overhead",82,"router-overhead-cli-vs-model-time",82,82,"thinking on · via Claude Code · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","range","p50-p95",[10508,32132],[1596,2583],"\u0001"],["Where an LLM router’s time goes: model vs CLI (CLI and harness time)",1698,973,"ms","1,698 ms","973 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 1,698 ms to 2,677 ms; Claude Sonnet 5.5 973 ms to 1,277 ms); not a confidence interval.","routing-overhead",82,"router-overhead-cli-vs-model-time",82,82,"thinking on · via Claude Code · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","range","p50-p95",[1698,2677],[973,1277],"\u0001"],["Routing calls that returned a decision",1,1,"rate","100% (82/82)","100% (82/82)","tie","The 95% intervals overlap (Claude Haiku 4.5 96% to 100%; Claude Sonnet 5.5 96% to 100%), so this sample cannot separate them.","routing-overhead",82,"router-overhead-completed",82,82,"thinking on · via Claude Code · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","ci95","ci95",[0.9552,1],[0.9552,1],"\u0001"],["Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))",441.74,247.3,"usd","$441.74","$247.30","unclear","No interval or range was recorded for either side, so the gap ($441.74 vs $247.30) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))",62.47,34.97,"usd","$62.47","$34.97","unclear","No interval or range was recorded for either side, so the gap ($62.47 vs $34.97) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Every model call routed (49.5 per task))",620.8785,128.5515,"seconds","620.9 s","128.6 s","unclear","No interval or range was recorded for either side, so the gap (620.9 s vs 128.6 s, 4.8x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Only System One decisions (7 per task))",87.801,18.179,"seconds","87.8 s","18.2 s","unclear","No interval or range was recorded for either side, so the gap (87.8 s vs 18.2 s, 4.8x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate: single call vs agent loop on eight hard tasks",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).","single-call-vs-agent-loop",24,"agent-loop-pass-rate",24,24,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Strict passes per task: single call vs agent loop: Interval merge fix",1,1,"rate","100% (3/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.4385,1],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: DST day-length fix",0.3333,1,"rate","33% (1/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 6% to 79%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.0615,0.7923],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: CSV parser",0.6667,1,"rate","67% (2/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.2077,0.9385],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: Event-loop order",0,1,"rate","0% (0/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0,0.5615],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: Room schedule",0,1,"rate","0% (0/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0,0.5615],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: SemVer regex",1,1,"rate","100% (3/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.4385,1],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: Money refactor",0.6667,1,"rate","67% (2/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0.2077,0.9385],[0.4385,1],"\u0001"],["Strict passes per task: single call vs agent loop: SQL report",0,1,"rate","0% (0/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop",3,"agent-loop-by-task",3,3,"Claude Code · single call","Claude Code · single call","ci95","ci95",[0,0.5615],[0.4385,1],"\u0001"],["Total time per attempt: single call vs agent loop",39.01,7.75,"seconds","39.0 s","7.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Sonnet 5.5 2.26 s to 34.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","single-call-vs-agent-loop",24,"agent-loop-total-time",24,24,"Claude Code · single call","Claude Code · single call","range","minmax",[15.27,75.13],[2.26,34.79],"\u0001"],["Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))",3941,2281,"tokens","3,941","2,281","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop",24,"agent-loop-tokens",24,24,"Claude Code · single call","Claude Code · single call","range","minmax",[3879,4221],[2234,2669],"\u0001"],["Tokens per attempt: single call vs agent loop (Output tokens)",5064,1050,"tokens","5,064","1,050","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop",24,"agent-loop-tokens",24,24,"Claude Code · single call","Claude Code · single call","range","minmax",[1899,9321],[176,3895],"\u0001"],["Tool calls per agent-loop attempt",3,0,"count","3","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","\u0001","agent-loop-tool-calls",24,16,"Claude Code · agent loop","Claude Code · agent loop","range","minmax",[2,18],[0,3],"\u0001"],["List-price cost per strict pass: single call vs agent loop (calculation)",0.0672,0.01435,"usd","$0.067","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.014, 4.7x) is not tested against run-to-run variation.","single-call-vs-agent-loop",24,"agent-loop-cost-per-pass",24,24,"Claude Code · single call","Claude Code · single call","\u0001","\u0001","\u0001","\u0001",true],["Haiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right))",0.8659,0.939,"rate","87% (71/82)","94% (77/82)","tie","The 95% intervals overlap (Claude Haiku 4.5 78% to 92%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","haiku-thinking-on-off",82,"haiku-thinking-router-exact",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","ci95","ci95",[0.7755,0.9234],[0.8651,0.9737],"\u0001"],["Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)",0.9124,0.9742,"rate","91% (177/194)","97% (189/194)","tie","The 95% intervals overlap (Claude Haiku 4.5 86% to 94%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them.","haiku-thinking-on-off",194,"haiku-thinking-router-exact",194,194,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","ci95","ci95",[0.8642,0.9446],[0.9411,0.9889],"\u0001"],["Haiku thinking study: time per routing decision (Wall time (CLI))",4.66,2.6,"seconds","4.66 s","2.60 s","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 4.66 s to 8.18 s; Claude Sonnet 5.5 2.60 s to 4.30 s); not a confidence interval.","haiku-thinking-on-off",82,"haiku-thinking-router-latency",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","range","p50-p95",[4.66,8.18],[2.6,4.3],"\u0001"],["Haiku thinking study: time per routing decision (Model time (API))",3.79,1.6,"seconds","3.79 s","1.60 s","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 3.79 s to 7.43 s; Claude Sonnet 5.5 1.60 s to 2.58 s); not a confidence interval.","haiku-thinking-on-off",82,"haiku-thinking-router-latency",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","range","p50-p95",[3.79,7.43],[1.6,2.58],"\u0001"],["Haiku thinking study: thinking and visible output tokens per routing decision (Thinking tokens)",0,2,"tokens","0","2","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","haiku-thinking-on-off",82,"haiku-thinking-router-tokens",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","\u0001","\u0001","\u0001","\u0001","\u0001"],["Haiku thinking study: thinking and visible output tokens per routing decision (Visible output tokens)",366,105,"tokens","366","105","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","haiku-thinking-on-off",82,"haiku-thinking-router-tokens",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","\u0001","\u0001","\u0001","\u0001","\u0001"],["Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)",3.364,7.324,"usd","$3.36","$7.32","unclear","No interval or range was recorded for either side, so the gap ($3.36 vs $7.32, 2.2x) is not tested against run-to-run variation.","haiku-thinking-on-off",82,"haiku-thinking-router-cost",82,82,"Claude Code · thinking off · typed routing decisions, thinking on vs off","Claude Code · effort low · typed routing decisions, thinking on vs off","\u0001","\u0001","\u0001","\u0001",true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",0,1,"rate","0% (0/24)","100% (12/12)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; Claude Sonnet 5.5 76% to 100%).","json-schema-vs-instructions","\u0001","structured-output-pass-rate",24,12,"Claude Code · instructions","Claude Code · instructions","ci95","ci95",[0,0.138],[0.7575,1],true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))",0.7083,1,"rate","71% (17/24)","100% (12/12)","tie","The 95% intervals overlap (Claude Haiku 4.5 51% to 85%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions","\u0001","structured-output-pass-rate",24,12,"Claude Code · instructions","Claude Code · instructions","ci95","ci95",[0.5083,0.8509],[0.7575,1],true],["What each call produced: strict pass, format miss, wrong values or error (Strict pass)",0,12,"count","0","12","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Format miss)",17,0,"count","17","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Wrong values)",7,0,"count","7","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Error)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per call, instructions vs schema mode",9.52,3.52,"seconds","9.52 s","3.52 s","b","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 5.67 s to 17.0 s; Claude Sonnet 5.5 2.67 s to 4.12 s). A range is not a confidence interval.","json-schema-vs-instructions","\u0001","structured-output-time",24,12,"Claude Code · instructions","Claude Code · instructions","range","minmax",[5.67,17],[2.67,4.12],"\u0001"],["Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)",1128,368,"tokens","1,128","368","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-tokens",24,12,"Claude Code · instructions","Claude Code · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Unseen routing decisions answered exactly right",0.7857,0.875,"rate","79% (44/56)","88% (49/56)","tie","The 95% intervals overlap (Claude Haiku 4.5 66% to 87%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-exact",56,56,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6618,0.8729],[0.7637,0.9381],"\u0001"],["Per-question accuracy on unseen decisions",0.816,0.92,"rate","82% (102/125)","92% (115/125)","tie","The 95% intervals overlap (Claude Haiku 4.5 74% to 87%; Claude Sonnet 5.5 86% to 96%), so this sample cannot separate them.","routing-holdout",125,"routing-holdout-key-accuracy",125,125,"Claude Code","Claude Code · effort low","ci95","ci95",[0.739,0.8741],[0.859,0.956],"\u0001"],["Exact rate on unseen decisions, by decision type: Failure class",0.9286,1,"rate","93% (13/14)","100% (14/14)","tie","The 95% intervals overlap (Claude Haiku 4.5 69% to 99%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.7847,1],"\u0001"],["Exact rate on unseen decisions, by decision type: Message intent",0.9286,1,"rate","93% (13/14)","100% (14/14)","tie","The 95% intervals overlap (Claude Haiku 4.5 69% to 99%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.7847,1],"\u0001"],["Exact rate on unseen decisions, by decision type: Is it a rule?",0.9286,0.9286,"rate","93% (13/14)","93% (13/14)","tie","The 95% intervals overlap (Claude Haiku 4.5 69% to 99%; Claude Sonnet 5.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Context shape",0.3571,0.5714,"rate","36% (5/14)","57% (8/14)","tie","The 95% intervals overlap (Claude Haiku 4.5 16% to 61%; Claude Sonnet 5.5 33% to 79%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"Claude Code","Claude Code · effort low","ci95","ci95",[0.1634,0.6124],[0.3259,0.7862],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm))",0.8902,0.939,"rate","89% (73/82)","94% (77/82)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 94%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","routing-holdout",82,"routing-holdout-tuned-vs-unseen",82,82,"Claude Code","Claude Code · effort low","ci95","ci95",[0.8044,0.9412],[0.8651,0.9737],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Unseen holdout)",0.7857,0.875,"rate","79% (44/56)","88% (49/56)","tie","The 95% intervals overlap (Claude Haiku 4.5 66% to 87%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-tuned-vs-unseen",56,56,"Claude Code","Claude Code · effort low","ci95","ci95",[0.6618,0.8729],[0.7637,0.9381],"\u0001"],["Time per routing decision, by route (Wall time)",9.444,2.359,"seconds","9.44 s","2.36 s","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 9.44 s to 25.4 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.","routing-holdout",56,"routing-holdout-latency",56,56,"Claude Code","Claude Code · effort low","range","p50-p95",[9.444,25.413],[2.359,3.657],"\u0001"],["Time per routing decision, by route (Model time (API, CLI-reported))",7.522,1.485,"seconds","7.52 s","1.49 s","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 7.52 s to 23.9 s; Claude Sonnet 5.5 1.49 s to 2.38 s); not a confidence interval.","routing-holdout",56,"routing-holdout-latency",56,56,"Claude Code","Claude Code · effort low","range","p50-p95",[7.522,23.913],[1.485,2.377],"\u0001"],["Cost per 1,000 unseen routing decisions",7.129,7.244,"usd","$7.13","$7.24","unclear","No interval or range was recorded for either side, so the gap ($7.13 vs $7.24) is not tested against run-to-run variation.","routing-holdout",56,"routing-holdout-cost-per-1000",56,56,"Claude Code","Claude Code · effort low","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",91.68,54.54,"percent","91.7%","54.5%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[0,95.91],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.024492,0.006665,"usd","$0.024","$0.0067","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.001795,0.003672,"usd","$0.0018","$0.0037","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.00451,0.004012,"usd","$0.0045","$0.0040","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",91.68,54.54,"percent","91.7%","54.5%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[0,95.91],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",90.19,0,"percent","90.2%","0%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[73.1,97.59],[0,72.75],true],["Time to first text: a 250-line answer, six models",4,1.96,"seconds","4.00 s","1.96 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; Claude Sonnet 5.5 0.88 s to 4.09 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[2.84,6.38],[0.88,4.09],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",153.2,231.7,"tokens","153","232","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[152.6,216.1],[230.3,233],true],["Output speed in characters per second after the first text (calculation)",547,517,"count","547","517","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","\u0001","speed-anatomy-chars-per-second",3,4,"Claude Code","Claude Code","range","minmax",[546,548],[513,519],true],["Time to first text as the prompt grows: 1k",1.93,1.45,"seconds","1.93 s","1.45 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 1.85 s to 2.04 s; Claude Sonnet 5.5 1.23 s to 1.72 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.85,2.04],[1.23,1.72],true],["Time to first text as the prompt grows: 16k",2.27,1.78,"seconds","2.27 s","1.78 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.47 s; Claude Sonnet 5.5 1.64 s to 2.11 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[2.22,2.47],[1.64,2.11],true],["Time to first text as the prompt grows: 64k",2.78,3.07,"seconds","2.78 s","3.07 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.45 s to 2.89 s; Claude Sonnet 5.5 1.38 s to 3.61 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[2.45,2.89],[1.38,3.61],true],["Total time per call by prompt size (1k prompt)",2.34,1.78,"seconds","2.34 s","1.78 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.46 s; Claude Sonnet 5.5 1.57 s to 2.12 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.22,2.46],[1.57,2.12],"\u0001"],["Total time per call by prompt size (16k prompt)",2.79,2.1,"seconds","2.79 s","2.10 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.58 s to 2.84 s; Claude Sonnet 5.5 1.98 s to 2.48 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.58,2.84],[1.98,2.48],"\u0001"],["Total time per call by prompt size (64k prompt)",3.13,3.44,"seconds","3.13 s","3.44 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 3.28 s; Claude Sonnet 5.5 1.74 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.84,3.28],[1.74,4.38],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,1,"rate","100% (9/9)","100% (9/9)","tie","The 95% intervals overlap (Claude Haiku 4.5 70% to 100%; Claude Sonnet 5.5 70% to 100%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Claude Code","ci95","ci95",[0.7009,1],[0.7009,1],"\u0001"],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Interval merge fix",0.01789,0.00557,"usd","$0.018","$0.0056","unclear","No interval or range was recorded for either side, so the gap ($0.018 vs $0.0056, 3.2x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): DST day length",0.0293,0.02532,"usd","$0.029","$0.025","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.025) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): CSV parser",0.02919,0.01464,"usd","$0.029","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.015) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Event-loop order",0.03708,0.01588,"usd","$0.037","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.037 vs $0.016, 2.3x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Room schedule",0.03578,0.01243,"usd","$0.036","$0.012","unclear","No interval or range was recorded for either side, so the gap ($0.036 vs $0.012, 2.9x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SemVer regex",0.04217,0.00514,"usd","$0.042","$0.0051","unclear","No interval or range was recorded for either side, so the gap ($0.042 vs $0.0051, 8.2x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Money refactor",0.02039,0.00961,"usd","$0.020","$0.0096","unclear","No interval or range was recorded for either side, so the gap ($0.020 vs $0.0096, 2.1x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SQLite report query",0.03648,0.01719,"usd","$0.036","$0.017","unclear","No interval or range was recorded for either side, so the gap ($0.036 vs $0.017, 2.1x) is not tested against run-to-run variation.","haiku-retry-or-escalate",3,"retry-escalate-call-cost-by-task",3,3,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on 4 harder tasks (Strict pass)",0,0.375,"rate","0% (0/12)","38% (6/16)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 24%; Claude Sonnet 5.5 18% to 61%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Claude Code","ci95","ci95",[0,0.2425],[0.1848,0.6136],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0,0.375,"rate","0% (0/12)","38% (6/16)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 24%; Claude Sonnet 5.5 18% to 61%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Claude Code","ci95","ci95",[0,0.2425],[0.1848,0.6136],"\u0001"],["Calls that tried a tool although tools were off",0.0833,0.3125,"rate","8% (1/12)","31% (5/16)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-tool-attempts",12,16,"Claude Code","Claude Code","ci95","ci95",[0.0149,0.3539],[0.1416,0.556],"\u0001"],["Strict pass rate by task: 10x10 nonogram",0,1,"rate","0% (0/3)","100% (4/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 51% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.5101,1],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0,"rate","0% (0/3)","0% (0/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 0% to 49%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0,0.4899],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,0,"rate","0% (0/3)","0% (0/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 0% to 49%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0,0.4899],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0,0.5,"rate","0% (0/3)","50% (2/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 15% to 85%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.15,0.85],"\u0001"],["Total time per call on harder tasks",108.98,70.43,"seconds","109.0 s","70.4 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 25.7 s to 223.9 s; Claude Sonnet 5.5 4.32 s to 210.1 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",10,12,"Claude Code","Claude Code","range","minmax",[25.73,223.95],[4.32,210.08],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",12508,9287,"tokens","12,508","9,287","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",10,12,"Claude Code","Claude Code","range","minmax",[2965,26532],[407,27921],"\u0001"]]}],["claude-opus-5-5-vs-claude-fable-5-1","claude-opus-5-5","claude-fable-5-1","Claude Opus 5.5 vs Claude Fable 5.1","Claude Opus 5.5 vs Claude Fable 5.1: measured benchmarks","Claude Opus 5.5 vs Claude Fable 5.1: 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Opus 5.5 and Claude Fable 5.1 share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 5 ties and 18 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 4 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 80% to 100%; Claude Fable 5.1 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.75,1.94,"seconds","2.75 s","1.94 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.47 s to 8.91 s; Claude Fable 5.1 1.41 s to 9.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.47,8.91],[1.41,9.83],"\u0001"],["Time to first useful output",1.92,1.2,"seconds","1.92 s","1.20 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.56 s to 7.23 s; Claude Fable 5.1 0.95 s to 7.90 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[1.56,7.23],[0.95,7.9],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1401,2760,"tokens","1,401","2,760","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",680,473,"tokens","680","473","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",64,64,"tokens","64","64","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00688,0.00987,"usd","$0.0069","$0.0099","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 $0.0059 to $0.022; Claude Fable 5.1 $0.0049 to $0.058); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00592,0.02226],[0.0049,0.05843],true],["List-price cost per passing answer (calculation)",0.01009,0.02054,"usd","$0.010","$0.021","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.021, 2.0x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; Claude Fable 5.1 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; Claude Fable 5.1 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",9.18,16.13,"seconds","9.18 s","16.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 4.24 s to 27.2 s; Claude Fable 5.1 4.46 s to 90.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[4.24,27.21],[4.46,90],"\u0001"],["Time to first useful output on hard tasks",6.78,11.63,"seconds","6.78 s","11.6 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.39 s to 21.8 s; Claude Fable 5.1 2.00 s to 85.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[2.39,21.77],[2,85.33],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",945,1366,"tokens","945","1,366","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.02824,0.09331,"usd","$0.028","$0.093","unclear","No interval or range was recorded for either side, so the gap ($0.028 vs $0.093, 3.3x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",54.79,64.24,"percent","54.8%","64.2%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[29.92,95.6],[23.44,97.19],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.012528,0.053696,"usd","$0.013","$0.054","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.008003,0.018702,"usd","$0.0080","$0.019","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.00771,0.02091,"usd","$0.0077","$0.021","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.79,64.24,"percent","54.8%","64.2%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[29.92,95.6],[23.44,97.19],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",0,0,"percent","0%","0%","tie","Same value. More or fewer is not better by itself for this metric.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[0,93.33],[0,74.01],true],["Time to first text: a 250-line answer, six models",1.97,4.43,"seconds","1.97 s","4.43 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.70 s to 2.35 s; Claude Fable 5.1 2.27 s to 4.64 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[1.7,2.35],[2.27,4.64],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",155.5,122.6,"tokens","156","123","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[154.6,156.4],[120.9,131.4],true],["Output speed in characters per second after the first text (calculation)",347,273,"count","347","273","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Code","Claude Code","range","minmax",[345,349],[270,293],true]]}],["claude-sonnet-5-5-vs-claude-fable-5-1","claude-sonnet-5-5","claude-fable-5-1","Claude Sonnet 5.5 vs Claude Fable 5.1","Claude Sonnet 5.5 vs Claude Fable 5.1: measured benchmarks","Claude Sonnet 5.5 vs Claude Fable 5.1: 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Sonnet 5.5 and Claude Fable 5.1 share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 4 ties and 19 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 4 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; Claude Fable 5.1 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,1.94,"seconds","2.31 s","1.94 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; Claude Fable 5.1 1.41 s to 9.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.17,7.73],[1.41,9.83],"\u0001"],["Time to first useful output",1.56,1.2,"seconds","1.56 s","1.20 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; Claude Fable 5.1 0.95 s to 7.90 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.99,6.39],[0.95,7.9],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1401,2760,"tokens","1,401","2,760","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",685,473,"tokens","685","473","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",107,64,"tokens","107","64","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.0036,0.00987,"usd","$0.0036","$0.0099","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 $0.0034 to $0.010; Claude Fable 5.1 $0.0049 to $0.058); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00342,0.01021],[0.0049,0.05843],true],["List-price cost per passing answer (calculation)",0.00624,0.02054,"usd","$0.0062","$0.021","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.021, 3.3x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Fable 5.1 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Fable 5.1 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,16.13,"seconds","7.75 s","16.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; Claude Fable 5.1 4.46 s to 90.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[2.26,34.79],[4.46,90],"\u0001"],["Time to first useful output on hard tasks",5.95,11.63,"seconds","5.95 s","11.6 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.86 s to 30.6 s; Claude Fable 5.1 2.00 s to 85.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[0.86,30.57],[2,85.33],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1050,1366,"tokens","1,050","1,366","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.09331,"usd","$0.014","$0.093","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.093, 6.5x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",54.54,64.24,"percent","54.5%","64.2%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[0,95.91],[23.44,97.19],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.006665,0.053696,"usd","$0.0067","$0.054","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.003672,0.018702,"usd","$0.0037","$0.019","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.004012,0.02091,"usd","$0.0040","$0.021","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.54,64.24,"percent","54.5%","64.2%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[0,95.91],[23.44,97.19],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",0,0,"percent","0%","0%","tie","Same value. More or fewer is not better by itself for this metric.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[0,72.75],[0,74.01],true],["Time to first text: a 250-line answer, six models",1.96,4.43,"seconds","1.96 s","4.43 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; Claude Fable 5.1 2.27 s to 4.64 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[0.88,4.09],[2.27,4.64],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",231.7,122.6,"tokens","232","123","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[230.3,233],[120.9,131.4],true],["Output speed in characters per second after the first text (calculation)",517,273,"count","517","273","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Code","Claude Code","range","minmax",[513,519],[270,293],true]]}],["claude-haiku-4-5-vs-claude-opus-5-5","claude-haiku-4-5","claude-opus-5-5","Claude Haiku 4.5 vs Claude Opus 5.5","Claude Haiku 4.5 vs Claude Opus 5.5: measured benchmarks","Claude Haiku 4.5 vs Claude Opus 5.5: 25 measured metrics from 4 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Haiku 4.5 and Claude Opus 5.5 share 25 measured metrics and 14 list-price calculations from 5 studies. Claude Opus 5.5 leads on 3 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Pass rate on 4 harder tasks (Lenient (format misses counted)), 50% (6/12) vs 0% (0/12). On those rows the 95% intervals do not overlap. The other rows are 7 ties and 29 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",4.43,2.75,"seconds","4.43 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[3.16,23.57],[2.47,8.91],"\u0001"],["Time to first useful output",3.63,1.92,"seconds","3.63 s","1.92 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; Claude Opus 5.5 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.78,22.27],[1.56,7.23],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",0,1401,"tokens","0","1,401","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",3790,680,"tokens","3,790","680","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",367,64,"tokens","367","64","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00566,0.00688,"usd","$0.0057","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; Claude Opus 5.5 $0.0059 to $0.022); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00513,0.01804],[0.00592,0.02226],true],["List-price cost per passing answer (calculation)",0.00836,0.01009,"usd","$0.0084","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Opus 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Opus 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,9.18,"seconds","39.0 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[15.27,75.13],[4.24,27.21],"\u0001"],["Time to first useful output on hard tasks",35.54,6.78,"seconds","35.5 s","6.78 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; Claude Opus 5.5 2.39 s to 21.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[12.88,70.31],[2.39,21.77],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",5064,945,"tokens","5,064","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.02824,"usd","$0.067","$0.028","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.028, 2.4x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",91.68,54.79,"percent","91.7%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[29.92,95.6],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.024492,0.012528,"usd","$0.024","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.001795,0.008003,"usd","$0.0018","$0.0080","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.00451,0.00771,"usd","$0.0045","$0.0077","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",91.68,54.79,"percent","91.7%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[29.92,95.6],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",90.19,0,"percent","90.2%","0%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[73.1,97.59],[0,93.33],true],["Time to first text: a 250-line answer, six models",4,1.97,"seconds","4.00 s","1.97 s","unclear","Only 4 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.84 s to 6.38 s; Claude Opus 5.5 1.70 s to 2.35 s), but 4 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[2.84,6.38],[1.7,2.35],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",153.2,155.5,"tokens","153","156","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[152.6,216.1],[154.6,156.4],true],["Output speed in characters per second after the first text (calculation)",547,347,"count","547","347","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","\u0001","speed-anatomy-chars-per-second",3,4,"Claude Code","Claude Code","range","minmax",[546,548],[345,349],true],["Time to first text as the prompt grows: 1k",1.93,1.51,"seconds","1.93 s","1.51 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 1.85 s to 2.04 s; Claude Opus 5.5 1.46 s to 2.01 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[1.85,2.04],[1.46,2.01],true],["Time to first text as the prompt grows: 16k",2.27,1.74,"seconds","2.27 s","1.74 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.22 s to 2.47 s; Claude Opus 5.5 1.70 s to 2.97 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[2.22,2.47],[1.7,2.97],true],["Time to first text as the prompt grows: 64k",2.78,1.79,"seconds","2.78 s","1.79 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.45 s to 2.89 s; Claude Opus 5.5 1.72 s to 3.72 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Claude Code","range","minmax",[2.45,2.89],[1.72,3.72],true],["Total time per call by prompt size (1k prompt)",2.34,1.83,"seconds","2.34 s","1.83 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.22 s to 2.46 s; Claude Opus 5.5 1.82 s to 2.41 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.22,2.46],[1.82,2.41],"\u0001"],["Total time per call by prompt size (16k prompt)",2.79,2.36,"seconds","2.79 s","2.36 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.58 s to 2.84 s; Claude Opus 5.5 2.11 s to 3.40 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.58,2.84],[2.11,3.4],"\u0001"],["Total time per call by prompt size (64k prompt)",3.13,2.35,"seconds","3.13 s","2.35 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 3.28 s; Claude Opus 5.5 2.26 s to 4.29 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Claude Code","range","minmax",[2.84,3.28],[2.26,4.29],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,0.5556,"rate","100% (9/9)","56% (5/9)","tie","The 95% intervals overlap (Claude Haiku 4.5 70% to 100%; Claude Opus 5.5 27% to 81%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Claude Code","ci95","ci95",[0.7009,1],[0.2667,0.8112],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0,0.4167,"rate","0% (0/12)","42% (5/12)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 24%; Claude Opus 5.5 19% to 68%), so this sample cannot separate them.","harder-tasks-head-to-head",12,"harder-h2h-pass-rate",12,12,"Claude Code","Claude Code","ci95","ci95",[0,0.2425],[0.1933,0.6805],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0,0.5,"rate","0% (0/12)","50% (6/12)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; Claude Opus 5.5 25% to 75%).","harder-tasks-head-to-head",12,"harder-h2h-pass-rate",12,12,"Claude Code","Claude Code","ci95","ci95",[0,0.2425],[0.2538,0.7462],"\u0001"],["Calls that tried a tool although tools were off",0.0833,0.4167,"rate","8% (1/12)","42% (5/12)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head",12,"harder-h2h-tool-attempts",12,12,"Claude Code","Claude Code","ci95","ci95",[0.0149,0.3539],[0.1933,0.6805],"\u0001"],["Strict pass rate by task: 10x10 nonogram",0,1,"rate","0% (0/3)","100% (3/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Opus 5.5 44% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head",3,"harder-h2h-pass-by-task",3,3,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.4385,1],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0,"rate","0% (0/3)","0% (0/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Opus 5.5 0% to 56%), so this sample cannot separate them.","harder-tasks-head-to-head",3,"harder-h2h-pass-by-task",3,3,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0,0.5615],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,0.3333,"rate","0% (0/3)","33% (1/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them.","harder-tasks-head-to-head",3,"harder-h2h-pass-by-task",3,3,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.0615,0.7923],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0,0.3333,"rate","0% (0/3)","33% (1/3)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them.","harder-tasks-head-to-head",3,"harder-h2h-pass-by-task",3,3,"Claude Code","Claude Code","ci95","ci95",[0,0.5615],[0.0615,0.7923],"\u0001"],["Total time per call on harder tasks",108.98,80.34,"seconds","109.0 s","80.3 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 25.7 s to 223.9 s; Claude Opus 5.5 3.82 s to 279.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",10,9,"Claude Code","Claude Code","range","minmax",[25.73,223.95],[3.82,279.5],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",12508,8420,"tokens","12,508","8,420","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",10,9,"Claude Code","Claude Code","range","minmax",[2965,26532],[279,40044],"\u0001"]]}],["claude-code-cli-vs-codex-cli","claude-code-cli","codex-cli","Claude Code vs Codex CLI","Claude Code vs Codex CLI: measured benchmarks","Claude Code vs Codex CLI: 66 measured metrics from 11 studies (Pass rate on five validated tasks; Total time per call; more), with sample sizes and intervals.","Claude Code and Codex CLI share 66 measured metrics and 22 list-price calculations from 12 studies. Claude Code leads on 7 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 4 more. Codex CLI leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 31 ties and 49 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Each run pairs a CLI with a model, so these rows cannot separate the CLI from the model; the contexts name both. Some rows rest on small samples (n = 2 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Code 55% to 93%; Codex CLI 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,5.65,"seconds","2.31 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.17 s to 7.73 s; Codex CLI 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","range","minmax",[2.17,7.73],[4.1,25.46],"\u0001"],["Time to first useful output",1.56,5.05,"seconds","1.56 s","5.05 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 0.99 s to 6.39 s; Codex CLI 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","range","minmax",[0.99,6.39],[3.36,17.82],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1401,5180,"tokens","1,401","5,180","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",685,6943,"tokens","685","6,943","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",107,42,"tokens","107","42","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.0036,0.01018,"usd","$0.0036","$0.010","unclear","The run ranges (fastest to slowest) overlap (Claude Code $0.0034 to $0.010; Codex CLI $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","range","minmax",[0.00342,0.01021],[0.0054,0.02686],true],["List-price cost per passing answer (calculation)",0.00624,0.01564,"usd","$0.0062","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,13.11,"seconds","7.75 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","range","minmax",[2.26,34.79],[8.54,61.6],"\u0001"],["Time to first useful output on hard tasks",5.95,10.23,"seconds","5.95 s","10.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 0.86 s to 30.6 s; Codex CLI 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-first-useful-latency",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","range","minmax",[0.86,30.57],[6.09,40.41],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1050,335,"tokens","1,050","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head","\u0001","hard-h2h-output-tokens",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",23.1,113.4,"seconds","23.1 s","113.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[18.7,44.5],[78.5,221.9],"\u0001"],["Tool calls per coding session",7.5,12.5,"calls","7.5","12.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","coding-agents-head-to-head",12,"coding-agents-tool-calls",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[3,14],[8,18],"\u0001"],["List-price cost per passing coding session (calculation)",0.085,0.0978,"usd","$0.085","$0.098","unclear","No interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Code 81% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.63,13.11,"seconds","7.63 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.71 s to 24.0 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","range","minmax",[2.71,24.01],[8.54,61.6],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",770,335,"tokens","770","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01352,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: how many different answers (Exact number)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (JSON object)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (Code fix)",3,6,"count","3","6","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: time per call (Exact number)",6.89,13.38,"seconds","6.89 s","13.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","range","minmax",[5.81,7.81],[12.29,17.97],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",2.89,6.42,"seconds","2.89 s","6.42 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.68 s to 5.30 s; Codex CLI 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","range","minmax",[2.68,5.3],[5.25,8.26],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",2.67,11.29,"seconds","2.67 s","11.3 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","range","minmax",[2.32,4.34],[9.08,14.85],"\u0001"],["CLI start-up tax on a one-word answer (First output event)",563,489,"ms","563 ms","489 ms","unclear","The run ranges (fastest to slowest) overlap (Claude Code 519 ms to 726 ms; Codex CLI 354 ms to 1,304 ms); the medians alone do not show a reliable difference. A range is not a confidence interval.","routing-overhead",5,"cli-startup-tax",5,5,"Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","default model · CLI start-up, one-word prompt, 5 runs","range","minmax",[519,726],[354,1304],"\u0001"],["CLI start-up tax on a one-word answer (First model output)",1461,5059,"ms","1,461 ms","5,059 ms","a","The run ranges (fastest to slowest) do not overlap (Claude Code 1,206 ms to 2,308 ms; Codex CLI 4,391 ms to 5,478 ms). A range is not a confidence interval. Samples are small (5 runs per side).","routing-overhead",5,"cli-startup-tax",5,5,"Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","default model · CLI start-up, one-word prompt, 5 runs","range","minmax",[1206,2308],[4391,5478],"\u0001"],["CLI start-up tax on a one-word answer (Total wall time)",2529,5999,"ms","2,529 ms","5,999 ms","a","The run ranges (fastest to slowest) do not overlap (Claude Code 2,273 ms to 3,382 ms; Codex CLI 5,367 ms to 6,506 ms). A range is not a confidence interval. Samples are small (5 runs per side).","routing-overhead",5,"cli-startup-tax",5,5,"Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","default model · CLI start-up, one-word prompt, 5 runs","range","minmax",[2273,3382],[5367,6506],"\u0001"],["Input tokens a CLI sends for a one-word answer",6761,17051,"tokens","6,761","17,051","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","routing-overhead",5,"cli-startup-input-tokens",5,5,"Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs","default model · CLI start-up, one-word prompt, 5 runs","\u0001","\u0001","\u0001","\u0001","\u0001"],["Repairing a scheduler: Claude Code vs Codex vs API (Total time)",15,61.16,"seconds","15.0 s","61.2 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 13.9 s to 15.9 s; Codex CLI 59.9 s to 69.5 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs","GPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[13.89,15.89],[59.9,69.51],"\u0001"],["Repairing a scheduler: Claude Code vs Codex vs API (First useful output)",7.55,15.56,"seconds","7.55 s","15.6 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 6.77 s to 7.63 s; Codex CLI 13.7 s to 23.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs","GPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[6.77,7.63],[13.65,23.04],"\u0001"],["Output tokens to repair the scheduler (Output tokens)",2227,1181,"tokens","2,227","1,181","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",3,"scheduler-repair-output-tokens",3,3,"Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs","GPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs","\u0001","\u0001","\u0001","\u0001","\u0001"],["Strict pass rate: single call vs agent loop on eight hard tasks",1,0.625,"rate","100% (24/24)","63% (10/16)","a","The 95% intervals do not overlap (Claude Code 86% to 100%; Codex CLI 39% to 82%).","single-call-vs-agent-loop","\u0001","agent-loop-pass-rate",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.862,1],[0.3864,0.8152],"\u0001"],["Strict passes per task: single call vs agent loop: Interval merge fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001"],["Strict passes per task: single call vs agent loop: DST day-length fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001"],["Strict passes per task: single call vs agent loop: CSV parser",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001"],["Strict passes per task: single call vs agent loop: Event-loop order",1,0,"rate","100% (3/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0,0.6576],"\u0001"],["Strict passes per task: single call vs agent loop: Room schedule",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001"],["Strict passes per task: single call vs agent loop: SemVer regex",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001"],["Strict passes per task: single call vs agent loop: Money refactor",1,0,"rate","100% (3/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0,0.6576],"\u0001"],["Strict passes per task: single call vs agent loop: SQL report",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","\u0001","agent-loop-by-task",3,2,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001"],["Total time per attempt: single call vs agent loop",7.75,5.16,"seconds","7.75 s","5.16 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","single-call-vs-agent-loop","\u0001","agent-loop-total-time",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","range","minmax",[2.26,34.79],[3.59,11.32],"\u0001"],["Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))",2281,11582,"tokens","2,281","11,582","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","\u0001","agent-loop-tokens",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","range","minmax",[2234,2669],[11526,11818],"\u0001"],["Tokens per attempt: single call vs agent loop (Output tokens)",1050,345,"tokens","1,050","345","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","\u0001","agent-loop-tokens",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","range","minmax",[176,3895],[36,634],"\u0001"],["Tool calls per agent-loop attempt",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","single-call-vs-agent-loop","\u0001","agent-loop-tool-calls",16,14,"Claude Sonnet 5.5 · agent loop","GPT-6 Luna · agent loop","range","minmax",[0,3],[0,1],"\u0001"],["List-price cost per strict pass: single call vs agent loop (calculation)",0.01435,0.00116,"usd","$0.014","$0.0012","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation.","single-call-vs-agent-loop","\u0001","agent-loop-cost-per-pass",24,16,"Claude Sonnet 5.5 · single call","GPT-6 Luna · single call","\u0001","\u0001","\u0001","\u0001",true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions",12,"structured-output-pass-rate",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","ci95","ci95",[0.7575,1],[0.7575,1],true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions",12,"structured-output-pass-rate",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","ci95","ci95",[0.7575,1],[0.7575,1],true],["What each call produced: strict pass, format miss, wrong values or error (Strict pass)",12,12,"count","12","12","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Format miss)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Wrong values)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Error)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per call, instructions vs schema mode",3.52,6.21,"seconds","3.52 s","6.21 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 2.67 s to 4.12 s; Codex CLI 4.20 s to 12.3 s). A range is not a confidence interval.","json-schema-vs-instructions",12,"structured-output-time",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","range","minmax",[2.67,4.12],[4.2,12.27],"\u0001"],["Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)",368,117,"tokens","368","117","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions",12,"structured-output-tokens",12,12,"Claude Sonnet 5.5 · instructions","GPT-6.1 Sol · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Reasoning share of output tokens per call on hard tasks (calculation)",54.54,46.33,"percent","54.5%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-share",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[0,95.91],[11.42,86.85],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.006665,0.002273,"usd","$0.0067","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.003672,0.003025,"usd","$0.0037","$0.0030","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.004012,0.020339,"usd","$0.0040","$0.020","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005946,0.002273,"usd","$0.0059","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Sonnet 5.5 · effort medium","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.01352,0.025637,"usd","$0.014","$0.026","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Sonnet 5.5 · effort medium","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.54,46.33,"percent","54.5%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-short-vs-hard",24,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[0,95.91],[11.42,86.85],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",0,41.05,"percent","0%","41%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[0,72.75],[0,71.43],true],["Time to first text: a 250-line answer, six models",1.96,3.52,"seconds","1.96 s","3.52 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 0.88 s to 4.09 s; Codex CLI 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[0.88,4.09],[2.75,4.42],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",231.7,79.6,"tokens","232","80","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[230.3,233],[71.6,80.5],true],["Output speed in characters per second after the first text (calculation)",517,323,"count","517","323","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[513,519],[291,327],true],["Time to first text as the prompt grows: 1k",1.45,3.36,"seconds","1.45 s","3.36 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.23 s to 1.72 s; Codex CLI 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.23,1.72],[3.36,4.75],true],["Time to first text as the prompt grows: 16k",1.78,4.02,"seconds","1.78 s","4.02 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.64 s to 2.11 s; Codex CLI 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.64,2.11],[3.3,4.28],true],["Time to first text as the prompt grows: 64k",3.07,3.93,"seconds","3.07 s","3.93 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 1.38 s to 3.61 s; Codex CLI 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.38,3.61],[3.42,4.38],true],["Total time per call by prompt size (1k prompt)",1.78,3.43,"seconds","1.78 s","3.43 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.57 s to 2.12 s; Codex CLI 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.57,2.12],[3.43,4.92],"\u0001"],["Total time per call by prompt size (16k prompt)",2.1,4.14,"seconds","2.10 s","4.14 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.98 s to 2.48 s; Codex CLI 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.98,2.48],[3.96,4.68],"\u0001"],["Total time per call by prompt size (64k prompt)",3.44,3.96,"seconds","3.44 s","3.96 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 1.74 s to 4.38 s; Codex CLI 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","range","minmax",[1.74,4.38],[3.47,4.44],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,1,"rate","100% (9/9)","100% (9/9)","tie","The 95% intervals overlap (Claude Code 70% to 100%; Codex CLI 70% to 100%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Sonnet 5.5","GPT-6.1 Sol · effort low","ci95","ci95",[0.7009,1],[0.7009,1],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0.375,0.6875,"rate","38% (6/16)","69% (11/16)","tie","The 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head",16,"harder-h2h-pass-rate",16,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.1848,0.6136],[0.444,0.8584],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0.375,0.6875,"rate","38% (6/16)","69% (11/16)","tie","The 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head",16,"harder-h2h-pass-rate",16,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.1848,0.6136],[0.444,0.8584],"\u0001"],["Calls that tried a tool although tools were off",0.3125,0,"rate","31% (5/16)","0% (0/16)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head",16,"harder-h2h-tool-attempts",16,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.1416,0.556],[0,0.1936],"\u0001"],["Strict pass rate by task: 10x10 nonogram",1,0.75,"rate","100% (4/4)","75% (3/4)","tie","The 95% intervals overlap (Claude Code 51% to 100%; Codex CLI 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.5101,1],[0.3006,0.9544],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0.25,"rate","0% (0/4)","25% (1/4)","tie","The 95% intervals overlap (Claude Code 0% to 49%; Codex CLI 5% to 70%), so this sample cannot separate them.","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0,0.4899],[0.0456,0.6994],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,1,"rate","0% (0/4)","100% (4/4)","b","The 95% intervals do not overlap (Claude Code 0% to 49%; Codex CLI 51% to 100%).","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0,0.4899],[0.5101,1],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0.5,0.75,"rate","50% (2/4)","75% (3/4)","tie","The 95% intervals overlap (Claude Code 15% to 85%; Codex CLI 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","ci95","ci95",[0.15,0.85],[0.3006,0.9544],"\u0001"],["Total time per call on harder tasks",70.43,120.24,"seconds","70.4 s","120.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 4.32 s to 210.1 s; Codex CLI 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",12,13,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[4.32,210.08],[46.24,273.46],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",9287,4994,"tokens","9,287","4,994","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",12,13,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","range","minmax",[407,27921],[2099,13413],"\u0001"],["List-price cost per strict pass on harder tasks (calculation)",0.23843,0.08293,"usd","$0.24","$0.083","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $0.083, 2.9x) is not tested against run-to-run variation.","harder-tasks-head-to-head",16,"harder-h2h-cost-per-pass",16,16,"Claude Sonnet 5.5","GPT-6.1 Sol · effort medium","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-vs-gpt-6-1-sol-codex-cli","claude-sonnet-5-5","gpt-6-1-sol-codex-cli","Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)","Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): benchmarks","Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): 49 measured metrics from 9 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) share 49 measured metrics and 21 list-price calculations from 10 studies. Claude Sonnet 5.5 leads on 4 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 1 more. GPT-6.1 Sol (Codex CLI) leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 22 ties and 43 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,5.65,"seconds","2.31 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[2.17,7.73],[4.1,25.46],"\u0001"],["Time to first useful output",1.56,5.05,"seconds","1.56 s","5.05 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.99,6.39],[3.36,17.82],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1401,5180,"tokens","1,401","5,180","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",685,6943,"tokens","685","6,943","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",107,42,"tokens","107","42","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.0036,0.01018,"usd","$0.0036","$0.010","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 $0.0034 to $0.010; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.00342,0.01021],[0.0054,0.02686],true],["List-price cost per passing answer (calculation)",0.00624,0.01564,"usd","$0.0062","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,13.11,"seconds","7.75 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[2.26,34.79],[8.54,61.6],"\u0001"],["Time to first useful output on hard tasks",5.95,10.23,"seconds","5.95 s","10.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.86 s to 30.6 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-first-useful-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[0.86,30.57],[6.09,40.41],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1050,335,"tokens","1,050","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head","\u0001","hard-h2h-output-tokens",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",23.1,113.4,"seconds","23.1 s","113.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s). A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[18.7,44.5],[78.5,221.9],"\u0001"],["Tool calls per coding session",7.5,12.5,"calls","7.5","12.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","coding-agents-head-to-head",12,"coding-agents-tool-calls",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[3,14],[8,18],"\u0001"],["List-price cost per passing coding session (calculation)",0.085,0.0978,"usd","$0.085","$0.098","unclear","No interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.63,13.11,"seconds","7.63 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.71 s to 24.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","range","minmax",[2.71,24.01],[8.54,61.6],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",770,335,"tokens","770","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01352,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: how many different answers (Exact number)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (JSON object)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (Code fix)",3,6,"count","3","6","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: time per call (Exact number)",6.89,13.38,"seconds","6.89 s","13.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 5.81 s to 7.81 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[5.81,7.81],[12.29,17.97],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",2.89,6.42,"seconds","2.89 s","6.42 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.68 s to 5.30 s; GPT-6.1 Sol (Codex CLI) 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[2.68,5.3],[5.25,8.26],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",2.67,11.29,"seconds","2.67 s","11.3 s","a","The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.32 s to 4.34 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[2.32,4.34],[9.08,14.85],"\u0001"],["Repairing a scheduler: Claude Code vs Codex vs API (Total time)",15,61.16,"seconds","15.0 s","61.2 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 13.9 s to 15.9 s; GPT-6.1 Sol (Codex CLI) 59.9 s to 69.5 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs","Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[13.89,15.89],[59.9,69.51],"\u0001"],["Repairing a scheduler: Claude Code vs Codex vs API (First useful output)",7.55,15.56,"seconds","7.55 s","15.6 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 6.77 s to 7.63 s; GPT-6.1 Sol (Codex CLI) 13.7 s to 23.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs","Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[6.77,7.63],[13.65,23.04],"\u0001"],["Output tokens to repair the scheduler (Output tokens)",2227,1181,"tokens","2,227","1,181","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",3,"scheduler-repair-output-tokens",3,3,"Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs","Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs","\u0001","\u0001","\u0001","\u0001","\u0001"],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions",12,"structured-output-pass-rate",12,12,"Claude Code · instructions","Codex CLI · effort low · instructions","ci95","ci95",[0.7575,1],[0.7575,1],true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions",12,"structured-output-pass-rate",12,12,"Claude Code · instructions","Codex CLI · effort low · instructions","ci95","ci95",[0.7575,1],[0.7575,1],true],["What each call produced: strict pass, format miss, wrong values or error (Strict pass)",12,12,"count","12","12","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Format miss)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Wrong values)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Error)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions",12,"structured-output-outcomes",12,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per call, instructions vs schema mode",3.52,6.21,"seconds","3.52 s","6.21 s","a","The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.67 s to 4.12 s; GPT-6.1 Sol (Codex CLI) 4.20 s to 12.3 s). A range is not a confidence interval.","json-schema-vs-instructions",12,"structured-output-time",12,12,"Claude Code · instructions","Codex CLI · effort low · instructions","range","minmax",[2.67,4.12],[4.2,12.27],"\u0001"],["Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)",368,117,"tokens","368","117","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions",12,"structured-output-tokens",12,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Reasoning share of output tokens per call on hard tasks (calculation)",54.54,46.33,"percent","54.5%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-share",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[0,95.91],[11.42,86.85],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.006665,0.002273,"usd","$0.0067","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.003672,0.003025,"usd","$0.0037","$0.0030","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.004012,0.020339,"usd","$0.0040","$0.020","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005946,0.002273,"usd","$0.0059","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.01352,0.025637,"usd","$0.014","$0.026","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.54,46.33,"percent","54.5%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-short-vs-hard",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[0,95.91],[11.42,86.85],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",0,41.05,"percent","0%","41%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Codex CLI · effort medium","range","minmax",[0,72.75],[0,71.43],true],["Time to first text: a 250-line answer, six models",1.96,3.52,"seconds","1.96 s","3.52 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[0.88,4.09],[2.75,4.42],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",231.7,79.6,"tokens","232","80","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[230.3,233],[71.6,80.5],true],["Output speed in characters per second after the first text (calculation)",517,323,"count","517","323","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[513,519],[291,327],true],["Time to first text as the prompt grows: 1k",1.45,3.36,"seconds","1.45 s","3.36 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.23 s to 1.72 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.23,1.72],[3.36,4.75],true],["Time to first text as the prompt grows: 16k",1.78,4.02,"seconds","1.78 s","4.02 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.64 s to 2.11 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.64,2.11],[3.3,4.28],true],["Time to first text as the prompt grows: 64k",3.07,3.93,"seconds","3.07 s","3.93 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.38 s to 3.61 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.38,3.61],[3.42,4.38],true],["Total time per call by prompt size (1k prompt)",1.78,3.43,"seconds","1.78 s","3.43 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.57 s to 2.12 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.57,2.12],[3.43,4.92],"\u0001"],["Total time per call by prompt size (16k prompt)",2.1,4.14,"seconds","2.10 s","4.14 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.98 s to 2.48 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.98,2.48],[3.96,4.68],"\u0001"],["Total time per call by prompt size (64k prompt)",3.44,3.96,"seconds","3.44 s","3.96 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.74 s to 4.38 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.74,4.38],[3.47,4.44],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,1,"rate","100% (9/9)","100% (9/9)","tie","The 95% intervals overlap (Claude Sonnet 5.5 70% to 100%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Codex CLI · effort low","ci95","ci95",[0.7009,1],[0.7009,1],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0.375,0.6875,"rate","38% (6/16)","69% (11/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head",16,"harder-h2h-pass-rate",16,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.1848,0.6136],[0.444,0.8584],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0.375,0.6875,"rate","38% (6/16)","69% (11/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head",16,"harder-h2h-pass-rate",16,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.1848,0.6136],[0.444,0.8584],"\u0001"],["Calls that tried a tool although tools were off",0.3125,0,"rate","31% (5/16)","0% (0/16)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head",16,"harder-h2h-tool-attempts",16,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.1416,0.556],[0,0.1936],"\u0001"],["Strict pass rate by task: 10x10 nonogram",1,0.75,"rate","100% (4/4)","75% (3/4)","tie","The 95% intervals overlap (Claude Sonnet 5.5 51% to 100%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.5101,1],[0.3006,0.9544],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0.25,"rate","0% (0/4)","25% (1/4)","tie","The 95% intervals overlap (Claude Sonnet 5.5 0% to 49%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them.","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.4899],[0.0456,0.6994],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,1,"rate","0% (0/4)","100% (4/4)","b","The 95% intervals do not overlap (Claude Sonnet 5.5 0% to 49%; GPT-6.1 Sol (Codex CLI) 51% to 100%).","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.4899],[0.5101,1],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0.5,0.75,"rate","50% (2/4)","75% (3/4)","tie","The 95% intervals overlap (Claude Sonnet 5.5 15% to 85%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head",4,"harder-h2h-pass-by-task",4,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.15,0.85],[0.3006,0.9544],"\u0001"],["Total time per call on harder tasks",70.43,120.24,"seconds","70.4 s","120.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.32 s to 210.1 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",12,13,"Claude Code","Codex CLI · effort medium","range","minmax",[4.32,210.08],[46.24,273.46],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",9287,4994,"tokens","9,287","4,994","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",12,13,"Claude Code","Codex CLI · effort medium","range","minmax",[407,27921],[2099,13413],"\u0001"],["List-price cost per strict pass on harder tasks (calculation)",0.23843,0.08293,"usd","$0.24","$0.083","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $0.083, 2.9x) is not tested against run-to-run variation.","harder-tasks-head-to-head",16,"harder-h2h-cost-per-pass",16,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true]]}],["jev-1-13-vs-claude-haiku-4-5","jev-1-13","claude-haiku-4-5","Jev 1.13 vs Claude Haiku 4.5","Jev 1.13 vs Claude Haiku 4.5: measured benchmarks","Jev 1.13 vs Claude Haiku 4.5: 17 measured metrics from 3 studies (Typed routing decisions answered exactly right; more), with sample sizes and intervals.","Jev 1.13 and Claude Haiku 4.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 12,543 ms; Time per routing decision, by route (Wall time), 0.14 s vs 9.44 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Typed routing decisions answered exactly right",0.8984,0.8902,"rate","90%","89% (73/82)","tie","The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them.","routing-jev-vs-llm",82,"routing-exact-decisions",82,82,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8191,0.9497],[0.8044,0.9412],"\u0001"],["Per-question accuracy",0.9485,0.9433,"rate","95% (184/194)","94% (183/194)","tie","The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Haiku 4.5 90% to 97%), so this sample cannot separate them.","routing-jev-vs-llm",194,"routing-key-accuracy",194,194,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.9077,0.9718],[0.9013,0.968],"\u0001"],["Exact rate by decision type: Failure class",1,0.9444,"rate","100% (18/18)","94% (17/18)","tie","The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Haiku 4.5 74% to 99%), so this sample cannot separate them.","routing-jev-vs-llm",18,"routing-exact-by-decision",18,18,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8241,1],[0.7424,0.9901],"\u0001"],["Exact rate by decision type: Message intent",1,1,"rate","100% (20/20)","100% (20/20)","tie","The 95% intervals overlap (Jev 1.13 84% to 100%; Claude Haiku 4.5 84% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",20,"routing-exact-by-decision",20,20,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8389,1],[0.8389,1],"\u0001"],["Exact rate by decision type: Is it a rule?",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Jev 1.13 76% to 100%; Claude Haiku 4.5 76% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",12,"routing-exact-by-decision",12,12,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Exact rate by decision type: Context shape",0.7396,0.75,"rate","74%","75% (24/32)","tie","The 95% intervals overlap (Jev 1.13 58% to 87%; Claude Haiku 4.5 58% to 87%), so this sample cannot separate them.","routing-jev-vs-llm",32,"routing-exact-by-decision",32,32,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.5789,0.8675],[0.5789,0.8675],"\u0001"],["Cost per 1,000 routing decisions",0.0337,8.924,"usd","$0.034","$8.92","unclear","No interval or range was recorded for either side, so the gap ($0.034 vs $8.92, 265x) is not tested against run-to-run variation.","routing-jev-vs-llm","\u0001","routing-cost-per-1000",246,82,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Time to make one routing decision",136.5,12543,"ms","137 ms","12,543 ms","a","Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Haiku 4.5 12,543 ms to 34,481 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.","routing-overhead","\u0001","router-overhead-decision-latency",246,82,"routing overhead per decision · TypeSafe API","thinking on · via Claude Code · routing overhead per decision","range","p50-p95",[136.5,195.7],[12543,34481],"\u0001"],["Routing calls that returned a decision",1,1,"rate","100% (246/246)","100% (82/82)","tie","The 95% intervals overlap (Jev 1.13 98% to 100%; Claude Haiku 4.5 96% to 100%), so this sample cannot separate them.","routing-overhead","\u0001","router-overhead-completed",246,82,"routing overhead per decision · TypeSafe API","thinking on · via Claude Code · routing overhead per decision","ci95","ci95",[0.9846,1],[0.9552,1],"\u0001"],["Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))",1.67,441.74,"usd","$1.67","$441.74","unclear","No interval or range was recorded for either side, so the gap ($1.67 vs $441.74, 265x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))",0.24,62.47,"usd","$0.24","$62.47","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $62.47, 260x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Every model call routed (49.5 per task))",6.7568,620.8785,"seconds","6.76 s","620.9 s","unclear","No interval or range was recorded for either side, so the gap (6.76 s vs 620.9 s, 92x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Only System One decisions (7 per task))",0.9555,87.801,"seconds","0.96 s","87.8 s","unclear","No interval or range was recorded for either side, so the gap (0.96 s vs 87.8 s, 92x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Unseen routing decisions answered exactly right",0.8214,0.7857,"rate","82% (46/56)","79% (44/56)","tie","The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Haiku 4.5 66% to 87%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-exact",56,56,"","Claude Code","ci95","ci95",[0.7016,0.9],[0.6618,0.8729],"\u0001"],["Per-question accuracy on unseen decisions",0.904,0.816,"rate","90% (113/125)","82% (102/125)","tie","The 95% intervals overlap (Jev 1.13 84% to 94%; Claude Haiku 4.5 74% to 87%), so this sample cannot separate them.","routing-holdout",125,"routing-holdout-key-accuracy",125,125,"","Claude Code","ci95","ci95",[0.8397,0.9442],[0.739,0.8741],"\u0001"],["Exact rate on unseen decisions, by decision type: Failure class",0.9286,0.9286,"rate","93% (13/14)","93% (13/14)","tie","The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Haiku 4.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code","ci95","ci95",[0.6853,0.9873],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Message intent",0.8571,0.9286,"rate","86% (12/14)","93% (13/14)","tie","The 95% intervals overlap (Jev 1.13 60% to 96%; Claude Haiku 4.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code","ci95","ci95",[0.6006,0.9599],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Is it a rule?",0.9286,0.9286,"rate","93% (13/14)","93% (13/14)","tie","The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Haiku 4.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code","ci95","ci95",[0.6853,0.9873],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Context shape",0.5714,0.3571,"rate","57% (8/14)","36% (5/14)","tie","The 95% intervals overlap (Jev 1.13 33% to 79%; Claude Haiku 4.5 16% to 61%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code","ci95","ci95",[0.3259,0.7862],[0.1634,0.6124],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm))",0.9024,0.8902,"rate","90% (74/82)","89% (73/82)","tie","The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them.","routing-holdout",82,"routing-holdout-tuned-vs-unseen",82,82,"","Claude Code","ci95","ci95",[0.8191,0.9497],[0.8044,0.9412],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Unseen holdout)",0.8214,0.7857,"rate","82% (46/56)","79% (44/56)","tie","The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Haiku 4.5 66% to 87%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-tuned-vs-unseen",56,56,"","Claude Code","ci95","ci95",[0.7016,0.9],[0.6618,0.8729],"\u0001"],["Time per routing decision, by route (Wall time)",0.139,9.444,"seconds","0.14 s","9.44 s","a","Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Haiku 4.5 9.44 s to 25.4 s); not a confidence interval.","routing-holdout","\u0001","routing-holdout-latency",168,56,"","Claude Code","range","p50-p95",[0.139,0.192],[9.444,25.413],"\u0001"],["Cost per 1,000 unseen routing decisions",0.03065,7.129,"usd","$0.031","$7.13","unclear","No interval or range was recorded for either side, so the gap ($0.031 vs $7.13, 233x) is not tested against run-to-run variation.","routing-holdout","\u0001","routing-holdout-cost-per-1000",168,56,"","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["jev-1-13-vs-claude-sonnet-5-5","jev-1-13","claude-sonnet-5-5","Jev 1.13 vs Claude Sonnet 5.5","Jev 1.13 vs Claude Sonnet 5.5: measured benchmarks","Jev 1.13 vs Claude Sonnet 5.5: 17 measured metrics from 3 studies (Typed routing decisions answered exactly right; more), with sample sizes and intervals.","Jev 1.13 and Claude Sonnet 5.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 2,597 ms; Time per routing decision, by route (Wall time), 0.14 s vs 2.36 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Typed routing decisions answered exactly right",0.8984,0.939,"rate","90%","94% (77/82)","tie","The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","routing-jev-vs-llm",82,"routing-exact-decisions",82,82,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8191,0.9497],[0.8651,0.9737],"\u0001"],["Per-question accuracy",0.9485,0.9742,"rate","95% (184/194)","97% (189/194)","tie","The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them.","routing-jev-vs-llm",194,"routing-key-accuracy",194,194,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.9077,0.9718],[0.9411,0.9889],"\u0001"],["Exact rate by decision type: Failure class",1,1,"rate","100% (18/18)","100% (18/18)","tie","The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Sonnet 5.5 82% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",18,"routing-exact-by-decision",18,18,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8241,1],[0.8241,1],"\u0001"],["Exact rate by decision type: Message intent",1,1,"rate","100% (20/20)","100% (20/20)","tie","The 95% intervals overlap (Jev 1.13 84% to 100%; Claude Sonnet 5.5 84% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",20,"routing-exact-by-decision",20,20,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8389,1],[0.8389,1],"\u0001"],["Exact rate by decision type: Is it a rule?",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Jev 1.13 76% to 100%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",12,"routing-exact-by-decision",12,12,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Exact rate by decision type: Context shape",0.7396,0.8438,"rate","74%","84% (27/32)","tie","The 95% intervals overlap (Jev 1.13 58% to 87%; Claude Sonnet 5.5 68% to 93%), so this sample cannot separate them.","routing-jev-vs-llm",32,"routing-exact-by-decision",32,32,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.5789,0.8675],[0.6825,0.9314],"\u0001"],["Cost per 1,000 routing decisions",0.0337,4.996,"usd","$0.034","$5.00","unclear","No interval or range was recorded for either side, so the gap ($0.034 vs $5.00, 148x) is not tested against run-to-run variation.","routing-jev-vs-llm","\u0001","routing-cost-per-1000",246,82,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Time to make one routing decision",136.5,2597,"ms","137 ms","2,597 ms","a","Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.","routing-overhead","\u0001","router-overhead-decision-latency",246,82,"routing overhead per decision · TypeSafe API","effort low · via Claude Code · routing overhead per decision","range","p50-p95",[136.5,195.7],[2597,4298],"\u0001"],["Routing calls that returned a decision",1,1,"rate","100% (246/246)","100% (82/82)","tie","The 95% intervals overlap (Jev 1.13 98% to 100%; Claude Sonnet 5.5 96% to 100%), so this sample cannot separate them.","routing-overhead","\u0001","router-overhead-completed",246,82,"routing overhead per decision · TypeSafe API","effort low · via Claude Code · routing overhead per decision","ci95","ci95",[0.9846,1],[0.9552,1],"\u0001"],["Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))",1.67,247.3,"usd","$1.67","$247.30","unclear","No interval or range was recorded for either side, so the gap ($1.67 vs $247.30, 148x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))",0.24,34.97,"usd","$0.24","$34.97","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $34.97, 146x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Every model call routed (49.5 per task))",6.7568,128.5515,"seconds","6.76 s","128.6 s","unclear","No interval or range was recorded for either side, so the gap (6.76 s vs 128.6 s, 19x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Only System One decisions (7 per task))",0.9555,18.179,"seconds","0.96 s","18.2 s","unclear","No interval or range was recorded for either side, so the gap (0.96 s vs 18.2 s, 19x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Unseen routing decisions answered exactly right",0.8214,0.875,"rate","82% (46/56)","88% (49/56)","tie","The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-exact",56,56,"","Claude Code · effort low","ci95","ci95",[0.7016,0.9],[0.7637,0.9381],"\u0001"],["Per-question accuracy on unseen decisions",0.904,0.92,"rate","90% (113/125)","92% (115/125)","tie","The 95% intervals overlap (Jev 1.13 84% to 94%; Claude Sonnet 5.5 86% to 96%), so this sample cannot separate them.","routing-holdout",125,"routing-holdout-key-accuracy",125,125,"","Claude Code · effort low","ci95","ci95",[0.8397,0.9442],[0.859,0.956],"\u0001"],["Exact rate on unseen decisions, by decision type: Failure class",0.9286,1,"rate","93% (13/14)","100% (14/14)","tie","The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.7847,1],"\u0001"],["Exact rate on unseen decisions, by decision type: Message intent",0.8571,1,"rate","86% (12/14)","100% (14/14)","tie","The 95% intervals overlap (Jev 1.13 60% to 96%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code · effort low","ci95","ci95",[0.6006,0.9599],[0.7847,1],"\u0001"],["Exact rate on unseen decisions, by decision type: Is it a rule?",0.9286,0.9286,"rate","93% (13/14)","93% (13/14)","tie","The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Sonnet 5.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Context shape",0.5714,0.5714,"rate","57% (8/14)","57% (8/14)","tie","The 95% intervals overlap (Jev 1.13 33% to 79%; Claude Sonnet 5.5 33% to 79%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code · effort low","ci95","ci95",[0.3259,0.7862],[0.3259,0.7862],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm))",0.9024,0.939,"rate","90% (74/82)","94% (77/82)","tie","The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","routing-holdout",82,"routing-holdout-tuned-vs-unseen",82,82,"","Claude Code · effort low","ci95","ci95",[0.8191,0.9497],[0.8651,0.9737],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Unseen holdout)",0.8214,0.875,"rate","82% (46/56)","88% (49/56)","tie","The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-tuned-vs-unseen",56,56,"","Claude Code · effort low","ci95","ci95",[0.7016,0.9],[0.7637,0.9381],"\u0001"],["Time per routing decision, by route (Wall time)",0.139,2.359,"seconds","0.14 s","2.36 s","a","Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.","routing-holdout","\u0001","routing-holdout-latency",168,56,"","Claude Code · effort low","range","p50-p95",[0.139,0.192],[2.359,3.657],"\u0001"],["Cost per 1,000 unseen routing decisions",0.03065,7.244,"usd","$0.031","$7.24","unclear","No interval or range was recorded for either side, so the gap ($0.031 vs $7.24, 236x) is not tested against run-to-run variation.","routing-holdout","\u0001","routing-holdout-cost-per-1000",168,56,"","Claude Code · effort low","\u0001","\u0001","\u0001","\u0001",true]]}],["gpt-6-1-sol-codex-cli-vs-gpt-6-1-sol-openai-api","gpt-6-1-sol-codex-cli","gpt-6-1-sol-openai-api","GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API)","GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API)","GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API): 8 measured metrics from one study, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) share 8 measured metrics from one study. GPT-6.1 Sol (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 1.52 s vs 4.19 s; CLI vs API: time for a one-line answer (First useful output), 1.34 s vs 3.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 6 unclear; each row says why. Every row ran the two sides through different routes (for example Codex CLI vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",4.19,1.52,"seconds","4.19 s","1.52 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","OpenAI API · effort high · fixed exact reply, 5 runs","range","minmax",[3.81,4.69],[1.35,2.23]],["CLI vs API: time for a one-line answer (First useful output)",3.79,1.34,"seconds","3.79 s","1.34 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","OpenAI API · effort high · fixed exact reply, 5 runs","range","minmax",[3.37,4.3],[1.26,2.12]],["CLI vs API: time for a small coding task (Total time)",17.85,9.56,"seconds","17.9 s","9.56 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6.1 Sol (OpenAI API) 9.44 s to 10.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","OpenAI API · effort high · small coding task, 3 runs","range","minmax",[17.68,22.42],[9.44,10.94]],["CLI vs API: time for a small coding task (First useful output)",17.27,5.31,"seconds","17.3 s","5.31 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6.1 Sol (OpenAI API) 4.99 s to 6.42 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","OpenAI API · effort high · small coding task, 3 runs","range","minmax",[17.13,21.86],[4.99,6.42]],["Hidden prompt: input tokens for the same one-line request",19555,17,"tokens","19,555","17","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort high · short fixed tasks","OpenAI API · effort high · short fixed tasks","\u0001","\u0001","\u0001","\u0001"],["Repairing a scheduler: Claude Code vs Codex vs API (Total time)",61.16,17.32,"seconds","61.2 s","17.3 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 59.9 s to 69.5 s; GPT-6.1 Sol (OpenAI API) 16.3 s to 18.6 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs","OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[59.9,69.51],[16.28,18.61]],["Repairing a scheduler: Claude Code vs Codex vs API (First useful output)",15.56,7.46,"seconds","15.6 s","7.46 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 13.7 s to 23.0 s; GPT-6.1 Sol (OpenAI API) 6.68 s to 9.05 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs","OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[13.65,23.04],[6.68,9.05]],["Output tokens to repair the scheduler (Output tokens)",1181,1313,"tokens","1,181","1,313","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",3,"scheduler-repair-output-tokens",3,3,"Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs","OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-5-5-vs-gpt-6-1-sol-codex-cli","claude-opus-5-5","gpt-6-1-sol-codex-cli","Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)","Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): benchmarks","Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): 31 measured metrics from 6 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) share 31 measured metrics and 19 list-price calculations from 7 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 12 ties and 38 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.71,5.6,"seconds","2.71 s","5.60 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.45 s to 11.8 s; GPT-6.1 Sol (Codex CLI) 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[2.45,11.78],[4.05,19.52],"\u0001"],["Time to first useful output",2.04,5.32,"seconds","2.04 s","5.32 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.40 s to 9.94 s; GPT-6.1 Sol (Codex CLI) 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[1.4,9.94],[3.64,16.37],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1463,6716,"tokens","1,463","6,716","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",619,5406,"tokens","619","5,406","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",78,42,"tokens","78","42","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00694,0.01047,"usd","$0.0069","$0.010","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 $0.0059 to $0.027; GPT-6.1 Sol (Codex CLI) $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[0.00592,0.02708],[0.0066,0.02812],true],["List-price cost per passing answer (calculation)",0.01049,0.01322,"usd","$0.010","$0.013","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.013) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",11.03,18.12,"seconds","11.0 s","18.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.63 s to 63.0 s; GPT-6.1 Sol (Codex CLI) 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","range","minmax",[3.63,63],[11.67,92.21],"\u0001"],["Time to first useful output on hard tasks",7.13,12.69,"seconds","7.13 s","12.7 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.15 s to 56.2 s; GPT-6.1 Sol (Codex CLI) 8.93 s to 75.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-first-useful-latency",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","range","minmax",[2.15,56.23],[8.93,75.91],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1052,436,"tokens","1,052","436","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head","\u0001","hard-h2h-output-tokens",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.03337,0.01514,"usd","$0.033","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.033 vs $0.015, 2.2x) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Opus 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",56.9,113.4,"seconds","56.9 s","113.4 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 29.8 s to 185.8 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[29.8,185.8],[78.5,221.9],"\u0001"],["Tool calls per coding session",7.5,12.5,"calls","7.5","12.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","coding-agents-head-to-head",12,"coding-agents-tool-calls",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[5,14],[8,18],"\u0001"],["List-price cost per passing coding session (calculation)",0.2229,0.0978,"usd","$0.22","$0.098","unclear","No interval or range was recorded for either side, so the gap ($0.22 vs $0.098, 2.3x) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",9.72,13.11,"seconds","9.72 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 4.78 s to 31.4 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","range","minmax",[4.78,31.36],[8.54,61.6],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",853,335,"tokens","853","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02947,0.02564,"usd","$0.029","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.026) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",54.43,57.01,"percent","54.4%","57%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-share",24,16,"Claude Code · effort high","Codex CLI · effort high","range","minmax",[36.14,96.23],[29.19,90.8],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.017969,0.004114,"usd","$0.018","$0.0041","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code · effort high","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.00799,0.002905,"usd","$0.0080","$0.0029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code · effort high","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.007407,0.008117,"usd","$0.0074","$0.0081","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code · effort high","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.01344,0.002273,"usd","$0.013","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.029475,0.025637,"usd","$0.029","$0.026","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.43,57.01,"percent","54.4%","57%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-short-vs-hard",24,16,"Claude Code · effort high","Codex CLI · effort high","range","minmax",[36.14,96.23],[29.19,90.8],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",43.59,58.06,"percent","43.6%","58.1%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code · effort high","Codex CLI · effort high","range","minmax",[0,93.33],[0,75.76],true],["Time to first text: a 250-line answer, six models",1.97,3.52,"seconds","1.97 s","3.52 s","unclear","Only 4 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.35 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s), but 4 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[1.7,2.35],[2.75,4.42],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",155.5,79.6,"tokens","156","80","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[154.6,156.4],[71.6,80.5],true],["Output speed in characters per second after the first text (calculation)",347,323,"count","347","323","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[345,349],[291,327],true],["Time to first text as the prompt grows: 1k",1.51,3.36,"seconds","1.51 s","3.36 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.46 s to 2.01 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.46,2.01],[3.36,4.75],true],["Time to first text as the prompt grows: 16k",1.74,4.02,"seconds","1.74 s","4.02 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.97 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.7,2.97],[3.3,4.28],true],["Time to first text as the prompt grows: 64k",1.79,3.93,"seconds","1.79 s","3.93 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.72 s to 3.72 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.72,3.72],[3.42,4.38],true],["Total time per call by prompt size (1k prompt)",1.83,3.43,"seconds","1.83 s","3.43 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.82 s to 2.41 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.82,2.41],[3.43,4.92],"\u0001"],["Total time per call by prompt size (16k prompt)",2.36,4.14,"seconds","2.36 s","4.14 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 2.11 s to 3.40 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.11,3.4],[3.96,4.68],"\u0001"],["Total time per call by prompt size (64k prompt)",2.35,3.96,"seconds","2.35 s","3.96 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.26 s to 4.29 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.26,4.29],[3.47,4.44],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",0.5556,1,"rate","56% (5/9)","100% (9/9)","tie","The 95% intervals overlap (Claude Opus 5.5 27% to 81%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Codex CLI · effort low","ci95","ci95",[0.2667,0.8112],[0.7009,1],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0.4167,0.6875,"rate","42% (5/12)","69% (11/16)","tie","The 95% intervals overlap (Claude Opus 5.5 19% to 68%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.1933,0.6805],[0.444,0.8584],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0.5,0.6875,"rate","50% (6/12)","69% (11/16)","tie","The 95% intervals overlap (Claude Opus 5.5 25% to 75%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.2538,0.7462],[0.444,0.8584],"\u0001"],["Calls that tried a tool although tools were off",0.4167,0,"rate","42% (5/12)","0% (0/16)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-tool-attempts",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.1933,0.6805],[0,0.1936],"\u0001"],["Strict pass rate by task: 10x10 nonogram",1,0.75,"rate","100% (3/3)","75% (3/4)","tie","The 95% intervals overlap (Claude Opus 5.5 44% to 100%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.4385,1],[0.3006,0.9544],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0.25,"rate","0% (0/3)","25% (1/4)","tie","The 95% intervals overlap (Claude Opus 5.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.0456,0.6994],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0.3333,1,"rate","33% (1/3)","100% (4/4)","tie","The 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 51% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.0615,0.7923],[0.5101,1],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0.3333,0.75,"rate","33% (1/3)","75% (3/4)","tie","The 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.0615,0.7923],[0.3006,0.9544],"\u0001"],["Total time per call on harder tasks",80.34,120.24,"seconds","80.3 s","120.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.82 s to 279.5 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",9,13,"Claude Code","Codex CLI · effort medium","range","minmax",[3.82,279.5],[46.24,273.46],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",8420,4994,"tokens","8,420","4,994","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",9,13,"Claude Code","Codex CLI · effort medium","range","minmax",[279,40044],[2099,13413],"\u0001"],["List-price cost per strict pass on harder tasks (calculation)",0.59333,0.08293,"usd","$0.59","$0.083","unclear","No interval or range was recorded for either side, so the gap ($0.59 vs $0.083, 7.2x) is not tested against run-to-run variation.","harder-tasks-head-to-head","\u0001","harder-h2h-cost-per-pass",12,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-haiku-4-5-vs-gpt-6-1-sol-codex-cli","claude-haiku-4-5","gpt-6-1-sol-codex-cli","Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)","Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI): benchmarks","Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI): 40 measured metrics from 6 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) share 40 measured metrics and 16 list-price calculations from 7 studies. Claude Haiku 4.5 leads on 2 rows: Same prompt, 10 times: time per call (Exact number), 5.06 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 5.95 s vs 11.3 s. GPT-6.1 Sol (Codex CLI) leads on 6 rows: Pass rate on eight hard tasks (Strict pass), 100% (16/16) vs 46% (11/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); Same prompt, 10 times: strict pass rate (JSON object), 100% (10/10) vs 10% (1/10); and 3 more. On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 13 ties and 35 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",4.43,5.65,"seconds","4.43 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[3.16,23.57],[4.1,25.46],"\u0001"],["Time to first useful output",3.63,5.05,"seconds","3.63 s","5.05 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[2.78,22.27],[3.36,17.82],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",0,5180,"tokens","0","5,180","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",3790,6943,"tokens","3,790","6,943","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",367,42,"tokens","367","42","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00566,0.01018,"usd","$0.0057","$0.010","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.00513,0.01804],[0.0054,0.02686],true],["List-price cost per passing answer (calculation)",0.00836,0.01564,"usd","$0.0084","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.016) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (16/16)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; GPT-6.1 Sol (Codex CLI) 81% to 100%).","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Haiku 4.5 47% to 82%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,13.11,"seconds","39.0 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[15.27,75.13],[8.54,61.6],"\u0001"],["Time to first useful output on hard tasks",35.54,10.23,"seconds","35.5 s","10.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-first-useful-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[12.88,70.31],[6.09,40.41],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",5064,335,"tokens","5,064","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head","\u0001","hard-h2h-output-tokens",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.02564,"usd","$0.067","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.026, 2.6x) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",0,1,"rate","0% (0/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; GPT-6.1 Sol (Codex CLI) 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0,0.2775],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",0.1,1,"rate","10% (1/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; GPT-6.1 Sol (Codex CLI) 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.0179,0.4042],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: how many different answers (Exact number)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (JSON object)",1,1,"count","1","1","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: how many different answers (Code fix)",6,6,"count","6","6","tie","Same value. More or fewer is not better by itself for this metric.","caching-consistency",10,"consistency-distinct-answers",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001","\u0001","\u0001","\u0001"],["Same prompt, 10 times: time per call (Exact number)",5.06,13.38,"seconds","5.06 s","13.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.42 s to 6.20 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[4.42,6.2],[12.29,17.97],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",7.03,6.42,"seconds","7.03 s","6.42 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.28 s to 12.3 s; GPT-6.1 Sol (Codex CLI) 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[5.28,12.27],[5.25,8.26],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",5.95,11.29,"seconds","5.95 s","11.3 s","a","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[4.89,7.33],[9.08,14.85],"\u0001"],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)",0,1,"rate","0% (0/24)","100% (12/12)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; GPT-6.1 Sol (Codex CLI) 76% to 100%).","json-schema-vs-instructions","\u0001","structured-output-pass-rate",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","ci95","ci95",[0,0.138],[0.7575,1],true],["Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))",0.7083,1,"rate","71% (17/24)","100% (12/12)","tie","The 95% intervals overlap (Claude Haiku 4.5 51% to 85%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.","json-schema-vs-instructions","\u0001","structured-output-pass-rate",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","ci95","ci95",[0.5083,0.8509],[0.7575,1],true],["What each call produced: strict pass, format miss, wrong values or error (Strict pass)",0,12,"count","0","12","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Format miss)",17,0,"count","17","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Wrong values)",7,0,"count","7","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["What each call produced: strict pass, format miss, wrong values or error (Error)",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","json-schema-vs-instructions","\u0001","structured-output-outcomes",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time per call, instructions vs schema mode",9.52,6.21,"seconds","9.52 s","6.21 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.67 s to 17.0 s; GPT-6.1 Sol (Codex CLI) 4.20 s to 12.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","json-schema-vs-instructions","\u0001","structured-output-time",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","range","minmax",[5.67,17],[4.2,12.27],"\u0001"],["Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)",1128,117,"tokens","1,128","117","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","json-schema-vs-instructions","\u0001","structured-output-tokens",24,12,"Claude Code · instructions","Codex CLI · effort low · instructions","\u0001","\u0001","\u0001","\u0001","\u0001"],["Reasoning share of output tokens per call on hard tasks (calculation)",91.68,46.33,"percent","91.7%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-share",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[76.46,99.27],[11.42,86.85],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.024492,0.002273,"usd","$0.024","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.001795,0.003025,"usd","$0.0018","$0.0030","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.00451,0.020339,"usd","$0.0045","$0.020","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",91.68,46.33,"percent","91.7%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-short-vs-hard",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[76.46,99.27],[11.42,86.85],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",90.19,41.05,"percent","90.2%","41%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Codex CLI · effort medium","range","minmax",[73.1,97.59],[0,71.43],true],["Time to first text: a 250-line answer, six models",4,3.52,"seconds","4.00 s","3.52 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[2.84,6.38],[2.75,4.42],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",153.2,79.6,"tokens","153","80","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[152.6,216.1],[71.6,80.5],true],["Output speed in characters per second after the first text (calculation)",547,323,"count","547","323","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","\u0001","speed-anatomy-chars-per-second",3,4,"Claude Code","Codex CLI · effort low","range","minmax",[546,548],[291,327],true],["Time to first text as the prompt grows: 1k",1.93,3.36,"seconds","1.93 s","3.36 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 1.85 s to 2.04 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[1.85,2.04],[3.36,4.75],true],["Time to first text as the prompt grows: 16k",2.27,4.02,"seconds","2.27 s","4.02 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.47 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.22,2.47],[3.3,4.28],true],["Time to first text as the prompt grows: 64k",2.78,3.93,"seconds","2.78 s","3.93 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.45 s to 2.89 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-prompt-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.45,2.89],[3.42,4.38],true],["Total time per call by prompt size (1k prompt)",2.34,3.43,"seconds","2.34 s","3.43 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.46 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.22,2.46],[3.43,4.92],"\u0001"],["Total time per call by prompt size (16k prompt)",2.79,4.14,"seconds","2.79 s","4.14 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.58 s to 2.84 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.58,2.84],[3.96,4.68],"\u0001"],["Total time per call by prompt size (64k prompt)",3.13,3.96,"seconds","3.13 s","3.96 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.84 s to 3.28 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",3,"speed-anatomy-total-by-size",3,3,"Claude Code","Codex CLI · effort low","range","minmax",[2.84,3.28],[3.47,4.44],"\u0001"],["Exact lookup answers at the 1k, 16k and 64k prompt-size targets",1,1,"rate","100% (9/9)","100% (9/9)","tie","The 95% intervals overlap (Claude Haiku 4.5 70% to 100%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them.","llm-speed-anatomy",9,"speed-anatomy-lookup-correct",9,9,"Claude Code","Codex CLI · effort low","ci95","ci95",[0.7009,1],[0.7009,1],"\u0001"],["Pass rate on 4 harder tasks (Strict pass)",0,0.6875,"rate","0% (0/12)","69% (11/16)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%).","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.2425],[0.444,0.8584],"\u0001"],["Pass rate on 4 harder tasks (Lenient (format misses counted))",0,0.6875,"rate","0% (0/12)","69% (11/16)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%).","harder-tasks-head-to-head","\u0001","harder-h2h-pass-rate",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.2425],[0.444,0.8584],"\u0001"],["Calls that tried a tool although tools were off",0.0833,0,"rate","8% (1/12)","0% (0/16)","unclear","More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-tool-attempts",12,16,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0.0149,0.3539],[0,0.1936],"\u0001"],["Strict pass rate by task: 10x10 nonogram",0,0.75,"rate","0% (0/3)","75% (3/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.3006,0.9544],"\u0001"],["Strict pass rate by task: Sudoku, 22 givens",0,0.25,"rate","0% (0/3)","25% (1/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.0456,0.6994],"\u0001"],["Strict pass rate by task: 6x6 Skyscrapers",0,1,"rate","0% (0/3)","100% (4/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 51% to 100%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.5101,1],"\u0001"],["Strict pass rate by task: Seeded shuffle output",0,0.75,"rate","0% (0/3)","75% (3/4)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.","harder-tasks-head-to-head","\u0001","harder-h2h-pass-by-task",3,4,"Claude Code","Codex CLI · effort medium","ci95","ci95",[0,0.5615],[0.3006,0.9544],"\u0001"],["Total time per call on harder tasks",108.98,120.24,"seconds","109.0 s","120.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 25.7 s to 223.9 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","harder-tasks-head-to-head","\u0001","harder-h2h-total-latency",10,13,"Claude Code","Codex CLI · effort medium","range","minmax",[25.73,223.95],[46.24,273.46],"\u0001"],["Output tokens per call on harder tasks (Output tokens)",12508,4994,"tokens","12,508","4,994","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","harder-tasks-head-to-head","\u0001","harder-h2h-output-tokens",10,13,"Claude Code","Codex CLI · effort medium","range","minmax",[2965,26532],[2099,13413],"\u0001"]]}],["claude-fable-5-1-vs-gpt-6-1-sol-codex-cli","claude-fable-5-1","gpt-6-1-sol-codex-cli","Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI)","Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI): benchmarks","Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI): 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 4 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Fable 5.1 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",1.94,5.65,"seconds","1.94 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 1.41 s to 9.83 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[1.41,9.83],[4.1,25.46],"\u0001"],["Time to first useful output",1.2,5.05,"seconds","1.20 s","5.05 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 0.95 s to 7.90 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.95,7.9],[3.36,17.82],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",2760,5180,"tokens","2,760","5,180","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",473,6943,"tokens","473","6,943","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",64,42,"tokens","64","42","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00987,0.01018,"usd","$0.0099","$0.010","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 $0.0049 to $0.058; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.0049,0.05843],[0.0054,0.02686],true],["List-price cost per passing answer (calculation)",0.02054,0.01564,"usd","$0.021","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.021 vs $0.016) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",16.13,13.11,"seconds","16.1 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 4.46 s to 90.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[4.46,90],[8.54,61.6],"\u0001"],["Time to first useful output on hard tasks",11.63,10.23,"seconds","11.6 s","10.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.00 s to 85.3 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-first-useful-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[2,85.33],[6.09,40.41],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1366,335,"tokens","1,366","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head","\u0001","hard-h2h-output-tokens",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.09331,0.02564,"usd","$0.093","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.093 vs $0.026, 3.6x) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",64.24,46.33,"percent","64.2%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-share",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[23.44,97.19],[11.42,86.85],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.053696,0.002273,"usd","$0.054","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.018702,0.003025,"usd","$0.019","$0.0030","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.02091,0.020339,"usd","$0.021","$0.020","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",64.24,46.33,"percent","64.2%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-short-vs-hard",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[23.44,97.19],[11.42,86.85],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",0,41.05,"percent","0%","41%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Codex CLI · effort medium","range","minmax",[0,74.01],[0,71.43],true],["Time to first text: a 250-line answer, six models",4.43,3.52,"seconds","4.43 s","3.52 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.27 s to 4.64 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[2.27,4.64],[2.75,4.42],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",122.6,79.6,"tokens","123","80","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[120.9,131.4],[71.6,80.5],true],["Output speed in characters per second after the first text (calculation)",273,323,"count","273","323","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[270,293],[291,327],true]]}],["jev-1-13-vs-deterministic-routing-policy","jev-1-13","deterministic-routing-policy","Jev 1.13 vs Deterministic routing policy","Jev 1.13 vs Deterministic routing policy: benchmarks","Jev 1.13 vs Deterministic routing policy: 3 measured metrics from one study (Time to make one routing decision; more), with sample sizes and intervals.","Jev 1.13 and Deterministic routing policy share 3 measured metrics and 4 list-price calculations from one study. Deterministic routing policy leads on 1 row: Time to make one routing decision, 1.42 µs vs 137 ms. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Time to make one routing decision",136.5,0.00142,"ms","137 ms","1.42 µs","b","Jev 1.13’s median is above Deterministic routing policy’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Deterministic routing policy 1.42 µs to 2.33 µs); not a confidence interval.","routing-overhead","router-overhead-decision-latency",246,20000,"routing overhead per decision · TypeSafe API","Agent · in process · routing overhead per decision","range","p50-p95",[136.5,195.7],[0.00142,0.00233],"\u0001"],["Routing calls that returned a decision",1,1,"rate","100% (246/246)","100% (20000/20000)","tie","The 95% intervals overlap (Jev 1.13 98% to 100%; Deterministic routing policy 100% to 100%), so this sample cannot separate them.","routing-overhead","router-overhead-completed",246,20000,"routing overhead per decision · TypeSafe API","Agent · in process · routing overhead per decision","ci95","ci95",[0.9846,1],[0.9998,1],"\u0001"],["Cost per 1,000 routing decisions: no model call vs provider-reported",0.0337,0,"usd","$0.034","$0.00","unclear","No interval or range was recorded for either side, so the gap ($0.034 vs $0.00) is not tested against run-to-run variation.","routing-overhead","router-overhead-cost-reported",82,20000,"routing overhead per decision · TypeSafe API","Agent · in process · routing overhead per decision","\u0001","\u0001","\u0001","\u0001","\u0001"],["Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))",1.67,0,"usd","$1.67","$0.00","unclear","No interval or range was recorded for either side, so the gap ($1.67 vs $0.00) is not tested against run-to-run variation.","routing-overhead","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","Agent · in process · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))",0.24,0,"usd","$0.24","$0.00","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $0.00) is not tested against run-to-run variation.","routing-overhead","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","Agent · in process · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Every model call routed (49.5 per task))",6.7568,0.0000703,"seconds","6.76 s","70.3 µs","unclear","No interval or range was recorded for either side, so the gap (6.76 s vs 70.3 µs, 96,114x) is not tested against run-to-run variation.","routing-overhead","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","Agent · in process · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Only System One decisions (7 per task))",0.9555,0.0000099,"seconds","0.96 s","9.9 µs","unclear","No interval or range was recorded for either side, so the gap (0.96 s vs 9.9 µs, 96,515x) is not tested against run-to-run variation.","routing-overhead","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","Agent · in process · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true]]}],["deterministic-routing-policy-vs-claude-sonnet-5-5","deterministic-routing-policy","claude-sonnet-5-5","Deterministic routing policy vs Claude Sonnet 5.5","Deterministic routing policy vs Claude Sonnet 5.5","Deterministic routing policy vs Claude Sonnet 5.5: 2 measured metrics from one study (Time to make one routing decision; more), with sample sizes and intervals.","Deterministic routing policy and Claude Sonnet 5.5 share 2 measured metrics and 4 list-price calculations from one study. Deterministic routing policy leads on 1 row: Time to make one routing decision, 1.42 µs vs 2,597 ms. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 1 tie and 4 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Time to make one routing decision",0.00142,2597,"ms","1.42 µs","2,597 ms","a","Claude Sonnet 5.5’s median is above Deterministic routing policy’s 95th percentile (p50–p95 bands: Deterministic routing policy 1.42 µs to 2.33 µs; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval.","routing-overhead","router-overhead-decision-latency",20000,82,"Agent · in process · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","range","p50-p95",[0.00142,0.00233],[2597,4298],"\u0001"],["Routing calls that returned a decision",1,1,"rate","100% (20000/20000)","100% (82/82)","tie","The 95% intervals overlap (Deterministic routing policy 100% to 100%; Claude Sonnet 5.5 96% to 100%), so this sample cannot separate them.","routing-overhead","router-overhead-completed",20000,82,"Agent · in process · routing overhead per decision","effort low · via Claude Code · routing overhead per decision","ci95","ci95",[0.9998,1],[0.9552,1],"\u0001"],["Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))",0,247.3,"usd","$0.00","$247.30","unclear","No interval or range was recorded for either side, so the gap ($0.00 vs $247.30) is not tested against run-to-run variation.","routing-overhead","router-overhead-cost-per-1000-tasks","\u0001","\u0001","Agent · in process · calculation per 1,000 tasks from recorded decision counts","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))",0,34.97,"usd","$0.00","$34.97","unclear","No interval or range was recorded for either side, so the gap ($0.00 vs $34.97) is not tested against run-to-run variation.","routing-overhead","router-overhead-cost-per-1000-tasks","\u0001","\u0001","Agent · in process · calculation per 1,000 tasks from recorded decision counts","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Every model call routed (49.5 per task))",0.0000703,128.5515,"seconds","70.3 µs","128.6 s","unclear","No interval or range was recorded for either side, so the gap (70.3 µs vs 128.6 s, 1.8 million times) is not tested against run-to-run variation.","routing-overhead","router-overhead-delay-per-task","\u0001","\u0001","Agent · in process · calculation per task from recorded decision counts, decisions in line","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Only System One decisions (7 per task))",0.0000099,18.179,"seconds","9.9 µs","18.2 s","unclear","No interval or range was recorded for either side, so the gap (9.9 µs vs 18.2 s, 1.8 million times) is not tested against run-to-run variation.","routing-overhead","router-overhead-delay-per-task","\u0001","\u0001","Agent · in process · calculation per task from recorded decision counts, decisions in line","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true]]}],["deterministic-routing-policy-vs-claude-haiku-4-5","deterministic-routing-policy","claude-haiku-4-5","Deterministic routing policy vs Claude Haiku 4.5","Deterministic routing policy vs Claude Haiku 4.5: benchmarks","Deterministic routing policy vs Claude Haiku 4.5: 2 measured metrics from one study (Time to make one routing decision; more), with sample sizes and intervals.","Deterministic routing policy and Claude Haiku 4.5 share 2 measured metrics and 4 list-price calculations from one study. Deterministic routing policy leads on 1 row: Time to make one routing decision, 1.42 µs vs 12,543 ms. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 1 tie and 4 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Time to make one routing decision",0.00142,12543,"ms","1.42 µs","12,543 ms","a","Claude Haiku 4.5’s median is above Deterministic routing policy’s 95th percentile (p50–p95 bands: Deterministic routing policy 1.42 µs to 2.33 µs; Claude Haiku 4.5 12,543 ms to 34,481 ms); not a confidence interval.","routing-overhead","router-overhead-decision-latency",20000,82,"Agent · in process · routing overhead per decision","thinking on · via Claude Code · routing overhead per decision","range","p50-p95",[0.00142,0.00233],[12543,34481],"\u0001"],["Routing calls that returned a decision",1,1,"rate","100% (20000/20000)","100% (82/82)","tie","The 95% intervals overlap (Deterministic routing policy 100% to 100%; Claude Haiku 4.5 96% to 100%), so this sample cannot separate them.","routing-overhead","router-overhead-completed",20000,82,"Agent · in process · routing overhead per decision","thinking on · via Claude Code · routing overhead per decision","ci95","ci95",[0.9998,1],[0.9552,1],"\u0001"],["Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))",0,441.74,"usd","$0.00","$441.74","unclear","No interval or range was recorded for either side, so the gap ($0.00 vs $441.74) is not tested against run-to-run variation.","routing-overhead","router-overhead-cost-per-1000-tasks","\u0001","\u0001","Agent · in process · calculation per 1,000 tasks from recorded decision counts","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))",0,62.47,"usd","$0.00","$62.47","unclear","No interval or range was recorded for either side, so the gap ($0.00 vs $62.47) is not tested against run-to-run variation.","routing-overhead","router-overhead-cost-per-1000-tasks","\u0001","\u0001","Agent · in process · calculation per 1,000 tasks from recorded decision counts","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Every model call routed (49.5 per task))",0.0000703,620.8785,"seconds","70.3 µs","620.9 s","unclear","No interval or range was recorded for either side, so the gap (70.3 µs vs 620.9 s, 8.8 million times) is not tested against run-to-run variation.","routing-overhead","router-overhead-delay-per-task","\u0001","\u0001","Agent · in process · calculation per task from recorded decision counts, decisions in line","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Only System One decisions (7 per task))",0.0000099,87.801,"seconds","9.9 µs","87.8 s","unclear","No interval or range was recorded for either side, so the gap (9.9 µs vs 87.8 s, 8.9 million times) is not tested against run-to-run variation.","routing-overhead","router-overhead-delay-per-task","\u0001","\u0001","Agent · in process · calculation per task from recorded decision counts, decisions in line","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true]]}],["openrouter-vs-anthropic","openrouter","anthropic","OpenRouter vs Anthropic","OpenRouter vs Anthropic: price per model","OpenRouter vs Anthropic: 14 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","OpenRouter and Anthropic share 14 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 14 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Haiku 4.5: OpenRouter vs Anthropic list price (Input)",1,1,"usd","$1.00","$1.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-haiku-4-5","Claude Haiku 4.5 · list price, snapshot 2026-10-06","first-party list price · Claude Haiku 4.5 · list price, snapshot 2026-10-06"],["Claude Haiku 4.5: OpenRouter vs Anthropic list price (Output)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-haiku-4-5","Claude Haiku 4.5 · list price, snapshot 2026-10-06","first-party list price · Claude Haiku 4.5 · list price, snapshot 2026-10-06"],["Claude Sonnet 5: OpenRouter vs Anthropic list price (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-sonnet-5","Claude Sonnet 5 · list price, snapshot 2026-10-06","first-party list price · Claude Sonnet 5 · list price, snapshot 2026-10-06"],["Claude Sonnet 5: OpenRouter vs Anthropic list price (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-sonnet-5","Claude Sonnet 5 · list price, snapshot 2026-10-06","first-party list price · Claude Sonnet 5 · list price, snapshot 2026-10-06"],["Claude Sonnet 5.5: OpenRouter vs Anthropic list price (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-sonnet-5-5","Claude Sonnet 5.5 · list price, snapshot 2026-10-06","first-party list price · Claude Sonnet 5.5 · list price, snapshot 2026-10-06"],["Claude Sonnet 5.5: OpenRouter vs Anthropic list price (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-sonnet-5-5","Claude Sonnet 5.5 · list price, snapshot 2026-10-06","first-party list price · Claude Sonnet 5.5 · list price, snapshot 2026-10-06"],["Claude Opus 4.8: OpenRouter vs Anthropic list price (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-opus-4-8","Claude Opus 4.8 · list price, snapshot 2026-10-06","first-party list price · Claude Opus 4.8 · list price, snapshot 2026-10-06"],["Claude Opus 4.8: OpenRouter vs Anthropic list price (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-opus-4-8","Claude Opus 4.8 · list price, snapshot 2026-10-06","first-party list price · Claude Opus 4.8 · list price, snapshot 2026-10-06"],["Claude Opus 5: OpenRouter vs Anthropic list price (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-opus-5","Claude Opus 5 · list price, snapshot 2026-10-06","first-party list price · Claude Opus 5 · list price, snapshot 2026-10-06"],["Claude Opus 5: OpenRouter vs Anthropic list price (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-opus-5","Claude Opus 5 · list price, snapshot 2026-10-06","first-party list price · Claude Opus 5 · list price, snapshot 2026-10-06"],["Claude Opus 5.5: OpenRouter vs Anthropic list price (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-opus-5-5","Claude Opus 5.5 · list price, snapshot 2026-10-06","first-party list price · Claude Opus 5.5 · list price, snapshot 2026-10-06"],["Claude Opus 5.5: OpenRouter vs Anthropic list price (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-opus-5-5","Claude Opus 5.5 · list price, snapshot 2026-10-06","first-party list price · Claude Opus 5.5 · list price, snapshot 2026-10-06"],["Claude Fable 5.1: OpenRouter vs Anthropic list price (Input)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-fable-5-1","Claude Fable 5.1 · list price, snapshot 2026-10-06","first-party list price · Claude Fable 5.1 · list price, snapshot 2026-10-06"],["Claude Fable 5.1: OpenRouter vs Anthropic list price (Output)",50,50,"usd","$50.00","$50.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-claude-fable-5-1","Claude Fable 5.1 · list price, snapshot 2026-10-06","first-party list price · Claude Fable 5.1 · list price, snapshot 2026-10-06"]]}],["openrouter-vs-google-ai-studio","openrouter","google-ai-studio","OpenRouter vs Google AI Studio","OpenRouter vs Google AI Studio: price per model","OpenRouter vs Google AI Studio: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","OpenRouter and Google AI Studio share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Gemini 3.8 Flash: OpenRouter vs Google list price (Input)",0.75,0.75,"usd","$0.75","$0.75","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-gemini-3-8-flash","Gemini 3.8 Flash · list price, snapshot 2026-10-06","first-party list price · Gemini 3.8 Flash · list price, snapshot 2026-10-06"],["Gemini 3.8 Flash: OpenRouter vs Google list price (Output)",3.75,3.75,"usd","$3.75","$3.75","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-gemini-3-8-flash","Gemini 3.8 Flash · list price, snapshot 2026-10-06","first-party list price · Gemini 3.8 Flash · list price, snapshot 2026-10-06"],["Gemini 3.5 Flash: OpenRouter vs Google list price (Input)",1.5,1.5,"usd","$1.50","$1.50","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-gemini-3-5-flash","Gemini 3.5 Flash · list price, snapshot 2026-10-06","first-party list price · Gemini 3.5 Flash · list price, snapshot 2026-10-06"],["Gemini 3.5 Flash: OpenRouter vs Google list price (Output)",9,9,"usd","$9.00","$9.00","tie","Same reported list price.","inference-provider-index","gateway-vs-direct-gemini-3-5-flash","Gemini 3.5 Flash · list price, snapshot 2026-10-06","first-party list price · Gemini 3.5 Flash · list price, snapshot 2026-10-06"]]}],["openrouter-vs-openai","openrouter","openai","OpenRouter vs OpenAI","OpenRouter vs OpenAI: price per model","OpenRouter vs OpenAI: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","OpenRouter and OpenAI share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties; each row says why.",[{"metric":"GPT-6 Luna: OpenRouter vs OpenAI list price (Input)","aValue":0.1,"bValue":0.1,"unit":"usd","aDisplay":"$0.10","bDisplay":"$0.10","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"gateway-vs-direct-gpt-6-luna","aContext":"GPT-6 Luna · list price, snapshot 2026-10-06","bContext":"first-party list price · GPT-6 Luna · list price, snapshot 2026-10-06"},{"metric":"GPT-6 Luna: OpenRouter vs OpenAI list price (Output)","aValue":0.5,"bValue":0.5,"unit":"usd","aDisplay":"$0.50","bDisplay":"$0.50","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"gateway-vs-direct-gpt-6-luna","aContext":"GPT-6 Luna · list price, snapshot 2026-10-06","bContext":"first-party list price · GPT-6 Luna · list price, snapshot 2026-10-06"}]],["groq-vs-together","groq","together","Groq vs Together AI","Groq vs Together AI: price per model","Groq vs Together AI: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and Together AI share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.15,"usd","$0.15","$0.15","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.6,"usd","$0.60","$0.60","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.59,1.04,"usd","$0.59","$1.04","unclear","A reported list price has no interval, so the gap ($0.59 vs $1.04, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.79,1.04,"usd","$0.79","$1.04","unclear","A reported list price has no interval, so the gap ($0.79 vs $1.04, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["groq-vs-cerebras","groq","cerebras","Groq vs Cerebras","Groq vs Cerebras: price per model","Groq vs Cerebras: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and Cerebras share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.35,"usd","$0.15","$0.35","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.35, 2.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.75,"usd","$0.60","$0.75","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.75, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.075,0.35,"usd","$0.075","$0.35","unclear","A reported list price has no interval, so the gap ($0.075 vs $0.35, 4.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["together-vs-fireworks","together","fireworks","Together AI vs Fireworks AI","Together AI vs Fireworks AI: price per model","Together AI vs Fireworks AI: 6 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and Fireworks AI share 6 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 ties and 3 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Kimi K3: price per million tokens by provider (Input)",2.7,3,"usd","$2.70","$3.00","unclear","A reported list price has no interval, so the gap ($2.70 vs $3.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",13.5,15,"usd","$13.50","$15.00","unclear","A reported list price has no interval, so the gap ($13.50 vs $15.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.27,0.3,"usd","$0.27","$0.30","unclear","A reported list price has no interval, so the gap ($0.27 vs $0.30, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.26,"usd","$0.26","$0.26","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["amazon-bedrock-vs-google-vertex","amazon-bedrock","google-vertex","Amazon Bedrock vs Google Vertex AI","Amazon Bedrock vs Google Vertex AI: price per model","Amazon Bedrock vs Google Vertex AI: 23 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Google Vertex AI share 23 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 21 ties and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Haiku 4.5: price per million tokens by provider (Input)",1,1,"usd","$1.00","$1.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Output)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Cache read)",0.1,0.1,"usd","$0.10","$0.10","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Input)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Output)",50,50,"usd","$50.00","$50.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Cache read)",0.25,0.25,"usd","$0.25","$0.25","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.09,"usd","$0.15","$0.090","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.090, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.36,"usd","$0.60","$0.36","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.36, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["amazon-bedrock-vs-azure","amazon-bedrock","azure","Amazon Bedrock vs Azure","Amazon Bedrock vs Azure: price per model","Amazon Bedrock vs Azure: 21 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Azure share 21 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 21 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Haiku 4.5: price per million tokens by provider (Input)",1,1,"usd","$1.00","$1.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Output)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Cache read)",0.1,0.1,"usd","$0.10","$0.10","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Input)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Output)",50,50,"usd","$50.00","$50.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Cache read)",0.25,0.25,"usd","$0.25","$0.25","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["anthropic-vs-amazon-bedrock","anthropic","amazon-bedrock","Anthropic vs Amazon Bedrock","Anthropic vs Amazon Bedrock: price per model","Anthropic vs Amazon Bedrock: 21 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Anthropic and Amazon Bedrock share 21 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 21 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Haiku 4.5: price per million tokens by provider (Input)",1,1,"usd","$1.00","$1.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Output)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Cache read)",0.1,0.1,"usd","$0.10","$0.10","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Input)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Output)",50,50,"usd","$50.00","$50.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Cache read)",0.25,0.25,"usd","$0.25","$0.25","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["anthropic-vs-azure","anthropic","azure","Anthropic vs Azure","Anthropic vs Azure: price per model","Anthropic vs Azure: 21 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Anthropic and Azure share 21 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 21 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Haiku 4.5: price per million tokens by provider (Input)",1,1,"usd","$1.00","$1.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Output)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Cache read)",0.1,0.1,"usd","$0.10","$0.10","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Input)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Output)",50,50,"usd","$50.00","$50.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Cache read)",0.25,0.25,"usd","$0.25","$0.25","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["anthropic-vs-google-vertex","anthropic","google-vertex","Anthropic vs Google Vertex AI","Anthropic vs Google Vertex AI: price per model","Anthropic vs Google Vertex AI: 21 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Anthropic and Google Vertex AI share 21 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 21 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Haiku 4.5: price per million tokens by provider (Input)",1,1,"usd","$1.00","$1.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Output)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Cache read)",0.1,0.1,"usd","$0.10","$0.10","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Input)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Output)",50,50,"usd","$50.00","$50.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Cache read)",0.25,0.25,"usd","$0.25","$0.25","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["google-vertex-vs-azure","google-vertex","azure","Google Vertex AI vs Azure","Google Vertex AI vs Azure: price per model","Google Vertex AI vs Azure: 21 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Azure share 21 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 21 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Haiku 4.5: price per million tokens by provider (Input)",1,1,"usd","$1.00","$1.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Output)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Haiku 4.5: price per million tokens by provider (Cache read)",0.1,0.1,"usd","$0.10","$0.10","tie","Same reported list price.","inference-provider-index","provider-prices-claude-haiku-4-5","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Input)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Output)",50,50,"usd","$50.00","$50.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Fable 5.1: price per million tokens by provider (Cache read)",0.25,0.25,"usd","$0.25","$0.25","tie","Same reported list price.","inference-provider-index","provider-prices-claude-fable-5-1","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Fable 5.1 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["deepinfra-vs-parasail","deepinfra","parasail","DeepInfra vs Parasail","DeepInfra vs Parasail: price per model","DeepInfra vs Parasail: 16 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","DeepInfra and Parasail share 16 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 1 tie and 15 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.037,0.1,"usd","$0.037","$0.10","unclear","A reported list price has no interval, so the gap ($0.037 vs $0.10, 2.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.17,0.75,"usd","$0.17","$0.75","unclear","A reported list price has no interval, so the gap ($0.17 vs $0.75, 4.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.1,0.22,"usd","$0.10","$0.22","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.22, 2.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.32,0.5,"usd","$0.32","$0.50","unclear","A reported list price has no interval, so the gap ($0.32 vs $0.50, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",1.3,0.45,"usd","$1.30","$0.45","unclear","A reported list price has no interval, so the gap ($1.30 vs $0.45, 2.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",2.6,3.48,"usd","$2.60","$3.48","unclear","A reported list price has no interval, so the gap ($2.60 vs $3.48, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.1,0.1,"usd","$0.10","$0.10","tie","Same reported list price.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.09,0.14,"usd","$0.090","$0.14","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.14, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",0.18,0.28,"usd","$0.18","$0.28","unclear","A reported list price has no interval, so the gap ($0.18 vs $0.28, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.018,0.07,"usd","$0.018","$0.070","unclear","A reported list price has no interval, so the gap ($0.018 vs $0.070, 3.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Input)",2.85,3,"usd","$2.85","$3.00","unclear","A reported list price has no interval, so the gap ($2.85 vs $3.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",14.25,15,"usd","$14.25","$15.00","unclear","A reported list price has no interval, so the gap ($14.25 vs $15.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.285,0.3,"usd","$0.28","$0.30","unclear","A reported list price has no interval, so the gap ($0.28 vs $0.30, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",0.5625,1.4,"usd","$0.56","$1.40","unclear","A reported list price has no interval, so the gap ($0.56 vs $1.40, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",2.5,4.4,"usd","$2.50","$4.40","unclear","A reported list price has no interval, so the gap ($2.50 vs $4.40, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.125,0.26,"usd","$0.13","$0.26","unclear","A reported list price has no interval, so the gap ($0.13 vs $0.26, 2.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["amazon-bedrock-vs-claude-platform-on-aws","amazon-bedrock","claude-platform-on-aws","Amazon Bedrock vs Claude Platform on AWS","Amazon Bedrock vs Claude Platform on AWS: price per model","Amazon Bedrock vs Claude Platform on AWS: 15 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Claude Platform on AWS share 15 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 15 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["anthropic-vs-claude-platform-on-aws","anthropic","claude-platform-on-aws","Anthropic vs Claude Platform on AWS","Anthropic vs Claude Platform on AWS: price per model","Anthropic vs Claude Platform on AWS: 15 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Anthropic and Claude Platform on AWS share 15 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 15 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["azure-vs-claude-platform-on-aws","azure","claude-platform-on-aws","Azure vs Claude Platform on AWS","Azure vs Claude Platform on AWS: price per model","Azure vs Claude Platform on AWS: 15 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Azure and Claude Platform on AWS share 15 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 15 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["google-vertex-vs-claude-platform-on-aws","google-vertex","claude-platform-on-aws","Google Vertex AI vs Claude Platform on AWS","Google Vertex AI vs Claude Platform on AWS: price per model","Google Vertex AI vs Claude Platform on AWS: 15 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Claude Platform on AWS share 15 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 15 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Claude Sonnet 5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Sonnet 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-sonnet-5-5","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Sonnet 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 4.8: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-4-8","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 4.8 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Output)",25,25,"usd","$25.00","$25.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Input)",4,4,"usd","$4.00","$4.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Output)",20,20,"usd","$20.00","$20.00","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Claude Opus 5.5: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-claude-opus-5-5","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","Claude Opus 5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["parasail-vs-novita","parasail","novita","Parasail vs Novita AI","Parasail vs Novita AI: price per model","Parasail vs Novita AI: 15 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Parasail and Novita AI share 15 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties and 13 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.1,0.05,"usd","$0.10","$0.050","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.050, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.75,0.25,"usd","$0.75","$0.25","unclear","A reported list price has no interval, so the gap ($0.75 vs $0.25, 3.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 4 Maverick: price per million tokens by provider (Input)",0.35,0.27,"usd","$0.35","$0.27","unclear","A reported list price has no interval, so the gap ($0.35 vs $0.27, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-4-maverick","fp8 · Llama 4 Maverick · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 4 Maverick · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 4 Maverick: price per million tokens by provider (Output)",1,0.85,"usd","$1.00","$0.85","unclear","A reported list price has no interval, so the gap ($1.00 vs $0.85, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-4-maverick","fp8 · Llama 4 Maverick · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 4 Maverick · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.22,0.135,"usd","$0.22","$0.14","unclear","A reported list price has no interval, so the gap ($0.22 vs $0.14, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.5,0.4,"usd","$0.50","$0.40","unclear","A reported list price has no interval, so the gap ($0.50 vs $0.40, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",0.45,1.6,"usd","$0.45","$1.60","unclear","A reported list price has no interval, so the gap ($0.45 vs $1.60, 3.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",3.48,3.2,"usd","$3.48","$3.20","unclear","A reported list price has no interval, so the gap ($3.48 vs $3.20, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.1,0.135,"usd","$0.10","$0.14","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.14, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.14,0.14,"usd","$0.14","$0.14","tie","Same reported list price.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",0.28,0.28,"usd","$0.28","$0.28","tie","Same reported list price.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.07,0.028,"usd","$0.070","$0.028","unclear","A reported list price has no interval, so the gap ($0.070 vs $0.028, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.42,"usd","$1.40","$0.42","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.42, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,1.32,"usd","$4.40","$1.32","unclear","A reported list price has no interval, so the gap ($4.40 vs $1.32, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.078,"usd","$0.26","$0.078","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.078, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["claude-haiku-4-5-vs-gpt-6-luna-codex-cli","claude-haiku-4-5","gpt-6-luna-codex-cli","Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI)","Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI): benchmarks","Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI): 14 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. GPT-6 Luna (Codex CLI) leads on 1 row: Total time per attempt: single call vs agent loop, 5.16 s vs 39.0 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation","n"],"$r":[["Strict pass rate: single call vs agent loop on eight hard tasks",0.4583,0.625,"rate","46% (11/24)","63% (10/16)","tie","The 95% intervals overlap (Claude Haiku 4.5 28% to 65%; GPT-6 Luna (Codex CLI) 39% to 82%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-pass-rate",24,16,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.2789,0.6493],[0.3864,0.8152],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Interval merge fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: DST day-length fix",0.3333,1,"rate","33% (1/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 6% to 79%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.0615,0.7923],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: CSV parser",0.6667,1,"rate","67% (2/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.2077,0.9385],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Event-loop order",0,0,"rate","0% (0/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0,0.5615],[0,0.6576],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Room schedule",0,0.5,"rate","0% (0/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0,0.5615],[0.0945,0.9055],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: SemVer regex",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Money refactor",0.6667,0,"rate","67% (2/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.2077,0.9385],[0,0.6576],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: SQL report",0,1,"rate","0% (0/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0,0.5615],[0.3424,1],"\u0001","\u0001"],["Total time per attempt: single call vs agent loop",39.01,5.16,"seconds","39.0 s","5.16 s","b","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s). A range is not a confidence interval.","single-call-vs-agent-loop","agent-loop-total-time",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[15.27,75.13],[3.59,11.32],"\u0001","\u0001"],["Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))",3941,11582,"tokens","3,941","11,582","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tokens",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[3879,4221],[11526,11818],"\u0001","\u0001"],["Tokens per attempt: single call vs agent loop (Output tokens)",5064,345,"tokens","5,064","345","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tokens",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[1899,9321],[36,634],"\u0001","\u0001"],["Tool calls per agent-loop attempt",3,0,"count","3","0","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tool-calls",24,14,"Claude Code · agent loop","Codex CLI · agent loop","range","minmax",[2,18],[0,1],"\u0001","\u0001"],["List-price cost per strict pass: single call vs agent loop (calculation)",0.0672,0.00116,"usd","$0.067","$0.0012","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.0012, 58x) is not tested against run-to-run variation.","single-call-vs-agent-loop","agent-loop-cost-per-pass",24,16,"Claude Code · single call","Codex CLI · single call","\u0001","\u0001","\u0001","\u0001",true,"\u0001"],["Time to first text: a 250-line answer, six models",4,3.3,"seconds","4.00 s","3.30 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy","speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[2.84,6.38],[3.19,3.47],"\u0001",4],["Output speed after the first text: visible tokens per second (calculation)",153.2,129.1,"tokens","153","129","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[152.6,216.1],[55.5,259.1],true,4],["Output speed in characters per second after the first text (calculation)",547,524,"count","547","524","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","speed-anatomy-chars-per-second",3,4,"Claude Code","Codex CLI · effort low","range","minmax",[546,548],[225,1052],true,"\u0001"]]}],["claude-sonnet-5-5-vs-gpt-6-luna-codex-cli","claude-sonnet-5-5","gpt-6-luna-codex-cli","Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI)","Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI): benchmarks","Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI): 14 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. Claude Sonnet 5.5 leads on 1 row: Strict pass rate: single call vs agent loop on eight hard tasks, 100% (24/24) vs 63% (10/16). On those rows the 95% intervals do not overlap. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation","n"],"$r":[["Strict pass rate: single call vs agent loop on eight hard tasks",1,0.625,"rate","100% (24/24)","63% (10/16)","a","The 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%).","single-call-vs-agent-loop","agent-loop-pass-rate",24,16,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.862,1],[0.3864,0.8152],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Interval merge fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: DST day-length fix",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: CSV parser",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Event-loop order",1,0,"rate","100% (3/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0,0.6576],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Room schedule",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: SemVer regex",1,0.5,"rate","100% (3/3)","50% (1/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.0945,0.9055],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: Money refactor",1,0,"rate","100% (3/3)","0% (0/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0,0.6576],"\u0001","\u0001"],["Strict passes per task: single call vs agent loop: SQL report",1,1,"rate","100% (3/3)","100% (2/2)","tie","The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.","single-call-vs-agent-loop","agent-loop-by-task",3,2,"Claude Code · single call","Codex CLI · single call","ci95","ci95",[0.4385,1],[0.3424,1],"\u0001","\u0001"],["Total time per attempt: single call vs agent loop",7.75,5.16,"seconds","7.75 s","5.16 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","single-call-vs-agent-loop","agent-loop-total-time",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[2.26,34.79],[3.59,11.32],"\u0001","\u0001"],["Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))",2281,11582,"tokens","2,281","11,582","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tokens",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[2234,2669],[11526,11818],"\u0001","\u0001"],["Tokens per attempt: single call vs agent loop (Output tokens)",1050,345,"tokens","1,050","345","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","single-call-vs-agent-loop","agent-loop-tokens",24,16,"Claude Code · single call","Codex CLI · single call","range","minmax",[176,3895],[36,634],"\u0001","\u0001"],["Tool calls per agent-loop attempt",0,0,"count","0","0","tie","Same value. More or fewer is not better by itself for this metric.","single-call-vs-agent-loop","agent-loop-tool-calls",16,14,"Claude Code · agent loop","Codex CLI · agent loop","range","minmax",[0,3],[0,1],"\u0001","\u0001"],["List-price cost per strict pass: single call vs agent loop (calculation)",0.01435,0.00116,"usd","$0.014","$0.0012","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation.","single-call-vs-agent-loop","agent-loop-cost-per-pass",24,16,"Claude Code · single call","Codex CLI · single call","\u0001","\u0001","\u0001","\u0001",true,"\u0001"],["Time to first text: a 250-line answer, six models",1.96,3.3,"seconds","1.96 s","3.30 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy","speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[0.88,4.09],[3.19,3.47],"\u0001",4],["Output speed after the first text: visible tokens per second (calculation)",231.7,129.1,"tokens","232","129","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[230.3,233],[55.5,259.1],true,4],["Output speed in characters per second after the first text (calculation)",517,524,"count","517","524","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","speed-anatomy-chars-per-second",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[513,519],[225,1052],true,4]]}],["deepinfra-vs-novita","deepinfra","novita","DeepInfra vs Novita AI","DeepInfra vs Novita AI: price per model","DeepInfra vs Novita AI: 13 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","DeepInfra and Novita AI share 13 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 13 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.037,0.05,"usd","$0.037","$0.050","unclear","A reported list price has no interval, so the gap ($0.037 vs $0.050, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.17,0.25,"usd","$0.17","$0.25","unclear","A reported list price has no interval, so the gap ($0.17 vs $0.25, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.1,0.135,"usd","$0.10","$0.14","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.14, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.32,0.4,"usd","$0.32","$0.40","unclear","A reported list price has no interval, so the gap ($0.32 vs $0.40, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",1.3,1.6,"usd","$1.30","$1.60","unclear","A reported list price has no interval, so the gap ($1.30 vs $1.60, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",2.6,3.2,"usd","$2.60","$3.20","unclear","A reported list price has no interval, so the gap ($2.60 vs $3.20, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.1,0.135,"usd","$0.10","$0.14","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.14, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.09,0.14,"usd","$0.090","$0.14","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.14, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",0.18,0.28,"usd","$0.18","$0.28","unclear","A reported list price has no interval, so the gap ($0.18 vs $0.28, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.018,0.028,"usd","$0.018","$0.028","unclear","A reported list price has no interval, so the gap ($0.018 vs $0.028, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",0.5625,0.42,"usd","$0.56","$0.42","unclear","A reported list price has no interval, so the gap ($0.56 vs $0.42, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",2.5,1.32,"usd","$2.50","$1.32","unclear","A reported list price has no interval, so the gap ($2.50 vs $1.32, 1.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.125,0.078,"usd","$0.13","$0.078","unclear","A reported list price has no interval, so the gap ($0.13 vs $0.078, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["claude-haiku-4-5-vs-claude-fable-5-1","claude-haiku-4-5","claude-fable-5-1","Claude Haiku 4.5 vs Claude Fable 5.1","Claude Haiku 4.5 vs Claude Fable 5.1: measured benchmarks","Claude Haiku 4.5 vs Claude Fable 5.1: 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Haiku 4.5 and Claude Fable 5.1 share 12 measured metrics and 11 list-price calculations from 4 studies. Claude Fable 5.1 leads on 2 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24). On those rows the 95% intervals do not overlap. The other rows are 1 tie and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Fable 5.1 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",4.43,1.94,"seconds","4.43 s","1.94 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Fable 5.1 1.41 s to 9.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[3.16,23.57],[1.41,9.83],"\u0001"],["Time to first useful output",3.63,1.2,"seconds","3.63 s","1.20 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; Claude Fable 5.1 0.95 s to 7.90 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.78,22.27],[0.95,7.9],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",0,2760,"tokens","0","2,760","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",3790,473,"tokens","3,790","473","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",367,64,"tokens","367","64","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00566,0.00987,"usd","$0.0057","$0.0099","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; Claude Fable 5.1 $0.0049 to $0.058); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00513,0.01804],[0.0049,0.05843],true],["List-price cost per passing answer (calculation)",0.00836,0.02054,"usd","$0.0084","$0.021","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.021, 2.5x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Fable 5.1 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Fable 5.1 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,16.13,"seconds","39.0 s","16.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Fable 5.1 4.46 s to 90.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[15.27,75.13],[4.46,90],"\u0001"],["Time to first useful output on hard tasks",35.54,11.63,"seconds","35.5 s","11.6 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; Claude Fable 5.1 2.00 s to 85.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[12.88,70.31],[2,85.33],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",5064,1366,"tokens","5,064","1,366","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.09331,"usd","$0.067","$0.093","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.093) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",91.68,64.24,"percent","91.7%","64.2%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[23.44,97.19],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.024492,0.053696,"usd","$0.024","$0.054","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.001795,0.018702,"usd","$0.0018","$0.019","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.00451,0.02091,"usd","$0.0045","$0.021","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",91.68,64.24,"percent","91.7%","64.2%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code","Claude Code","range","minmax",[76.46,99.27],[23.44,97.19],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",90.19,0,"percent","90.2%","0%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Claude Code","range","minmax",[73.1,97.59],[0,74.01],true],["Time to first text: a 250-line answer, six models",4,4.43,"seconds","4.00 s","4.43 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; Claude Fable 5.1 2.27 s to 4.64 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Claude Code","range","minmax",[2.84,6.38],[2.27,4.64],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",153.2,122.6,"tokens","153","123","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Claude Code","range","minmax",[152.6,216.1],[120.9,131.4],true],["Output speed in characters per second after the first text (calculation)",547,273,"count","547","273","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy","\u0001","speed-anatomy-chars-per-second",3,4,"Claude Code","Claude Code","range","minmax",[546,548],[270,293],true]]}],["google-ai-studio-vs-google-vertex","google-ai-studio","google-vertex","Google AI Studio vs Google Vertex AI","Google AI Studio vs Google Vertex AI: price per model","Google AI Studio vs Google Vertex AI: 12 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google AI Studio and Google Vertex AI share 12 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 12 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Gemini 3.8 Flash: price per million tokens by provider (Input)",0.75,0.75,"usd","$0.75","$0.75","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-8-flash","Gemini 3.8 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.8 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.8 Flash: price per million tokens by provider (Output)",3.75,3.75,"usd","$3.75","$3.75","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-8-flash","Gemini 3.8 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.8 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.8 Flash: price per million tokens by provider (Cache read)",0.075,0.075,"usd","$0.075","$0.075","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-8-flash","Gemini 3.8 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.8 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.5 Flash: price per million tokens by provider (Input)",1.5,1.5,"usd","$1.50","$1.50","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-5-flash","Gemini 3.5 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.5 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.5 Flash: price per million tokens by provider (Output)",9,9,"usd","$9.00","$9.00","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-5-flash","Gemini 3.5 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.5 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.5 Flash: price per million tokens by provider (Cache read)",0.15,0.15,"usd","$0.15","$0.15","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-5-flash","Gemini 3.5 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.5 Flash · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.5 Flash Lite: price per million tokens by provider (Input)",0.3,0.3,"usd","$0.30","$0.30","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-5-flash-lite","Gemini 3.5 Flash Lite · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.5 Flash Lite · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.5 Flash Lite: price per million tokens by provider (Output)",2.5,2.5,"usd","$2.50","$2.50","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-5-flash-lite","Gemini 3.5 Flash Lite · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.5 Flash Lite · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.5 Flash Lite: price per million tokens by provider (Cache read)",0.03,0.03,"usd","$0.030","$0.030","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-5-flash-lite","Gemini 3.5 Flash Lite · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.5 Flash Lite · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.1 Pro Preview: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-1-pro-preview","Gemini 3.1 Pro Preview · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.1 Pro Preview · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.1 Pro Preview: price per million tokens by provider (Output)",12,12,"usd","$12.00","$12.00","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-1-pro-preview","Gemini 3.1 Pro Preview · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.1 Pro Preview · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Gemini 3.1 Pro Preview: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-gemini-3-1-pro-preview","Gemini 3.1 Pro Preview · reported by OpenRouter’s public API, snapshot 2026-10-06","Gemini 3.1 Pro Preview · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["openai-vs-azure","openai","azure","OpenAI vs Azure","OpenAI vs Azure: price per model","OpenAI vs Azure: 12 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","OpenAI and Azure share 12 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 12 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["GPT-6 Sol: price per million tokens by provider (Input)",2,2,"usd","$2.00","$2.00","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-6-sol","GPT-6 Sol · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-6 Sol · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-6 Sol: price per million tokens by provider (Output)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-6-sol","GPT-6 Sol · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-6 Sol · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-6 Sol: price per million tokens by provider (Cache read)",0.2,0.2,"usd","$0.20","$0.20","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-6-sol","GPT-6 Sol · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-6 Sol · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-6 Luna: price per million tokens by provider (Input)",0.1,0.1,"usd","$0.10","$0.10","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-6-luna","GPT-6 Luna · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-6 Luna · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-6 Luna: price per million tokens by provider (Output)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-6-luna","GPT-6 Luna · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-6 Luna · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-6 Luna: price per million tokens by provider (Cache read)",0.01,0.01,"usd","$0.010","$0.010","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-6-luna","GPT-6 Luna · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-6 Luna · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-6 Astra: price per million tokens by provider (Input)",10,10,"usd","$10.00","$10.00","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-6-astra","GPT-6 Astra · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-6 Astra · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-6 Astra: price per million tokens by provider (Output)",50,50,"usd","$50.00","$50.00","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-6-astra","GPT-6 Astra · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-6 Astra · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-6 Astra: price per million tokens by provider (Cache read)",1,1,"usd","$1.00","$1.00","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-6-astra","GPT-6 Astra · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-6 Astra · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-5.5: price per million tokens by provider (Input)",5,5,"usd","$5.00","$5.00","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-5-5","GPT-5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-5.5: price per million tokens by provider (Output)",30,30,"usd","$30.00","$30.00","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-5-5","GPT-5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GPT-5.5: price per million tokens by provider (Cache read)",0.5,0.5,"usd","$0.50","$0.50","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-5-5","GPT-5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06","GPT-5.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["parasail-vs-siliconflow","parasail","siliconflow","Parasail vs SiliconFlow","Parasail vs SiliconFlow: price per model","Parasail vs SiliconFlow: 12 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Parasail and SiliconFlow share 12 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 1 tie and 11 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.1,0.15,"usd","$0.10","$0.15","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.15, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.75,0.6,"usd","$0.75","$0.60","unclear","A reported list price has no interval, so the gap ($0.75 vs $0.60, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.055,0.075,"usd","$0.055","$0.075","unclear","A reported list price has no interval, so the gap ($0.055 vs $0.075, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",0.45,1.50162,"usd","$0.45","$1.50","unclear","A reported list price has no interval, so the gap ($0.45 vs $1.50, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",3.48,3.135,"usd","$3.48","$3.13","unclear","A reported list price has no interval, so the gap ($3.48 vs $3.13, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.1,0.135,"usd","$0.10","$0.14","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.14, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.14,0.13,"usd","$0.14","$0.13","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.13, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",0.28,0.28,"usd","$0.28","$0.28","tie","Same reported list price.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.07,0.028,"usd","$0.070","$0.028","unclear","A reported list price has no interval, so the gap ($0.070 vs $0.028, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.7,"usd","$1.40","$0.70","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.70, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,2.2,"usd","$4.40","$2.20","unclear","A reported list price has no interval, so the gap ($4.40 vs $2.20, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.13,"usd","$0.26","$0.13","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.13, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["deepinfra-vs-cloudflare","deepinfra","cloudflare","DeepInfra vs Cloudflare Workers AI","DeepInfra vs Cloudflare Workers AI: price per model","DeepInfra vs Cloudflare Workers AI: 11 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","DeepInfra and Cloudflare Workers AI share 11 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 11 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.1,0.293,"usd","$0.10","$0.29","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.29, 2.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.32,2.253,"usd","$0.32","$2.25","unclear","A reported list price has no interval, so the gap ($0.32 vs $2.25, 7.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",1.3,1.15,"usd","$1.30","$1.15","unclear","A reported list price has no interval, so the gap ($1.30 vs $1.15, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",2.6,2.55,"usd","$2.60","$2.55","unclear","A reported list price has no interval, so the gap ($2.60 vs $2.55, 1.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.1,0.2,"usd","$0.10","$0.20","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.20, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.09,0.44,"usd","$0.090","$0.44","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.44, 4.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",0.18,1.32,"usd","$0.18","$1.32","unclear","A reported list price has no interval, so the gap ($0.18 vs $1.32, 7.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.018,0.014,"usd","$0.018","$0.014","unclear","A reported list price has no interval, so the gap ($0.018 vs $0.014, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",0.5625,1.4,"usd","$0.56","$1.40","unclear","A reported list price has no interval, so the gap ($0.56 vs $1.40, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",2.5,4.4,"usd","$2.50","$4.40","unclear","A reported list price has no interval, so the gap ($2.50 vs $4.40, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.125,0.26,"usd","$0.13","$0.26","unclear","A reported list price has no interval, so the gap ($0.13 vs $0.26, 2.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["deepinfra-vs-siliconflow","deepinfra","siliconflow","DeepInfra vs SiliconFlow","DeepInfra vs SiliconFlow: price per model","DeepInfra vs SiliconFlow: 11 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","DeepInfra and SiliconFlow share 11 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 11 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.037,0.15,"usd","$0.037","$0.15","unclear","A reported list price has no interval, so the gap ($0.037 vs $0.15, 4.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.17,0.6,"usd","$0.17","$0.60","unclear","A reported list price has no interval, so the gap ($0.17 vs $0.60, 3.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",1.3,1.50162,"usd","$1.30","$1.50","unclear","A reported list price has no interval, so the gap ($1.30 vs $1.50, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",2.6,3.135,"usd","$2.60","$3.13","unclear","A reported list price has no interval, so the gap ($2.60 vs $3.13, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.1,0.135,"usd","$0.10","$0.14","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.14, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.09,0.13,"usd","$0.090","$0.13","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.13, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",0.18,0.28,"usd","$0.18","$0.28","unclear","A reported list price has no interval, so the gap ($0.18 vs $0.28, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.018,0.028,"usd","$0.018","$0.028","unclear","A reported list price has no interval, so the gap ($0.018 vs $0.028, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",0.5625,0.7,"usd","$0.56","$0.70","unclear","A reported list price has no interval, so the gap ($0.56 vs $0.70, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",2.5,2.2,"usd","$2.50","$2.20","unclear","A reported list price has no interval, so the gap ($2.50 vs $2.20, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.125,0.13,"usd","$0.13","$0.13","unclear","A reported list price has no interval, so the gap ($0.13 vs $0.13, 1.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["novita-vs-cloudflare","novita","cloudflare","Novita AI vs Cloudflare Workers AI","Novita AI vs Cloudflare Workers AI: price per model","Novita AI vs Cloudflare Workers AI: 11 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Novita AI and Cloudflare Workers AI share 11 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 11 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.135,0.293,"usd","$0.14","$0.29","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.29, 2.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.4,2.253,"usd","$0.40","$2.25","unclear","A reported list price has no interval, so the gap ($0.40 vs $2.25, 5.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",1.6,1.15,"usd","$1.60","$1.15","unclear","A reported list price has no interval, so the gap ($1.60 vs $1.15, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",3.2,2.55,"usd","$3.20","$2.55","unclear","A reported list price has no interval, so the gap ($3.20 vs $2.55, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.135,0.2,"usd","$0.14","$0.20","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.20, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.14,0.44,"usd","$0.14","$0.44","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.44, 3.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",0.28,1.32,"usd","$0.28","$1.32","unclear","A reported list price has no interval, so the gap ($0.28 vs $1.32, 4.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.028,0.014,"usd","$0.028","$0.014","unclear","A reported list price has no interval, so the gap ($0.028 vs $0.014, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",0.42,1.4,"usd","$0.42","$1.40","unclear","A reported list price has no interval, so the gap ($0.42 vs $1.40, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",1.32,4.4,"usd","$1.32","$4.40","unclear","A reported list price has no interval, so the gap ($1.32 vs $4.40, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.078,0.26,"usd","$0.078","$0.26","unclear","A reported list price has no interval, so the gap ($0.078 vs $0.26, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["novita-vs-siliconflow","novita","siliconflow","Novita AI vs SiliconFlow","Novita AI vs SiliconFlow: price per model","Novita AI vs SiliconFlow: 11 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Novita AI and SiliconFlow share 11 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 ties and 8 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.05,0.15,"usd","$0.050","$0.15","unclear","A reported list price has no interval, so the gap ($0.050 vs $0.15, 3.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.25,0.6,"usd","$0.25","$0.60","unclear","A reported list price has no interval, so the gap ($0.25 vs $0.60, 2.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",1.6,1.50162,"usd","$1.60","$1.50","unclear","A reported list price has no interval, so the gap ($1.60 vs $1.50, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",3.2,3.135,"usd","$3.20","$3.13","unclear","A reported list price has no interval, so the gap ($3.20 vs $3.13, 1.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.135,0.135,"usd","$0.14","$0.14","tie","Same reported list price.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.14,0.13,"usd","$0.14","$0.13","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.13, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",0.28,0.28,"usd","$0.28","$0.28","tie","Same reported list price.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.028,0.028,"usd","$0.028","$0.028","tie","Same reported list price.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",0.42,0.7,"usd","$0.42","$0.70","unclear","A reported list price has no interval, so the gap ($0.42 vs $0.70, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",1.32,2.2,"usd","$1.32","$2.20","unclear","A reported list price has no interval, so the gap ($1.32 vs $2.20, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.078,0.13,"usd","$0.078","$0.13","unclear","A reported list price has no interval, so the gap ($0.078 vs $0.13, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["parasail-vs-cloudflare","parasail","cloudflare","Parasail vs Cloudflare Workers AI","Parasail vs Cloudflare Workers AI: price per model","Parasail vs Cloudflare Workers AI: 11 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Parasail and Cloudflare Workers AI share 11 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 ties and 8 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.22,0.293,"usd","$0.22","$0.29","unclear","A reported list price has no interval, so the gap ($0.22 vs $0.29, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.5,2.253,"usd","$0.50","$2.25","unclear","A reported list price has no interval, so the gap ($0.50 vs $2.25, 4.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",0.45,1.15,"usd","$0.45","$1.15","unclear","A reported list price has no interval, so the gap ($0.45 vs $1.15, 2.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",3.48,2.55,"usd","$3.48","$2.55","unclear","A reported list price has no interval, so the gap ($3.48 vs $2.55, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.1,0.2,"usd","$0.10","$0.20","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.20, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.14,0.44,"usd","$0.14","$0.44","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.44, 3.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",0.28,1.32,"usd","$0.28","$1.32","unclear","A reported list price has no interval, so the gap ($0.28 vs $1.32, 4.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.07,0.014,"usd","$0.070","$0.014","unclear","A reported list price has no interval, so the gap ($0.070 vs $0.014, 5.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.26,"usd","$0.26","$0.26","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["together-vs-deepinfra","together","deepinfra","Together AI vs DeepInfra","Together AI vs DeepInfra: price per model","Together AI vs DeepInfra: 10 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and DeepInfra share 10 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 10 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.037,"usd","$0.15","$0.037","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.037, 4.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.17,"usd","$0.60","$0.17","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.17, 3.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",1.04,0.1,"usd","$1.04","$0.10","unclear","A reported list price has no interval, so the gap ($1.04 vs $0.10, 10x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",1.04,0.32,"usd","$1.04","$0.32","unclear","A reported list price has no interval, so the gap ($1.04 vs $0.32, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Input)",2.7,2.85,"usd","$2.70","$2.85","unclear","A reported list price has no interval, so the gap ($2.70 vs $2.85, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",13.5,14.25,"usd","$13.50","$14.25","unclear","A reported list price has no interval, so the gap ($13.50 vs $14.25, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.27,0.285,"usd","$0.27","$0.28","unclear","A reported list price has no interval, so the gap ($0.27 vs $0.28, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.5625,"usd","$1.40","$0.56","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.56, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,2.5,"usd","$4.40","$2.50","unclear","A reported list price has no interval, so the gap ($4.40 vs $2.50, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.125,"usd","$0.26","$0.13","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.13, 2.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["together-vs-parasail","together","parasail","Together AI vs Parasail","Together AI vs Parasail: price per model","Together AI vs Parasail: 10 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and Parasail share 10 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 ties and 7 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.1,"usd","$0.15","$0.10","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.10, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.75,"usd","$0.60","$0.75","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.75, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",1.04,0.22,"usd","$1.04","$0.22","unclear","A reported list price has no interval, so the gap ($1.04 vs $0.22, 4.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",1.04,0.5,"usd","$1.04","$0.50","unclear","A reported list price has no interval, so the gap ($1.04 vs $0.50, 2.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Input)",2.7,3,"usd","$2.70","$3.00","unclear","A reported list price has no interval, so the gap ($2.70 vs $3.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",13.5,15,"usd","$13.50","$15.00","unclear","A reported list price has no interval, so the gap ($13.50 vs $15.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.27,0.3,"usd","$0.27","$0.30","unclear","A reported list price has no interval, so the gap ($0.27 vs $0.30, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.26,"usd","$0.26","$0.26","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["cloudflare-vs-siliconflow","cloudflare","siliconflow","Cloudflare Workers AI vs SiliconFlow","Cloudflare Workers AI vs SiliconFlow: price per model","Cloudflare Workers AI vs SiliconFlow: 9 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Cloudflare Workers AI and SiliconFlow share 9 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 9 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["DeepSeek V4 Pro 0423: price per million tokens by provider (Input)",1.15,1.50162,"usd","$1.15","$1.50","unclear","A reported list price has no interval, so the gap ($1.15 vs $1.50, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Output)",2.55,3.135,"usd","$2.55","$3.13","unclear","A reported list price has no interval, so the gap ($2.55 vs $3.13, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Pro 0423: price per million tokens by provider (Cache read)",0.2,0.135,"usd","$0.20","$0.14","unclear","A reported list price has no interval, so the gap ($0.20 vs $0.14, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-pro","DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Pro 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Input)",0.44,0.13,"usd","$0.44","$0.13","unclear","A reported list price has no interval, so the gap ($0.44 vs $0.13, 3.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Output)",1.32,0.28,"usd","$1.32","$0.28","unclear","A reported list price has no interval, so the gap ($1.32 vs $0.28, 4.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["DeepSeek V4 Flash 0423: price per million tokens by provider (Cache read)",0.014,0.028,"usd","$0.014","$0.028","unclear","A reported list price has no interval, so the gap ($0.014 vs $0.028, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-deepseek-v4-flash","DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · DeepSeek V4 Flash 0423 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.7,"usd","$1.40","$0.70","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.70, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,2.2,"usd","$4.40","$2.20","unclear","A reported list price has no interval, so the gap ($4.40 vs $2.20, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.13,"usd","$0.26","$0.13","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.13, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["parasail-vs-baseten","parasail","baseten","Parasail vs Baseten","Parasail vs Baseten: price per model","Parasail vs Baseten: 9 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Parasail and Baseten share 9 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 6 ties and 3 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.1,0.1,"usd","$0.10","$0.10","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.75,0.5,"usd","$0.75","$0.50","unclear","A reported list price has no interval, so the gap ($0.75 vs $0.50, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.055,0.1,"usd","$0.055","$0.10","unclear","A reported list price has no interval, so the gap ($0.055 vs $0.10, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Input)",3,3,"usd","$3.00","$3.00","tie","Same reported list price.","inference-provider-index","provider-prices-kimi-k3","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",15,15,"usd","$15.00","$15.00","tie","Same reported list price.","inference-provider-index","provider-prices-kimi-k3","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.3,0.3,"usd","$0.30","$0.30","tie","Same reported list price.","inference-provider-index","provider-prices-kimi-k3","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.14,"usd","$0.26","$0.14","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.14, 1.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["deepinfra-vs-baseten","deepinfra","baseten","DeepInfra vs Baseten","DeepInfra vs Baseten: price per model","DeepInfra vs Baseten: 8 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","DeepInfra and Baseten share 8 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 8 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.037,0.1,"usd","$0.037","$0.10","unclear","A reported list price has no interval, so the gap ($0.037 vs $0.10, 2.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.17,0.5,"usd","$0.17","$0.50","unclear","A reported list price has no interval, so the gap ($0.17 vs $0.50, 2.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Input)",2.85,3,"usd","$2.85","$3.00","unclear","A reported list price has no interval, so the gap ($2.85 vs $3.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",14.25,15,"usd","$14.25","$15.00","unclear","A reported list price has no interval, so the gap ($14.25 vs $15.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.285,0.3,"usd","$0.28","$0.30","unclear","A reported list price has no interval, so the gap ($0.28 vs $0.30, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",0.5625,1.4,"usd","$0.56","$1.40","unclear","A reported list price has no interval, so the gap ($0.56 vs $1.40, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",2.5,4.4,"usd","$2.50","$4.40","unclear","A reported list price has no interval, so the gap ($2.50 vs $4.40, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.125,0.14,"usd","$0.13","$0.14","unclear","A reported list price has no interval, so the gap ($0.13 vs $0.14, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["together-vs-baseten","together","baseten","Together AI vs Baseten","Together AI vs Baseten: price per model","Together AI vs Baseten: 8 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and Baseten share 8 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties and 6 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.1,"usd","$0.15","$0.10","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.10, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.5,"usd","$0.60","$0.50","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.50, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Input)",2.7,3,"usd","$2.70","$3.00","unclear","A reported list price has no interval, so the gap ($2.70 vs $3.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",13.5,15,"usd","$13.50","$15.00","unclear","A reported list price has no interval, so the gap ($13.50 vs $15.00, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.27,0.3,"usd","$0.27","$0.30","unclear","A reported list price has no interval, so the gap ($0.27 vs $0.30, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.14,"usd","$0.26","$0.14","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.14, 1.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["together-vs-novita","together","novita","Together AI vs Novita AI","Together AI vs Novita AI: price per model","Together AI vs Novita AI: 7 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and Novita AI share 7 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 7 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.05,"usd","$0.15","$0.050","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.050, 3.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.25,"usd","$0.60","$0.25","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.25, 2.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",1.04,0.135,"usd","$1.04","$0.14","unclear","A reported list price has no interval, so the gap ($1.04 vs $0.14, 7.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",1.04,0.4,"usd","$1.04","$0.40","unclear","A reported list price has no interval, so the gap ($1.04 vs $0.40, 2.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.42,"usd","$1.40","$0.42","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.42, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,1.32,"usd","$4.40","$1.32","unclear","A reported list price has no interval, so the gap ($4.40 vs $1.32, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.078,"usd","$0.26","$0.078","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.078, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["baseten-vs-siliconflow","baseten","siliconflow","Baseten vs SiliconFlow","Baseten vs SiliconFlow: price per model","Baseten vs SiliconFlow: 6 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Baseten and SiliconFlow share 6 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 6 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.1,0.15,"usd","$0.10","$0.15","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.15, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.5,0.6,"usd","$0.50","$0.60","unclear","A reported list price has no interval, so the gap ($0.50 vs $0.60, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.1,0.075,"usd","$0.10","$0.075","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.075, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.7,"usd","$1.40","$0.70","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.70, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,2.2,"usd","$4.40","$2.20","unclear","A reported list price has no interval, so the gap ($4.40 vs $2.20, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.14,0.13,"usd","$0.14","$0.13","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.13, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["fireworks-vs-baseten","fireworks","baseten","Fireworks AI vs Baseten","Fireworks AI vs Baseten: price per model","Fireworks AI vs Baseten: 6 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Fireworks AI and Baseten share 6 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 5 ties and 1 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Kimi K3: price per million tokens by provider (Input)",3,3,"usd","$3.00","$3.00","tie","Same reported list price.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",15,15,"usd","$15.00","$15.00","tie","Same reported list price.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.3,0.3,"usd","$0.30","$0.30","tie","Same reported list price.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.14,"usd","$0.26","$0.14","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.14, 1.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["fireworks-vs-deepinfra","fireworks","deepinfra","Fireworks AI vs DeepInfra","Fireworks AI vs DeepInfra: price per model","Fireworks AI vs DeepInfra: 6 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Fireworks AI and DeepInfra share 6 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 6 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Kimi K3: price per million tokens by provider (Input)",3,2.85,"usd","$3.00","$2.85","unclear","A reported list price has no interval, so the gap ($3.00 vs $2.85, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",15,14.25,"usd","$15.00","$14.25","unclear","A reported list price has no interval, so the gap ($15.00 vs $14.25, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.3,0.285,"usd","$0.30","$0.28","unclear","A reported list price has no interval, so the gap ($0.30 vs $0.28, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","mxfp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.5625,"usd","$1.40","$0.56","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.56, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,2.5,"usd","$4.40","$2.50","unclear","A reported list price has no interval, so the gap ($4.40 vs $2.50, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.125,"usd","$0.26","$0.13","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.13, 2.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["fireworks-vs-parasail","fireworks","parasail","Fireworks AI vs Parasail","Fireworks AI vs Parasail: price per model","Fireworks AI vs Parasail: 6 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Fireworks AI and Parasail share 6 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 6 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Kimi K3: price per million tokens by provider (Input)",3,3,"usd","$3.00","$3.00","tie","Same reported list price.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Output)",15,15,"usd","$15.00","$15.00","tie","Same reported list price.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Kimi K3: price per million tokens by provider (Cache read)",0.3,0.3,"usd","$0.30","$0.30","tie","Same reported list price.","inference-provider-index","provider-prices-kimi-k3","Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.26,"usd","$0.26","$0.26","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["gpt-6-1-sol-codex-cli-vs-gpt-6-luna-codex-cli","gpt-6-1-sol-codex-cli","gpt-6-luna-codex-cli","GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (Codex CLI)","GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (Codex CLI)","GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (Codex CLI): 6 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) share 6 measured metrics and 2 list-price calculations from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 8 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["CLI vs API: time for a one-line answer (Total time)",4.19,3.19,"seconds","4.19 s","3.19 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6 Luna (Codex CLI) 2.88 s to 3.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","Codex CLI · effort none · fixed exact reply, 5 runs","range","minmax",[3.81,4.69],[2.88,3.83],"\u0001"],["CLI vs API: time for a one-line answer (First useful output)",3.79,2.79,"seconds","3.79 s","2.79 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6 Luna (Codex CLI) 2.46 s to 3.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","Codex CLI · effort none · fixed exact reply, 5 runs","range","minmax",[3.37,4.3],[2.46,3.42],"\u0001"],["CLI vs API: time for a small coding task (Total time)",17.85,9.23,"seconds","17.9 s","9.23 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6 Luna (Codex CLI) 8.99 s to 11.7 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","Codex CLI · effort none · small coding task, 3 runs","range","minmax",[17.68,22.42],[8.99,11.68],"\u0001"],["CLI vs API: time for a small coding task (First useful output)",17.27,8.68,"seconds","17.3 s","8.68 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6 Luna (Codex CLI) 8.27 s to 11.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","Codex CLI · effort none · small coding task, 3 runs","range","minmax",[17.13,21.86],[8.27,11.01],"\u0001"],["Hidden prompt: input tokens for the same one-line request",19555,18859,"tokens","19,555","18,859","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort high · short fixed tasks","Codex CLI · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time to first text: a 250-line answer, six models",3.52,3.3,"seconds","3.52 s","3.30 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Codex CLI · effort low","Codex CLI · effort low","range","minmax",[2.75,4.42],[3.19,3.47],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",79.6,129.1,"tokens","80","129","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Codex CLI · effort low","Codex CLI · effort low","range","minmax",[71.6,80.5],[55.5,259.1],true],["Output speed in characters per second after the first text (calculation)",323,524,"count","323","524","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Codex CLI · effort low","Codex CLI · effort low","range","minmax",[291,327],[225,1052],true]]}],["groq-vs-parasail","groq","parasail","Groq vs Parasail","Groq vs Parasail: price per model","Groq vs Parasail: 6 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and Parasail share 6 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 6 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.1,"usd","$0.15","$0.10","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.10, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.75,"usd","$0.60","$0.75","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.75, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.075,0.055,"usd","$0.075","$0.055","unclear","A reported list price has no interval, so the gap ($0.075 vs $0.055, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.59,0.22,"usd","$0.59","$0.22","unclear","A reported list price has no interval, so the gap ($0.59 vs $0.22, 2.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.79,0.5,"usd","$0.79","$0.50","unclear","A reported list price has no interval, so the gap ($0.79 vs $0.50, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Cache read)",0.295,0.11,"usd","$0.29","$0.11","unclear","A reported list price has no interval, so the gap ($0.29 vs $0.11, 2.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["gpt-6-1-sol-codex-cli-vs-gpt-6-luna-openai-api","gpt-6-1-sol-codex-cli","gpt-6-luna-openai-api","GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (OpenAI API)","GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (OpenAI API)","GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (OpenAI API): 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (OpenAI API) share 5 measured metrics from one study. GPT-6 Luna (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 0.97 s vs 4.19 s; CLI vs API: time for a one-line answer (First useful output), 0.82 s vs 3.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example Codex CLI vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",4.19,0.97,"seconds","4.19 s","0.97 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6 Luna (OpenAI API) 0.65 s to 1.50 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[3.81,4.69],[0.65,1.5]],["CLI vs API: time for a one-line answer (First useful output)",3.79,0.82,"seconds","3.79 s","0.82 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6 Luna (OpenAI API) 0.51 s to 1.37 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[3.37,4.3],[0.51,1.37]],["CLI vs API: time for a small coding task (Total time)",17.85,4.01,"seconds","17.9 s","4.01 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6 Luna (OpenAI API) 3.83 s to 4.35 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[17.68,22.42],[3.83,4.35]],["CLI vs API: time for a small coding task (First useful output)",17.27,0.67,"seconds","17.3 s","0.67 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6 Luna (OpenAI API) 0.62 s to 0.81 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[17.13,21.86],[0.62,0.81]],["Hidden prompt: input tokens for the same one-line request",19555,17,"tokens","19,555","17","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort high · short fixed tasks","OpenAI API · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-6-1-sol-openai-api-vs-gpt-6-luna-codex-cli","gpt-6-1-sol-openai-api","gpt-6-luna-codex-cli","GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI)","GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI)","GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI): 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (OpenAI API) and GPT-6 Luna (Codex CLI) share 5 measured metrics from one study. GPT-6.1 Sol (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 1.52 s vs 3.19 s; CLI vs API: time for a one-line answer (First useful output), 1.34 s vs 2.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example OpenAI API vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",1.52,3.19,"seconds","1.52 s","3.19 s","a","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s; GPT-6 Luna (Codex CLI) 2.88 s to 3.83 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort high · fixed exact reply, 5 runs","Codex CLI · effort none · fixed exact reply, 5 runs","range","minmax",[1.35,2.23],[2.88,3.83]],["CLI vs API: time for a one-line answer (First useful output)",1.34,2.79,"seconds","1.34 s","2.79 s","a","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s; GPT-6 Luna (Codex CLI) 2.46 s to 3.42 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort high · fixed exact reply, 5 runs","Codex CLI · effort none · fixed exact reply, 5 runs","range","minmax",[1.26,2.12],[2.46,3.42]],["CLI vs API: time for a small coding task (Total time)",9.56,9.23,"seconds","9.56 s","9.23 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) 9.44 s to 10.9 s; GPT-6 Luna (Codex CLI) 8.99 s to 11.7 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort high · small coding task, 3 runs","Codex CLI · effort none · small coding task, 3 runs","range","minmax",[9.44,10.94],[8.99,11.68]],["CLI vs API: time for a small coding task (First useful output)",5.31,8.68,"seconds","5.31 s","8.68 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 4.99 s to 6.42 s; GPT-6 Luna (Codex CLI) 8.27 s to 11.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort high · small coding task, 3 runs","Codex CLI · effort none · small coding task, 3 runs","range","minmax",[4.99,6.42],[8.27,11.01]],["Hidden prompt: input tokens for the same one-line request",17,18859,"tokens","17","18,859","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"OpenAI API · effort high · short fixed tasks","Codex CLI · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-6-1-sol-openai-api-vs-gpt-6-luna-openai-api","gpt-6-1-sol-openai-api","gpt-6-luna-openai-api","GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (OpenAI API)","GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (OpenAI API)","GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (OpenAI API): 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (OpenAI API) and GPT-6 Luna (OpenAI API) share 5 measured metrics from one study. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 4 unclear; each row says why. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",1.52,0.97,"seconds","1.52 s","0.97 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s; GPT-6 Luna (OpenAI API) 0.65 s to 1.50 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort high · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[1.35,2.23],[0.65,1.5]],["CLI vs API: time for a one-line answer (First useful output)",1.34,0.82,"seconds","1.34 s","0.82 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s; GPT-6 Luna (OpenAI API) 0.51 s to 1.37 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort high · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[1.26,2.12],[0.51,1.37]],["CLI vs API: time for a small coding task (Total time)",9.56,4.01,"seconds","9.56 s","4.01 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 9.44 s to 10.9 s; GPT-6 Luna (OpenAI API) 3.83 s to 4.35 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort high · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[9.44,10.94],[3.83,4.35]],["CLI vs API: time for a small coding task (First useful output)",5.31,0.67,"seconds","5.31 s","0.67 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 4.99 s to 6.42 s; GPT-6 Luna (OpenAI API) 0.62 s to 0.81 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort high · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[4.99,6.42],[0.62,0.81]],["Hidden prompt: input tokens for the same one-line request",17,17,"tokens","17","17","tie","Same value. More or fewer is not better by itself for this metric.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"OpenAI API · effort high · short fixed tasks","OpenAI API · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-6-luna-codex-cli-vs-gpt-6-luna-openai-api","gpt-6-luna-codex-cli","gpt-6-luna-openai-api","GPT-6 Luna (Codex CLI) vs GPT-6 Luna (OpenAI API)","GPT-6 Luna (Codex CLI) vs GPT-6 Luna (OpenAI API)","GPT-6 Luna (Codex CLI) vs GPT-6 Luna (OpenAI API): 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","GPT-6 Luna (Codex CLI) and GPT-6 Luna (OpenAI API) share 5 measured metrics from one study. GPT-6 Luna (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 0.97 s vs 3.19 s; CLI vs API: time for a one-line answer (First useful output), 0.82 s vs 2.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example Codex CLI vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",3.19,0.97,"seconds","3.19 s","0.97 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6 Luna (Codex CLI) 2.88 s to 3.83 s; GPT-6 Luna (OpenAI API) 0.65 s to 1.50 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort none · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[2.88,3.83],[0.65,1.5]],["CLI vs API: time for a one-line answer (First useful output)",2.79,0.82,"seconds","2.79 s","0.82 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6 Luna (Codex CLI) 2.46 s to 3.42 s; GPT-6 Luna (OpenAI API) 0.51 s to 1.37 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort none · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[2.46,3.42],[0.51,1.37]],["CLI vs API: time for a small coding task (Total time)",9.23,4.01,"seconds","9.23 s","4.01 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6 Luna (Codex CLI) 8.99 s to 11.7 s; GPT-6 Luna (OpenAI API) 3.83 s to 4.35 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort none · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[8.99,11.68],[3.83,4.35]],["CLI vs API: time for a small coding task (First useful output)",8.68,0.67,"seconds","8.68 s","0.67 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6 Luna (Codex CLI) 8.27 s to 11.0 s; GPT-6 Luna (OpenAI API) 0.62 s to 0.81 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort none · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[8.27,11.01],[0.62,0.81]],["Hidden prompt: input tokens for the same one-line request",18859,17,"tokens","18,859","17","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort none · short fixed tasks","OpenAI API · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}],["novita-vs-baseten","novita","baseten","Novita AI vs Baseten","Novita AI vs Baseten: price per model","Novita AI vs Baseten: 5 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Novita AI and Baseten share 5 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 5 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.05,0.1,"usd","$0.050","$0.10","unclear","A reported list price has no interval, so the gap ($0.050 vs $0.10, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.25,0.5,"usd","$0.25","$0.50","unclear","A reported list price has no interval, so the gap ($0.25 vs $0.50, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",0.42,1.4,"usd","$0.42","$1.40","unclear","A reported list price has no interval, so the gap ($0.42 vs $1.40, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",1.32,4.4,"usd","$1.32","$4.40","unclear","A reported list price has no interval, so the gap ($1.32 vs $4.40, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.078,0.14,"usd","$0.078","$0.14","unclear","A reported list price has no interval, so the gap ($0.078 vs $0.14, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["together-vs-cloudflare","together","cloudflare","Together AI vs Cloudflare Workers AI","Together AI vs Cloudflare Workers AI: price per model","Together AI vs Cloudflare Workers AI: 5 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and Cloudflare Workers AI share 5 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 ties and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",1.04,0.293,"usd","$1.04","$0.29","unclear","A reported list price has no interval, so the gap ($1.04 vs $0.29, 3.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",1.04,2.253,"usd","$1.04","$2.25","unclear","A reported list price has no interval, so the gap ($1.04 vs $2.25, 2.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.26,"usd","$0.26","$0.26","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["together-vs-siliconflow","together","siliconflow","Together AI vs SiliconFlow","Together AI vs SiliconFlow: price per model","Together AI vs SiliconFlow: 5 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and SiliconFlow share 5 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties and 3 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.15,"usd","$0.15","$0.15","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.6,"usd","$0.60","$0.60","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.7,"usd","$1.40","$0.70","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.70, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,2.2,"usd","$4.40","$2.20","unclear","A reported list price has no interval, so the gap ($4.40 vs $2.20, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.13,"usd","$0.26","$0.13","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.13, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["deepinfra-vs-nebius","deepinfra","nebius","DeepInfra vs Nebius","DeepInfra vs Nebius: price per model","DeepInfra vs Nebius: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","DeepInfra and Nebius share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.037,0.15,"usd","$0.037","$0.15","unclear","A reported list price has no interval, so the gap ($0.037 vs $0.15, 4.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.17,0.6,"usd","$0.17","$0.60","unclear","A reported list price has no interval, so the gap ($0.17 vs $0.60, 3.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",0.5625,1.4,"usd","$0.56","$1.40","unclear","A reported list price has no interval, so the gap ($0.56 vs $1.40, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",2.5,4.4,"usd","$2.50","$4.40","unclear","A reported list price has no interval, so the gap ($2.50 vs $4.40, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["deepinfra-vs-sambanova","deepinfra","sambanova","DeepInfra vs SambaNova","DeepInfra vs SambaNova: price per model","DeepInfra vs SambaNova: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","DeepInfra and SambaNova share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.037,0.14,"usd","$0.037","$0.14","unclear","A reported list price has no interval, so the gap ($0.037 vs $0.14, 3.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.17,0.95,"usd","$0.17","$0.95","unclear","A reported list price has no interval, so the gap ($0.17 vs $0.95, 5.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.1,0.45,"usd","$0.10","$0.45","unclear","A reported list price has no interval, so the gap ($0.10 vs $0.45, 4.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.32,0.9,"usd","$0.32","$0.90","unclear","A reported list price has no interval, so the gap ($0.32 vs $0.90, 2.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["google-vertex-vs-deepinfra","google-vertex","deepinfra","Google Vertex AI vs DeepInfra","Google Vertex AI vs DeepInfra: price per model","Google Vertex AI vs DeepInfra: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and DeepInfra share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.09,0.037,"usd","$0.090","$0.037","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.037, 2.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.36,0.17,"usd","$0.36","$0.17","unclear","A reported list price has no interval, so the gap ($0.36 vs $0.17, 2.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.72,0.1,"usd","$0.72","$0.10","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.10, 7.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.72,0.32,"usd","$0.72","$0.32","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.32, 2.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["google-vertex-vs-groq","google-vertex","groq","Google Vertex AI vs Groq","Google Vertex AI vs Groq: price per model","Google Vertex AI vs Groq: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Groq share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.09,0.15,"usd","$0.090","$0.15","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.15, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.36,0.6,"usd","$0.36","$0.60","unclear","A reported list price has no interval, so the gap ($0.36 vs $0.60, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.72,0.59,"usd","$0.72","$0.59","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.59, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.72,0.79,"usd","$0.72","$0.79","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.79, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["google-vertex-vs-novita","google-vertex","novita","Google Vertex AI vs Novita AI","Google Vertex AI vs Novita AI: price per model","Google Vertex AI vs Novita AI: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Novita AI share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.09,0.05,"usd","$0.090","$0.050","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.050, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.36,0.25,"usd","$0.36","$0.25","unclear","A reported list price has no interval, so the gap ($0.36 vs $0.25, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.72,0.135,"usd","$0.72","$0.14","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.14, 5.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.72,0.4,"usd","$0.72","$0.40","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.40, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["google-vertex-vs-parasail","google-vertex","parasail","Google Vertex AI vs Parasail","Google Vertex AI vs Parasail: price per model","Google Vertex AI vs Parasail: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Parasail share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.09,0.1,"usd","$0.090","$0.10","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.10, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.36,0.75,"usd","$0.36","$0.75","unclear","A reported list price has no interval, so the gap ($0.36 vs $0.75, 2.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.72,0.22,"usd","$0.72","$0.22","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.22, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.72,0.5,"usd","$0.72","$0.50","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.50, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["google-vertex-vs-sambanova","google-vertex","sambanova","Google Vertex AI vs SambaNova","Google Vertex AI vs SambaNova: price per model","Google Vertex AI vs SambaNova: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and SambaNova share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.09,0.14,"usd","$0.090","$0.14","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.14, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.36,0.95,"usd","$0.36","$0.95","unclear","A reported list price has no interval, so the gap ($0.36 vs $0.95, 2.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.72,0.45,"usd","$0.72","$0.45","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.45, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.72,0.9,"usd","$0.72","$0.90","unclear","A reported list price has no interval, so the gap ($0.72 vs $0.90, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["google-vertex-vs-together","google-vertex","together","Google Vertex AI vs Together AI","Google Vertex AI vs Together AI: price per model","Google Vertex AI vs Together AI: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Together AI share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.09,0.15,"usd","$0.090","$0.15","unclear","A reported list price has no interval, so the gap ($0.090 vs $0.15, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.36,0.6,"usd","$0.36","$0.60","unclear","A reported list price has no interval, so the gap ($0.36 vs $0.60, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.72,1.04,"usd","$0.72","$1.04","unclear","A reported list price has no interval, so the gap ($0.72 vs $1.04, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.72,1.04,"usd","$0.72","$1.04","unclear","A reported list price has no interval, so the gap ($0.72 vs $1.04, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["groq-vs-deepinfra","groq","deepinfra","Groq vs DeepInfra","Groq vs DeepInfra: price per model","Groq vs DeepInfra: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and DeepInfra share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.037,"usd","$0.15","$0.037","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.037, 4.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.17,"usd","$0.60","$0.17","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.17, 3.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.59,0.1,"usd","$0.59","$0.10","unclear","A reported list price has no interval, so the gap ($0.59 vs $0.10, 5.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.79,0.32,"usd","$0.79","$0.32","unclear","A reported list price has no interval, so the gap ($0.79 vs $0.32, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["groq-vs-novita","groq","novita","Groq vs Novita AI","Groq vs Novita AI: price per model","Groq vs Novita AI: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and Novita AI share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.05,"usd","$0.15","$0.050","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.050, 3.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.25,"usd","$0.60","$0.25","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.25, 2.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.59,0.135,"usd","$0.59","$0.14","unclear","A reported list price has no interval, so the gap ($0.59 vs $0.14, 4.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.79,0.4,"usd","$0.79","$0.40","unclear","A reported list price has no interval, so the gap ($0.79 vs $0.40, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["groq-vs-sambanova","groq","sambanova","Groq vs SambaNova","Groq vs SambaNova: price per model","Groq vs SambaNova: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and SambaNova share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.14,"usd","$0.15","$0.14","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.14, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.95,"usd","$0.60","$0.95","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.95, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.59,0.45,"usd","$0.59","$0.45","unclear","A reported list price has no interval, so the gap ($0.59 vs $0.45, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.79,0.9,"usd","$0.79","$0.90","unclear","A reported list price has no interval, so the gap ($0.79 vs $0.90, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["nebius-vs-baseten","nebius","baseten","Nebius vs Baseten","Nebius vs Baseten: price per model","Nebius vs Baseten: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Nebius and Baseten share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.1,"usd","$0.15","$0.10","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.10, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.5,"usd","$0.60","$0.50","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.50, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["nebius-vs-novita","nebius","novita","Nebius vs Novita AI","Nebius vs Novita AI: price per model","Nebius vs Novita AI: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Nebius and Novita AI share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.05,"usd","$0.15","$0.050","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.050, 3.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.25,"usd","$0.60","$0.25","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.25, 2.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.42,"usd","$1.40","$0.42","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.42, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,1.32,"usd","$4.40","$1.32","unclear","A reported list price has no interval, so the gap ($4.40 vs $1.32, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["nebius-vs-parasail","nebius","parasail","Nebius vs Parasail","Nebius vs Parasail: price per model","Nebius vs Parasail: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Nebius and Parasail share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.1,"usd","$0.15","$0.10","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.10, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.75,"usd","$0.60","$0.75","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.75, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["nebius-vs-siliconflow","nebius","siliconflow","Nebius vs SiliconFlow","Nebius vs SiliconFlow: price per model","Nebius vs SiliconFlow: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Nebius and SiliconFlow share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.15,"usd","$0.15","$0.15","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.6,"usd","$0.60","$0.60","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,0.7,"usd","$1.40","$0.70","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.70, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,2.2,"usd","$4.40","$2.20","unclear","A reported list price has no interval, so the gap ($4.40 vs $2.20, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["sambanova-vs-novita","sambanova","novita","SambaNova vs Novita AI","SambaNova vs Novita AI: price per model","SambaNova vs Novita AI: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","SambaNova and Novita AI share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.14,0.05,"usd","$0.14","$0.050","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.050, 2.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.95,0.25,"usd","$0.95","$0.25","unclear","A reported list price has no interval, so the gap ($0.95 vs $0.25, 3.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.45,0.135,"usd","$0.45","$0.14","unclear","A reported list price has no interval, so the gap ($0.45 vs $0.14, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.9,0.4,"usd","$0.90","$0.40","unclear","A reported list price has no interval, so the gap ($0.90 vs $0.40, 2.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bf16 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["sambanova-vs-parasail","sambanova","parasail","SambaNova vs Parasail","SambaNova vs Parasail: price per model","SambaNova vs Parasail: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","SambaNova and Parasail share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.14,0.1,"usd","$0.14","$0.10","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.10, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.95,0.75,"usd","$0.95","$0.75","unclear","A reported list price has no interval, so the gap ($0.95 vs $0.75, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",0.45,0.22,"usd","$0.45","$0.22","unclear","A reported list price has no interval, so the gap ($0.45 vs $0.22, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",0.9,0.5,"usd","$0.90","$0.50","unclear","A reported list price has no interval, so the gap ($0.90 vs $0.50, 1.8x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["together-vs-nebius","together","nebius","Together AI vs Nebius","Together AI vs Nebius: price per model","Together AI vs Nebius: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and Nebius share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.15,"usd","$0.15","$0.15","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.6,"usd","$0.60","$0.60","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["together-vs-sambanova","together","sambanova","Together AI vs SambaNova","Together AI vs SambaNova: price per model","Together AI vs SambaNova: 4 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and SambaNova share 4 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 4 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.14,"usd","$0.15","$0.14","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.14, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.95,"usd","$0.60","$0.95","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.95, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Input)",1.04,0.45,"usd","$1.04","$0.45","unclear","A reported list price has no interval, so the gap ($1.04 vs $0.45, 2.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"],["Llama 3.3 70B Instruct: price per million tokens by provider (Output)",1.04,0.9,"usd","$1.04","$0.90","unclear","A reported list price has no interval, so the gap ($1.04 vs $0.90, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-llama-3-3-70b-instruct","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["baseten-vs-cloudflare","baseten","cloudflare","Baseten vs Cloudflare Workers AI","Baseten vs Cloudflare Workers AI: price per model","Baseten vs Cloudflare Workers AI: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Baseten and Cloudflare Workers AI share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties and 1 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.14,0.26,"usd","$0.14","$0.26","unclear","A reported list price has no interval, so the gap ($0.14 vs $0.26, 1.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["cerebras-vs-baseten","cerebras","baseten","Cerebras vs Baseten","Cerebras vs Baseten: price per model","Cerebras vs Baseten: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Cerebras and Baseten share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.35,0.1,"usd","$0.35","$0.10","unclear","A reported list price has no interval, so the gap ($0.35 vs $0.10, 3.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.75,0.5,"usd","$0.75","$0.50","unclear","A reported list price has no interval, so the gap ($0.75 vs $0.50, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.35,0.1,"usd","$0.35","$0.10","unclear","A reported list price has no interval, so the gap ($0.35 vs $0.10, 3.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["cerebras-vs-parasail","cerebras","parasail","Cerebras vs Parasail","Cerebras vs Parasail: price per model","Cerebras vs Parasail: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Cerebras and Parasail share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.35,0.1,"usd","$0.35","$0.10","unclear","A reported list price has no interval, so the gap ($0.35 vs $0.10, 3.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.75,0.75,"usd","$0.75","$0.75","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.35,0.055,"usd","$0.35","$0.055","unclear","A reported list price has no interval, so the gap ($0.35 vs $0.055, 6.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["cerebras-vs-siliconflow","cerebras","siliconflow","Cerebras vs SiliconFlow","Cerebras vs SiliconFlow: price per model","Cerebras vs SiliconFlow: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Cerebras and SiliconFlow share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.35,0.15,"usd","$0.35","$0.15","unclear","A reported list price has no interval, so the gap ($0.35 vs $0.15, 2.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.75,0.6,"usd","$0.75","$0.60","unclear","A reported list price has no interval, so the gap ($0.75 vs $0.60, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.35,0.075,"usd","$0.35","$0.075","unclear","A reported list price has no interval, so the gap ($0.35 vs $0.075, 4.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["claude-haiku-4-5-vs-claude-opus-4-5","claude-haiku-4-5","claude-opus-4-5","Claude Haiku 4.5 vs Claude Opus 4.5","Claude Haiku 4.5 vs Claude Opus 4.5: measured benchmarks","Claude Haiku 4.5 vs Claude Opus 4.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and Claude Opus 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7273,"rate","76% (25/33)","73% (24/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; Claude Opus 4.5 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5578,0.8493]],["Model calls per instance",68.5,35.9,"calls","68.5","35.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,1.184,"usd","$0.48","$1.18","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $1.18, 2.5x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-haiku-4-5-vs-claude-opus-4-6","claude-haiku-4-5","claude-opus-4-6","Claude Haiku 4.5 vs Claude Opus 4.6","Claude Haiku 4.5 vs Claude Opus 4.6: measured benchmarks","Claude Haiku 4.5 vs Claude Opus 4.6: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and Claude Opus 4.6 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; Claude Opus 4.6 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262]],["Model calls per instance",68.5,28.9,"calls","68.5","28.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,0.875,"usd","$0.48","$0.88","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $0.88) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-haiku-4-5-vs-claude-sonnet-4-5","claude-haiku-4-5","claude-sonnet-4-5","Claude Haiku 4.5 vs Claude Sonnet 4.5","Claude Haiku 4.5 vs Claude Sonnet 4.5: measured benchmarks","Claude Haiku 4.5 vs Claude Sonnet 4.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and Claude Sonnet 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7576,"rate","76% (25/33)","76% (25/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; Claude Sonnet 4.5 59% to 87%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5898,0.8717]],["Model calls per instance",68.5,51,"calls","68.5","51","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,0.913,"usd","$0.48","$0.91","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $0.91) is not tested against run-to-run variation.","cost-thought-experiments",25,"cost-per-resolved-agent-vs-panel",25,25,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-haiku-4-5-vs-deepseek-v3-2","claude-haiku-4-5","deepseek-v3-2","Claude Haiku 4.5 vs DeepSeek V3.2","Claude Haiku 4.5 vs DeepSeek V3.2: measured benchmarks","Claude Haiku 4.5 vs DeepSeek V3.2: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and DeepSeek V3.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7273,"rate","76% (25/33)","73% (24/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; DeepSeek V3.2 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5578,0.8493]],["Model calls per instance",68.5,88.2,"calls","68.5","88.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,0.637,"usd","$0.48","$0.64","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $0.64) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-haiku-4-5-vs-gemini-3-flash","claude-haiku-4-5","gemini-3-flash","Claude Haiku 4.5 vs Gemini 3 Flash","Claude Haiku 4.5 vs Gemini 3 Flash: measured benchmarks","Claude Haiku 4.5 vs Gemini 3 Flash: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and Gemini 3 Flash share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.8182,"rate","76% (25/33)","82% (27/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; Gemini 3 Flash 66% to 91%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.6561,0.9139]],["Model calls per instance",68.5,54.2,"calls","68.5","54.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,0.436,"usd","$0.48","$0.44","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $0.44) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,27,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-haiku-4-5-vs-glm-5","claude-haiku-4-5","glm-5","Claude Haiku 4.5 vs GLM 5","Claude Haiku 4.5 vs GLM 5: measured benchmarks","Claude Haiku 4.5 vs GLM 5: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Claude Haiku 4.5 and GLM 5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7879,"rate","76% (25/33)","79% (26/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; GLM 5 62% to 89%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.6225,0.8932]],["Model calls per instance",68.5,77.5,"calls","68.5","77.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,0.667,"usd","$0.48","$0.67","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $0.67) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,26,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-haiku-4-5-vs-gpt-5-2","claude-haiku-4-5","gpt-5-2","Claude Haiku 4.5 vs GPT 5.2","Claude Haiku 4.5 vs GPT 5.2: measured benchmarks","Claude Haiku 4.5 vs GPT 5.2: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and GPT 5.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.8485,"rate","76% (25/33)","85% (28/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; GPT 5.2 69% to 93%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.6908,0.9335]],["Model calls per instance",68.5,35.6,"calls","68.5","35.6","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,0.628,"usd","$0.48","$0.63","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $0.63) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,28,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-haiku-4-5-vs-gpt-5-mini","claude-haiku-4-5","gpt-5-mini","Claude Haiku 4.5 vs GPT 5 mini","Claude Haiku 4.5 vs GPT 5 mini: measured benchmarks","Claude Haiku 4.5 vs GPT 5 mini: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.6364,"rate","76% (25/33)","64% (21/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.4662,0.7781]],["Model calls per instance",68.5,20.8,"calls","68.5","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,0.08,"usd","$0.48","$0.080","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $0.080, 6.0x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,21,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-haiku-4-5-vs-kimi-k2-5","claude-haiku-4-5","kimi-k2-5","Claude Haiku 4.5 vs Kimi K2.5","Claude Haiku 4.5 vs Kimi K2.5: measured benchmarks","Claude Haiku 4.5 vs Kimi K2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262]],["Model calls per instance",68.5,56.7,"calls","68.5","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,0.256,"usd","$0.48","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $0.26) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-haiku-4-5-vs-minimax-m2-5","claude-haiku-4-5","minimax-m2-5","Claude Haiku 4.5 vs MiniMax M2.5","Claude Haiku 4.5 vs MiniMax M2.5: measured benchmarks","Claude Haiku 4.5 vs MiniMax M2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Haiku 4.5 and MiniMax M2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; MiniMax M2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262]],["Model calls per instance",68.5,58.4,"calls","68.5","58.4","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.479,0.107,"usd","$0.48","$0.11","unclear","No interval or range was recorded for either side, so the gap ($0.48 vs $0.11, 4.5x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-4-5-vs-claude-opus-4-6","claude-opus-4-5","claude-opus-4-6","Claude Opus 4.5 vs Claude Opus 4.6","Claude Opus 4.5 vs Claude Opus 4.6: measured benchmarks","Claude Opus 4.5 vs Claude Opus 4.6: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 4.5 and Claude Opus 4.6 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7273,0.697,"rate","73% (24/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Opus 4.5 56% to 85%; Claude Opus 4.6 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5578,0.8493],[0.5266,0.8262]],["Model calls per instance",35.9,28.9,"calls","35.9","28.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",1.184,0.875,"usd","$1.18","$0.88","unclear","No interval or range was recorded for either side, so the gap ($1.18 vs $0.88) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",24,23,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-4-5-vs-deepseek-v3-2","claude-opus-4-5","deepseek-v3-2","Claude Opus 4.5 vs DeepSeek V3.2","Claude Opus 4.5 vs DeepSeek V3.2: measured benchmarks","Claude Opus 4.5 vs DeepSeek V3.2: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 4.5 and DeepSeek V3.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7273,0.7273,"rate","73% (24/33)","73% (24/33)","tie","The 95% intervals overlap (Claude Opus 4.5 56% to 85%; DeepSeek V3.2 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5578,0.8493],[0.5578,0.8493]],["Model calls per instance",35.9,88.2,"calls","35.9","88.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",1.184,0.637,"usd","$1.18","$0.64","unclear","No interval or range was recorded for either side, so the gap ($1.18 vs $0.64) is not tested against run-to-run variation.","cost-thought-experiments",24,"cost-per-resolved-agent-vs-panel",24,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-4-5-vs-gpt-5-mini","claude-opus-4-5","gpt-5-mini","Claude Opus 4.5 vs GPT 5 mini","Claude Opus 4.5 vs GPT 5 mini: measured benchmarks","Claude Opus 4.5 vs GPT 5 mini: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 4.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7273,0.6364,"rate","73% (24/33)","64% (21/33)","tie","The 95% intervals overlap (Claude Opus 4.5 56% to 85%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5578,0.8493],[0.4662,0.7781]],["Model calls per instance",35.9,20.8,"calls","35.9","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",1.184,0.08,"usd","$1.18","$0.080","unclear","No interval or range was recorded for either side, so the gap ($1.18 vs $0.080, 15x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",24,21,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-4-5-vs-kimi-k2-5","claude-opus-4-5","kimi-k2-5","Claude Opus 4.5 vs Kimi K2.5","Claude Opus 4.5 vs Kimi K2.5: measured benchmarks","Claude Opus 4.5 vs Kimi K2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 4.5 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7273,0.697,"rate","73% (24/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Opus 4.5 56% to 85%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5578,0.8493],[0.5266,0.8262]],["Model calls per instance",35.9,56.7,"calls","35.9","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",1.184,0.256,"usd","$1.18","$0.26","unclear","No interval or range was recorded for either side, so the gap ($1.18 vs $0.26, 4.6x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",24,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-4-5-vs-minimax-m2-5","claude-opus-4-5","minimax-m2-5","Claude Opus 4.5 vs MiniMax M2.5","Claude Opus 4.5 vs MiniMax M2.5: measured benchmarks","Claude Opus 4.5 vs MiniMax M2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 4.5 and MiniMax M2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7273,0.697,"rate","73% (24/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Opus 4.5 56% to 85%; MiniMax M2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5578,0.8493],[0.5266,0.8262]],["Model calls per instance",35.9,58.4,"calls","35.9","58.4","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",1.184,0.107,"usd","$1.18","$0.11","unclear","No interval or range was recorded for either side, so the gap ($1.18 vs $0.11, 11x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",24,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-4-6-vs-deepseek-v3-2","claude-opus-4-6","deepseek-v3-2","Claude Opus 4.6 vs DeepSeek V3.2","Claude Opus 4.6 vs DeepSeek V3.2: measured benchmarks","Claude Opus 4.6 vs DeepSeek V3.2: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 4.6 and DeepSeek V3.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.697,0.7273,"rate","70% (23/33)","73% (24/33)","tie","The 95% intervals overlap (Claude Opus 4.6 53% to 83%; DeepSeek V3.2 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5266,0.8262],[0.5578,0.8493]],["Model calls per instance",28.9,88.2,"calls","28.9","88.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.875,0.637,"usd","$0.88","$0.64","unclear","No interval or range was recorded for either side, so the gap ($0.88 vs $0.64) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",23,24,"public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-4-6-vs-gpt-5-mini","claude-opus-4-6","gpt-5-mini","Claude Opus 4.6 vs GPT 5 mini","Claude Opus 4.6 vs GPT 5 mini: measured benchmarks","Claude Opus 4.6 vs GPT 5 mini: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 4.6 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.697,0.6364,"rate","70% (23/33)","64% (21/33)","tie","The 95% intervals overlap (Claude Opus 4.6 53% to 83%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5266,0.8262],[0.4662,0.7781]],["Model calls per instance",28.9,20.8,"calls","28.9","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.875,0.08,"usd","$0.88","$0.080","unclear","No interval or range was recorded for either side, so the gap ($0.88 vs $0.080, 11x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",23,21,"public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-4-6-vs-kimi-k2-5","claude-opus-4-6","kimi-k2-5","Claude Opus 4.6 vs Kimi K2.5","Claude Opus 4.6 vs Kimi K2.5: measured benchmarks","Claude Opus 4.6 vs Kimi K2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 4.6 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.697,0.697,"rate","70% (23/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Opus 4.6 53% to 83%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5266,0.8262],[0.5266,0.8262]],["Model calls per instance",28.9,56.7,"calls","28.9","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.875,0.256,"usd","$0.88","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.88 vs $0.26, 3.4x) is not tested against run-to-run variation.","cost-thought-experiments",23,"cost-per-resolved-agent-vs-panel",23,23,"public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-opus-4-6-vs-minimax-m2-5","claude-opus-4-6","minimax-m2-5","Claude Opus 4.6 vs MiniMax M2.5","Claude Opus 4.6 vs MiniMax M2.5: measured benchmarks","Claude Opus 4.6 vs MiniMax M2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 4.6 and MiniMax M2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.697,0.697,"rate","70% (23/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Opus 4.6 53% to 83%; MiniMax M2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5266,0.8262],[0.5266,0.8262]],["Model calls per instance",28.9,58.4,"calls","28.9","58.4","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.875,0.107,"usd","$0.88","$0.11","unclear","No interval or range was recorded for either side, so the gap ($0.88 vs $0.11, 8.2x) is not tested against run-to-run variation.","cost-thought-experiments",23,"cost-per-resolved-agent-vs-panel",23,23,"public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-sonnet-4-5-vs-claude-opus-4-5","claude-sonnet-4-5","claude-opus-4-5","Claude Sonnet 4.5 vs Claude Opus 4.5","Claude Sonnet 4.5 vs Claude Opus 4.5: measured benchmarks","Claude Sonnet 4.5 vs Claude Opus 4.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Sonnet 4.5 and Claude Opus 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7273,"rate","76% (25/33)","73% (24/33)","tie","The 95% intervals overlap (Claude Sonnet 4.5 59% to 87%; Claude Opus 4.5 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5578,0.8493]],["Model calls per instance",51,35.9,"calls","51","35.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.913,1.184,"usd","$0.91","$1.18","unclear","No interval or range was recorded for either side, so the gap ($0.91 vs $1.18) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-sonnet-4-5-vs-claude-opus-4-6","claude-sonnet-4-5","claude-opus-4-6","Claude Sonnet 4.5 vs Claude Opus 4.6","Claude Sonnet 4.5 vs Claude Opus 4.6: measured benchmarks","Claude Sonnet 4.5 vs Claude Opus 4.6: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Sonnet 4.5 and Claude Opus 4.6 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Sonnet 4.5 59% to 87%; Claude Opus 4.6 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262]],["Model calls per instance",51,28.9,"calls","51","28.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.913,0.875,"usd","$0.91","$0.88","unclear","No interval or range was recorded for either side, so the gap ($0.91 vs $0.88) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-sonnet-4-5-vs-deepseek-v3-2","claude-sonnet-4-5","deepseek-v3-2","Claude Sonnet 4.5 vs DeepSeek V3.2","Claude Sonnet 4.5 vs DeepSeek V3.2: measured benchmarks","Claude Sonnet 4.5 vs DeepSeek V3.2: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Sonnet 4.5 and DeepSeek V3.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7273,"rate","76% (25/33)","73% (24/33)","tie","The 95% intervals overlap (Claude Sonnet 4.5 59% to 87%; DeepSeek V3.2 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5578,0.8493]],["Model calls per instance",51,88.2,"calls","51","88.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.913,0.637,"usd","$0.91","$0.64","unclear","No interval or range was recorded for either side, so the gap ($0.91 vs $0.64) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-sonnet-4-5-vs-gpt-5-mini","claude-sonnet-4-5","gpt-5-mini","Claude Sonnet 4.5 vs GPT 5 mini","Claude Sonnet 4.5 vs GPT 5 mini: measured benchmarks","Claude Sonnet 4.5 vs GPT 5 mini: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Sonnet 4.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.6364,"rate","76% (25/33)","64% (21/33)","tie","The 95% intervals overlap (Claude Sonnet 4.5 59% to 87%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.4662,0.7781]],["Model calls per instance",51,20.8,"calls","51","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.913,0.08,"usd","$0.91","$0.080","unclear","No interval or range was recorded for either side, so the gap ($0.91 vs $0.080, 11x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,21,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-sonnet-4-5-vs-kimi-k2-5","claude-sonnet-4-5","kimi-k2-5","Claude Sonnet 4.5 vs Kimi K2.5","Claude Sonnet 4.5 vs Kimi K2.5: measured benchmarks","Claude Sonnet 4.5 vs Kimi K2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Sonnet 4.5 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Sonnet 4.5 59% to 87%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262]],["Model calls per instance",51,56.7,"calls","51","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.913,0.256,"usd","$0.91","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.91 vs $0.26, 3.6x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-sonnet-4-5-vs-minimax-m2-5","claude-sonnet-4-5","minimax-m2-5","Claude Sonnet 4.5 vs MiniMax M2.5","Claude Sonnet 4.5 vs MiniMax M2.5: measured benchmarks","Claude Sonnet 4.5 vs MiniMax M2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Sonnet 4.5 and MiniMax M2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Claude Sonnet 4.5 59% to 87%; MiniMax M2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262]],["Model calls per instance",51,58.4,"calls","51","58.4","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.913,0.107,"usd","$0.91","$0.11","unclear","No interval or range was recorded for either side, so the gap ($0.91 vs $0.11, 8.5x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["claude-sonnet-5-5-vs-gpt-6-1-sol-openai-api","claude-sonnet-5-5","gpt-6-1-sol-openai-api","Claude Sonnet 5.5 vs GPT-6.1 Sol (OpenAI API)","Claude Sonnet 5.5 vs GPT-6.1 Sol (OpenAI API): benchmarks","Claude Sonnet 5.5 vs GPT-6.1 Sol (OpenAI API): 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 and GPT-6.1 Sol (OpenAI API) share 3 measured metrics from one study. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example Claude Code vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Repairing a scheduler: Claude Code vs Codex vs API (Total time)",15,17.32,"seconds","15.0 s","17.3 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 13.9 s to 15.9 s; GPT-6.1 Sol (OpenAI API) 16.3 s to 18.6 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs","OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[13.89,15.89],[16.28,18.61]],["Repairing a scheduler: Claude Code vs Codex vs API (First useful output)",7.55,7.46,"seconds","7.55 s","7.46 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 6.77 s to 7.63 s; GPT-6.1 Sol (OpenAI API) 6.68 s to 9.05 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs","OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[6.77,7.63],[6.68,9.05]],["Output tokens to repair the scheduler (Output tokens)",2227,1313,"tokens","2,227","1,313","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",3,"scheduler-repair-output-tokens",3,3,"Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs","OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs","\u0001","\u0001","\u0001","\u0001"]]}],["deepseek-v3-2-vs-gpt-5-mini","deepseek-v3-2","gpt-5-mini","DeepSeek V3.2 vs GPT 5 mini","DeepSeek V3.2 vs GPT 5 mini: measured benchmarks","DeepSeek V3.2 vs GPT 5 mini: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","DeepSeek V3.2 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7273,0.6364,"rate","73% (24/33)","64% (21/33)","tie","The 95% intervals overlap (DeepSeek V3.2 56% to 85%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5578,0.8493],[0.4662,0.7781]],["Model calls per instance",88.2,20.8,"calls","88.2","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.637,0.08,"usd","$0.64","$0.080","unclear","No interval or range was recorded for either side, so the gap ($0.64 vs $0.080, 8.0x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",24,21,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["deepseek-v3-2-vs-kimi-k2-5","deepseek-v3-2","kimi-k2-5","DeepSeek V3.2 vs Kimi K2.5","DeepSeek V3.2 vs Kimi K2.5: measured benchmarks","DeepSeek V3.2 vs Kimi K2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","DeepSeek V3.2 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7273,0.697,"rate","73% (24/33)","70% (23/33)","tie","The 95% intervals overlap (DeepSeek V3.2 56% to 85%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5578,0.8493],[0.5266,0.8262]],["Model calls per instance",88.2,56.7,"calls","88.2","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.637,0.256,"usd","$0.64","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.64 vs $0.26, 2.5x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",24,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["deepseek-v3-2-vs-minimax-m2-5","deepseek-v3-2","minimax-m2-5","DeepSeek V3.2 vs MiniMax M2.5","DeepSeek V3.2 vs MiniMax M2.5: measured benchmarks","DeepSeek V3.2 vs MiniMax M2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","DeepSeek V3.2 and MiniMax M2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7273,0.697,"rate","73% (24/33)","70% (23/33)","tie","The 95% intervals overlap (DeepSeek V3.2 56% to 85%; MiniMax M2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5578,0.8493],[0.5266,0.8262]],["Model calls per instance",88.2,58.4,"calls","88.2","58.4","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.637,0.107,"usd","$0.64","$0.11","unclear","No interval or range was recorded for either side, so the gap ($0.64 vs $0.11, 6.0x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",24,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["fireworks-vs-cloudflare","fireworks","cloudflare","Fireworks AI vs Cloudflare Workers AI","Fireworks AI vs Cloudflare Workers AI: price per model","Fireworks AI vs Cloudflare Workers AI: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Fireworks AI and Cloudflare Workers AI share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["GLM 5.3: price per million tokens by provider (Input)",1.4,1.4,"usd","$1.40","$1.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,4.4,"usd","$4.40","$4.40","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.26,"usd","$0.26","$0.26","tie","Same reported list price.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["fireworks-vs-novita","fireworks","novita","Fireworks AI vs Novita AI","Fireworks AI vs Novita AI: price per model","Fireworks AI vs Novita AI: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Fireworks AI and Novita AI share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["GLM 5.3: price per million tokens by provider (Input)",1.4,0.42,"usd","$1.40","$0.42","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.42, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,1.32,"usd","$4.40","$1.32","unclear","A reported list price has no interval, so the gap ($4.40 vs $1.32, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.078,"usd","$0.26","$0.078","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.078, 3.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["fireworks-vs-siliconflow","fireworks","siliconflow","Fireworks AI vs SiliconFlow","Fireworks AI vs SiliconFlow: price per model","Fireworks AI vs SiliconFlow: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Fireworks AI and SiliconFlow share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["GLM 5.3: price per million tokens by provider (Input)",1.4,0.7,"usd","$1.40","$0.70","unclear","A reported list price has no interval, so the gap ($1.40 vs $0.70, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Output)",4.4,2.2,"usd","$4.40","$2.20","unclear","A reported list price has no interval, so the gap ($4.40 vs $2.20, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"],["GLM 5.3: price per million tokens by provider (Cache read)",0.26,0.13,"usd","$0.26","$0.13","unclear","A reported list price has no interval, so the gap ($0.26 vs $0.13, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-glm-5-3","GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["gemini-3-flash-vs-claude-opus-4-5","gemini-3-flash","claude-opus-4-5","Gemini 3 Flash vs Claude Opus 4.5","Gemini 3 Flash vs Claude Opus 4.5: measured benchmarks","Gemini 3 Flash vs Claude Opus 4.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Gemini 3 Flash and Claude Opus 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8182,0.7273,"rate","82% (27/33)","73% (24/33)","tie","The 95% intervals overlap (Gemini 3 Flash 66% to 91%; Claude Opus 4.5 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6561,0.9139],[0.5578,0.8493]],["Model calls per instance",54.2,35.9,"calls","54.2","35.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.436,1.184,"usd","$0.44","$1.18","unclear","No interval or range was recorded for either side, so the gap ($0.44 vs $1.18, 2.7x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",27,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gemini-3-flash-vs-claude-opus-4-6","gemini-3-flash","claude-opus-4-6","Gemini 3 Flash vs Claude Opus 4.6","Gemini 3 Flash vs Claude Opus 4.6: measured benchmarks","Gemini 3 Flash vs Claude Opus 4.6: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Gemini 3 Flash and Claude Opus 4.6 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8182,0.697,"rate","82% (27/33)","70% (23/33)","tie","The 95% intervals overlap (Gemini 3 Flash 66% to 91%; Claude Opus 4.6 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6561,0.9139],[0.5266,0.8262]],["Model calls per instance",54.2,28.9,"calls","54.2","28.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.436,0.875,"usd","$0.44","$0.88","unclear","No interval or range was recorded for either side, so the gap ($0.44 vs $0.88, 2.0x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",27,23,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gemini-3-flash-vs-claude-sonnet-4-5","gemini-3-flash","claude-sonnet-4-5","Gemini 3 Flash vs Claude Sonnet 4.5","Gemini 3 Flash vs Claude Sonnet 4.5: measured benchmarks","Gemini 3 Flash vs Claude Sonnet 4.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Gemini 3 Flash and Claude Sonnet 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8182,0.7576,"rate","82% (27/33)","76% (25/33)","tie","The 95% intervals overlap (Gemini 3 Flash 66% to 91%; Claude Sonnet 4.5 59% to 87%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6561,0.9139],[0.5898,0.8717]],["Model calls per instance",54.2,51,"calls","54.2","51","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.436,0.913,"usd","$0.44","$0.91","unclear","No interval or range was recorded for either side, so the gap ($0.44 vs $0.91, 2.1x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",27,25,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gemini-3-flash-vs-deepseek-v3-2","gemini-3-flash","deepseek-v3-2","Gemini 3 Flash vs DeepSeek V3.2","Gemini 3 Flash vs DeepSeek V3.2: measured benchmarks","Gemini 3 Flash vs DeepSeek V3.2: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Gemini 3 Flash and DeepSeek V3.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8182,0.7273,"rate","82% (27/33)","73% (24/33)","tie","The 95% intervals overlap (Gemini 3 Flash 66% to 91%; DeepSeek V3.2 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6561,0.9139],[0.5578,0.8493]],["Model calls per instance",54.2,88.2,"calls","54.2","88.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.436,0.637,"usd","$0.44","$0.64","unclear","No interval or range was recorded for either side, so the gap ($0.44 vs $0.64) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",27,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gemini-3-flash-vs-glm-5","gemini-3-flash","glm-5","Gemini 3 Flash vs GLM 5","Gemini 3 Flash vs GLM 5: measured benchmarks","Gemini 3 Flash vs GLM 5: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Gemini 3 Flash and GLM 5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8182,0.7879,"rate","82% (27/33)","79% (26/33)","tie","The 95% intervals overlap (Gemini 3 Flash 66% to 91%; GLM 5 62% to 89%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6561,0.9139],[0.6225,0.8932]],["Model calls per instance",54.2,77.5,"calls","54.2","77.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.436,0.667,"usd","$0.44","$0.67","unclear","No interval or range was recorded for either side, so the gap ($0.44 vs $0.67) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",27,26,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gemini-3-flash-vs-gpt-5-mini","gemini-3-flash","gpt-5-mini","Gemini 3 Flash vs GPT 5 mini","Gemini 3 Flash vs GPT 5 mini: measured benchmarks","Gemini 3 Flash vs GPT 5 mini: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Gemini 3 Flash and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8182,0.6364,"rate","82% (27/33)","64% (21/33)","tie","The 95% intervals overlap (Gemini 3 Flash 66% to 91%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6561,0.9139],[0.4662,0.7781]],["Model calls per instance",54.2,20.8,"calls","54.2","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.436,0.08,"usd","$0.44","$0.080","unclear","No interval or range was recorded for either side, so the gap ($0.44 vs $0.080, 5.5x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",27,21,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gemini-3-flash-vs-kimi-k2-5","gemini-3-flash","kimi-k2-5","Gemini 3 Flash vs Kimi K2.5","Gemini 3 Flash vs Kimi K2.5: measured benchmarks","Gemini 3 Flash vs Kimi K2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Gemini 3 Flash and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8182,0.697,"rate","82% (27/33)","70% (23/33)","tie","The 95% intervals overlap (Gemini 3 Flash 66% to 91%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6561,0.9139],[0.5266,0.8262]],["Model calls per instance",54.2,56.7,"calls","54.2","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.436,0.256,"usd","$0.44","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.44 vs $0.26) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",27,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gemini-3-flash-vs-minimax-m2-5","gemini-3-flash","minimax-m2-5","Gemini 3 Flash vs MiniMax M2.5","Gemini 3 Flash vs MiniMax M2.5: measured benchmarks","Gemini 3 Flash vs MiniMax M2.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Gemini 3 Flash and MiniMax M2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8182,0.697,"rate","82% (27/33)","70% (23/33)","tie","The 95% intervals overlap (Gemini 3 Flash 66% to 91%; MiniMax M2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6561,0.9139],[0.5266,0.8262]],["Model calls per instance",54.2,58.4,"calls","54.2","58.4","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.436,0.107,"usd","$0.44","$0.11","unclear","No interval or range was recorded for either side, so the gap ($0.44 vs $0.11, 4.1x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",27,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["glm-5-vs-claude-opus-4-5","glm-5","claude-opus-4-5","GLM 5 vs Claude Opus 4.5","GLM 5 vs Claude Opus 4.5: measured benchmarks","GLM 5 vs Claude Opus 4.5: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GLM 5 and Claude Opus 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7879,0.7273,"rate","79% (26/33)","73% (24/33)","tie","The 95% intervals overlap (GLM 5 62% to 89%; Claude Opus 4.5 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6225,0.8932],[0.5578,0.8493]],["Model calls per instance",77.5,35.9,"calls","77.5","35.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.667,1.184,"usd","$0.67","$1.18","unclear","No interval or range was recorded for either side, so the gap ($0.67 vs $1.18) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",26,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["glm-5-vs-claude-opus-4-6","glm-5","claude-opus-4-6","GLM 5 vs Claude Opus 4.6","GLM 5 vs Claude Opus 4.6: measured benchmarks","GLM 5 vs Claude Opus 4.6: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GLM 5 and Claude Opus 4.6 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7879,0.697,"rate","79% (26/33)","70% (23/33)","tie","The 95% intervals overlap (GLM 5 62% to 89%; Claude Opus 4.6 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6225,0.8932],[0.5266,0.8262]],["Model calls per instance",77.5,28.9,"calls","77.5","28.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.667,0.875,"usd","$0.67","$0.88","unclear","No interval or range was recorded for either side, so the gap ($0.67 vs $0.88) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",26,23,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["glm-5-vs-claude-sonnet-4-5","glm-5","claude-sonnet-4-5","GLM 5 vs Claude Sonnet 4.5","GLM 5 vs Claude Sonnet 4.5: measured benchmarks","GLM 5 vs Claude Sonnet 4.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","GLM 5 and Claude Sonnet 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7879,0.7576,"rate","79% (26/33)","76% (25/33)","tie","The 95% intervals overlap (GLM 5 62% to 89%; Claude Sonnet 4.5 59% to 87%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6225,0.8932],[0.5898,0.8717]],["Model calls per instance",77.5,51,"calls","77.5","51","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.667,0.913,"usd","$0.67","$0.91","unclear","No interval or range was recorded for either side, so the gap ($0.67 vs $0.91) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",26,25,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["glm-5-vs-deepseek-v3-2","glm-5","deepseek-v3-2","GLM 5 vs DeepSeek V3.2","GLM 5 vs DeepSeek V3.2: measured benchmarks","GLM 5 vs DeepSeek V3.2: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GLM 5 and DeepSeek V3.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7879,0.7273,"rate","79% (26/33)","73% (24/33)","tie","The 95% intervals overlap (GLM 5 62% to 89%; DeepSeek V3.2 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6225,0.8932],[0.5578,0.8493]],["Model calls per instance",77.5,88.2,"calls","77.5","88.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.667,0.637,"usd","$0.67","$0.64","unclear","No interval or range was recorded for either side, so the gap ($0.67 vs $0.64) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",26,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["glm-5-vs-gpt-5-mini","glm-5","gpt-5-mini","GLM 5 vs GPT 5 mini","GLM 5 vs GPT 5 mini: measured benchmarks","GLM 5 vs GPT 5 mini: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GLM 5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7879,0.6364,"rate","79% (26/33)","64% (21/33)","tie","The 95% intervals overlap (GLM 5 62% to 89%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6225,0.8932],[0.4662,0.7781]],["Model calls per instance",77.5,20.8,"calls","77.5","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.667,0.08,"usd","$0.67","$0.080","unclear","No interval or range was recorded for either side, so the gap ($0.67 vs $0.080, 8.3x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",26,21,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["glm-5-vs-kimi-k2-5","glm-5","kimi-k2-5","GLM 5 vs Kimi K2.5","GLM 5 vs Kimi K2.5: measured benchmarks","GLM 5 vs Kimi K2.5: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GLM 5 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7879,0.697,"rate","79% (26/33)","70% (23/33)","tie","The 95% intervals overlap (GLM 5 62% to 89%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6225,0.8932],[0.5266,0.8262]],["Model calls per instance",77.5,56.7,"calls","77.5","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.667,0.256,"usd","$0.67","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.67 vs $0.26, 2.6x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",26,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["glm-5-vs-minimax-m2-5","glm-5","minimax-m2-5","GLM 5 vs MiniMax M2.5","GLM 5 vs MiniMax M2.5: measured benchmarks","GLM 5 vs MiniMax M2.5: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GLM 5 and MiniMax M2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7879,0.697,"rate","79% (26/33)","70% (23/33)","tie","The 95% intervals overlap (GLM 5 62% to 89%; MiniMax M2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6225,0.8932],[0.5266,0.8262]],["Model calls per instance",77.5,58.4,"calls","77.5","58.4","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.667,0.107,"usd","$0.67","$0.11","unclear","No interval or range was recorded for either side, so the gap ($0.67 vs $0.11, 6.2x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",26,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-5-2-vs-claude-opus-4-5","gpt-5-2","claude-opus-4-5","GPT 5.2 vs Claude Opus 4.5","GPT 5.2 vs Claude Opus 4.5: measured benchmarks","GPT 5.2 vs Claude Opus 4.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","GPT 5.2 and Claude Opus 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.7273,"rate","85% (28/33)","73% (24/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; Claude Opus 4.5 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.5578,0.8493]],["Model calls per instance",35.6,35.9,"calls","35.6","35.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,1.184,"usd","$0.63","$1.18","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $1.18) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-5-2-vs-claude-opus-4-6","gpt-5-2","claude-opus-4-6","GPT 5.2 vs Claude Opus 4.6","GPT 5.2 vs Claude Opus 4.6: measured benchmarks","GPT 5.2 vs Claude Opus 4.6: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","GPT 5.2 and Claude Opus 4.6 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.697,"rate","85% (28/33)","70% (23/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; Claude Opus 4.6 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.5266,0.8262]],["Model calls per instance",35.6,28.9,"calls","35.6","28.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,0.875,"usd","$0.63","$0.88","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $0.88) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,23,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-5-2-vs-claude-sonnet-4-5","gpt-5-2","claude-sonnet-4-5","GPT 5.2 vs Claude Sonnet 4.5","GPT 5.2 vs Claude Sonnet 4.5: measured benchmarks","GPT 5.2 vs Claude Sonnet 4.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","GPT 5.2 and Claude Sonnet 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.7576,"rate","85% (28/33)","76% (25/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; Claude Sonnet 4.5 59% to 87%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.5898,0.8717]],["Model calls per instance",35.6,51,"calls","35.6","51","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,0.913,"usd","$0.63","$0.91","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $0.91) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,25,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-5-2-vs-deepseek-v3-2","gpt-5-2","deepseek-v3-2","GPT 5.2 vs DeepSeek V3.2","GPT 5.2 vs DeepSeek V3.2: measured benchmarks","GPT 5.2 vs DeepSeek V3.2: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GPT 5.2 and DeepSeek V3.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.7273,"rate","85% (28/33)","73% (24/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; DeepSeek V3.2 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.5578,0.8493]],["Model calls per instance",35.6,88.2,"calls","35.6","88.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,0.637,"usd","$0.63","$0.64","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $0.64) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-5-2-vs-gemini-3-flash","gpt-5-2","gemini-3-flash","GPT 5.2 vs Gemini 3 Flash","GPT 5.2 vs Gemini 3 Flash: measured benchmarks","GPT 5.2 vs Gemini 3 Flash: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GPT 5.2 and Gemini 3 Flash share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.8182,"rate","85% (28/33)","82% (27/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; Gemini 3 Flash 66% to 91%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.6561,0.9139]],["Model calls per instance",35.6,54.2,"calls","35.6","54.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,0.436,"usd","$0.63","$0.44","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $0.44) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,27,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-5-2-vs-glm-5","gpt-5-2","glm-5","GPT 5.2 vs GLM 5","GPT 5.2 vs GLM 5: measured benchmarks","GPT 5.2 vs GLM 5: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GPT 5.2 and GLM 5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.7879,"rate","85% (28/33)","79% (26/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; GLM 5 62% to 89%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.6225,0.8932]],["Model calls per instance",35.6,77.5,"calls","35.6","77.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,0.667,"usd","$0.63","$0.67","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $0.67) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,26,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-5-2-vs-gpt-5-mini","gpt-5-2","gpt-5-mini","GPT 5.2 vs GPT 5 mini","GPT 5.2 vs GPT 5 mini: measured benchmarks","GPT 5.2 vs GPT 5 mini: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GPT 5.2 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.6364,"rate","85% (28/33)","64% (21/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.4662,0.7781]],["Model calls per instance",35.6,20.8,"calls","35.6","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,0.08,"usd","$0.63","$0.080","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $0.080, 7.8x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,21,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-5-2-vs-kimi-k2-5","gpt-5-2","kimi-k2-5","GPT 5.2 vs Kimi K2.5","GPT 5.2 vs Kimi K2.5: measured benchmarks","GPT 5.2 vs Kimi K2.5: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GPT 5.2 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.697,"rate","85% (28/33)","70% (23/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.5266,0.8262]],["Model calls per instance",35.6,56.7,"calls","35.6","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,0.256,"usd","$0.63","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $0.26, 2.5x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["gpt-5-2-vs-minimax-m2-5","gpt-5-2","minimax-m2-5","GPT 5.2 vs MiniMax M2.5","GPT 5.2 vs MiniMax M2.5: measured benchmarks","GPT 5.2 vs MiniMax M2.5: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","GPT 5.2 and MiniMax M2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.697,"rate","85% (28/33)","70% (23/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; MiniMax M2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.5266,0.8262]],["Model calls per instance",35.6,58.4,"calls","35.6","58.4","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,0.107,"usd","$0.63","$0.11","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $0.11, 5.9x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["groq-vs-baseten","groq","baseten","Groq vs Baseten","Groq vs Baseten: price per model","Groq vs Baseten: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and Baseten share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.1,"usd","$0.15","$0.10","unclear","A reported list price has no interval, so the gap ($0.15 vs $0.10, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.5,"usd","$0.60","$0.50","unclear","A reported list price has no interval, so the gap ($0.60 vs $0.50, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.075,0.1,"usd","$0.075","$0.10","unclear","A reported list price has no interval, so the gap ($0.075 vs $0.10, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["groq-vs-siliconflow","groq","siliconflow","Groq vs SiliconFlow","Groq vs SiliconFlow: price per model","Groq vs SiliconFlow: 3 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and SiliconFlow share 3 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 3 ties; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aContext","bContext"],"$r":[["gpt-oss-120b: price per million tokens by provider (Input)",0.15,0.15,"usd","$0.15","$0.15","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Output)",0.6,0.6,"usd","$0.60","$0.60","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"],["gpt-oss-120b: price per million tokens by provider (Cache read)",0.075,0.075,"usd","$0.075","$0.075","tie","Same reported list price.","inference-provider-index","provider-prices-gpt-oss-120b","gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"]]}],["kimi-k2-5-vs-gpt-5-mini","kimi-k2-5","gpt-5-mini","Kimi K2.5 vs GPT 5 mini","Kimi K2.5 vs GPT 5 mini: measured benchmarks","Kimi K2.5 vs GPT 5 mini: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Kimi K2.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.697,0.6364,"rate","70% (23/33)","64% (21/33)","tie","The 95% intervals overlap (Kimi K2.5 53% to 83%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5266,0.8262],[0.4662,0.7781]],["Model calls per instance",56.7,20.8,"calls","56.7","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.256,0.08,"usd","$0.26","$0.080","unclear","No interval or range was recorded for either side, so the gap ($0.26 vs $0.080, 3.2x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",23,21,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["minimax-m2-5-vs-gpt-5-mini","minimax-m2-5","gpt-5-mini","MiniMax M2.5 vs GPT 5 mini","MiniMax M2.5 vs GPT 5 mini: measured benchmarks","MiniMax M2.5 vs GPT 5 mini: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","MiniMax M2.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.697,0.6364,"rate","70% (23/33)","64% (21/33)","tie","The 95% intervals overlap (MiniMax M2.5 53% to 83%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5266,0.8262],[0.4662,0.7781]],["Model calls per instance",58.4,20.8,"calls","58.4","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.107,0.08,"usd","$0.11","$0.080","unclear","No interval or range was recorded for either side, so the gap ($0.11 vs $0.080) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",23,21,"effort high · public mini-SWE-agent v2 run, same instances","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["minimax-m2-5-vs-kimi-k2-5","minimax-m2-5","kimi-k2-5","MiniMax M2.5 vs Kimi K2.5","MiniMax M2.5 vs Kimi K2.5: measured benchmarks","MiniMax M2.5 vs Kimi K2.5: 3 measured metrics from 2 studies (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","MiniMax M2.5 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.697,0.697,"rate","70% (23/33)","70% (23/33)","tie","The 95% intervals overlap (MiniMax M2.5 53% to 83%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5266,0.8262],[0.5266,0.8262]],["Model calls per instance",58.4,56.7,"calls","58.4","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.107,0.256,"usd","$0.11","$0.26","unclear","No interval or range was recorded for either side, so the gap ($0.11 vs $0.26, 2.4x) is not tested against run-to-run variation.","cost-thought-experiments",23,"cost-per-resolved-agent-vs-panel",23,23,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}],["agent-harness-vs-claude-haiku-4-5","agent-harness","claude-haiku-4-5","Agent vs Claude Haiku 4.5","Agent vs Claude Haiku 4.5: measured benchmarks","Agent vs Claude Haiku 4.5: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and Claude Haiku 4.5 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; Claude Haiku 4.5 ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7576,"rate","76% (25/33)","76% (25/33)","tie","The 95% intervals overlap (Agent 59% to 87%; Claude Haiku 4.5 59% to 87%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5898,0.8717],"\u0001"],["Model calls per instance",49.5,68.5,"calls","49.5","68.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.479,"usd","$3.71","$0.48","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.48, 7.7x) is not tested against run-to-run variation.","cost-thought-experiments",25,"cost-per-resolved-agent-vs-panel",25,25,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-claude-opus-4-5","agent-harness","claude-opus-4-5","Agent vs Claude Opus 4.5","Agent vs Claude Opus 4.5: measured benchmarks","Agent vs Claude Opus 4.5: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and Claude Opus 4.5 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; Claude Opus 4.5 ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7273,"rate","76% (25/33)","73% (24/33)","tie","The 95% intervals overlap (Agent 59% to 87%; Claude Opus 4.5 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5578,0.8493],"\u0001"],["Model calls per instance",49.5,35.9,"calls","49.5","35.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,1.184,"usd","$3.71","$1.18","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $1.18, 3.1x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,24,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-claude-opus-4-6","agent-harness","claude-opus-4-6","Agent vs Claude Opus 4.6","Agent vs Claude Opus 4.6: measured benchmarks","Agent vs Claude Opus 4.6: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and Claude Opus 4.6 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; Claude Opus 4.6 ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Agent 59% to 87%; Claude Opus 4.6 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262],"\u0001"],["Model calls per instance",49.5,28.9,"calls","49.5","28.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.875,"usd","$3.71","$0.88","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.88, 4.2x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"full pipeline on Claude Sonnet 5.5 · notional","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-claude-sonnet-4-5","agent-harness","claude-sonnet-4-5","Agent vs Claude Sonnet 4.5","Agent vs Claude Sonnet 4.5: measured benchmarks","Agent vs Claude Sonnet 4.5: 2 measured metrics from one study, with sample sizes, intervals and every failure counted.","Agent and Claude Sonnet 4.5 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; Claude Sonnet 4.5 ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7576,"rate","76% (25/33)","76% (25/33)","tie","The 95% intervals overlap (Agent 59% to 87%; Claude Sonnet 4.5 59% to 87%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5898,0.8717],"\u0001"],["Model calls per instance",49.5,51,"calls","49.5","51","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.913,"usd","$3.71","$0.91","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.91, 4.1x) is not tested against run-to-run variation.","cost-thought-experiments",25,"cost-per-resolved-agent-vs-panel",25,25,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-deepseek-v3-2","agent-harness","deepseek-v3-2","Agent vs DeepSeek V3.2","Agent vs DeepSeek V3.2: measured benchmarks","Agent vs DeepSeek V3.2: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and DeepSeek V3.2 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; DeepSeek V3.2 ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7273,"rate","76% (25/33)","73% (24/33)","tie","The 95% intervals overlap (Agent 59% to 87%; DeepSeek V3.2 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5578,0.8493],"\u0001"],["Model calls per instance",49.5,88.2,"calls","49.5","88.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.637,"usd","$3.71","$0.64","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.64, 5.8x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,24,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-gemini-3-flash","agent-harness","gemini-3-flash","Agent vs Gemini 3 Flash","Agent vs Gemini 3 Flash: measured benchmarks","Agent vs Gemini 3 Flash: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and Gemini 3 Flash share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; Gemini 3 Flash ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.8182,"rate","76% (25/33)","82% (27/33)","tie","The 95% intervals overlap (Agent 59% to 87%; Gemini 3 Flash 66% to 91%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.6561,0.9139],"\u0001"],["Model calls per instance",49.5,54.2,"calls","49.5","54.2","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.436,"usd","$3.71","$0.44","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.44, 8.5x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,27,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-glm-5","agent-harness","glm-5","Agent vs GLM 5","Agent vs GLM 5: measured benchmarks","Agent vs GLM 5: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and GLM 5 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; GLM 5 ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.7879,"rate","76% (25/33)","79% (26/33)","tie","The 95% intervals overlap (Agent 59% to 87%; GLM 5 62% to 89%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.6225,0.8932],"\u0001"],["Model calls per instance",49.5,77.5,"calls","49.5","77.5","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.667,"usd","$3.71","$0.67","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.67, 5.6x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,26,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-gpt-5-2","agent-harness","gpt-5-2","Agent vs GPT 5.2","Agent vs GPT 5.2: measured benchmarks","Agent vs GPT 5.2: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and GPT 5.2 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; GPT 5.2 ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.8485,"rate","76% (25/33)","85% (28/33)","tie","The 95% intervals overlap (Agent 59% to 87%; GPT 5.2 69% to 93%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.6908,0.9335],"\u0001"],["Model calls per instance",49.5,35.6,"calls","49.5","35.6","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.628,"usd","$3.71","$0.63","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.63, 5.9x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,28,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-gpt-5-mini","agent-harness","gpt-5-mini","Agent vs GPT 5 mini","Agent vs GPT 5 mini: measured benchmarks","Agent vs GPT 5 mini: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and GPT 5 mini share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; GPT 5 mini ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.6364,"rate","76% (25/33)","64% (21/33)","tie","The 95% intervals overlap (Agent 59% to 87%; GPT 5 mini 47% to 78%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.4662,0.7781],"\u0001"],["Model calls per instance",49.5,20.8,"calls","49.5","20.8","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.08,"usd","$3.71","$0.080","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.080, 46x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,21,"full pipeline on Claude Sonnet 5.5 · notional","public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-kimi-k2-5","agent-harness","kimi-k2-5","Agent vs Kimi K2.5","Agent vs Kimi K2.5: measured benchmarks","Agent vs Kimi K2.5: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and Kimi K2.5 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; Kimi K2.5 ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Agent 59% to 87%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262],"\u0001"],["Model calls per instance",49.5,56.7,"calls","49.5","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.256,"usd","$3.71","$0.26","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.26, 14x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["agent-harness-vs-minimax-m2-5","agent-harness","minimax-m2-5","Agent vs MiniMax M2.5","Agent vs MiniMax M2.5: measured benchmarks","Agent vs MiniMax M2.5: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","Agent and MiniMax M2.5 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; MiniMax M2.5 ran under a different, simpler harness, so this compares systems, not models.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Agent 59% to 87%; MiniMax M2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262],"\u0001"],["Model calls per instance",49.5,58.4,"calls","49.5","58.4","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.107,"usd","$3.71","$0.11","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.11, 35x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}],["amazon-bedrock-vs-baseten","amazon-bedrock","baseten","Amazon Bedrock vs Baseten","Amazon Bedrock vs Baseten: price per model","Amazon Bedrock vs Baseten: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Baseten share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.1,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.10","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.15 vs $0.10, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.5,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.50","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.60 vs $0.50, 1.2x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock-vs-cerebras","amazon-bedrock","cerebras","Amazon Bedrock vs Cerebras","Amazon Bedrock vs Cerebras: price per model","Amazon Bedrock vs Cerebras: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Cerebras share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.35,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.35","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.15 vs $0.35, 2.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.75,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.75","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.60 vs $0.75, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock-vs-deepinfra","amazon-bedrock","deepinfra","Amazon Bedrock vs DeepInfra","Amazon Bedrock vs DeepInfra: price per model","Amazon Bedrock vs DeepInfra: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and DeepInfra share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.037,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.037","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.15 vs $0.037, 4.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.17,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.17","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.60 vs $0.17, 3.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock-vs-groq","amazon-bedrock","groq","Amazon Bedrock vs Groq","Amazon Bedrock vs Groq: price per model","Amazon Bedrock vs Groq: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Groq share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.15,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.15","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.6,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.60","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock-vs-nebius","amazon-bedrock","nebius","Amazon Bedrock vs Nebius","Amazon Bedrock vs Nebius: price per model","Amazon Bedrock vs Nebius: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Nebius share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.15,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.15","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.6,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.60","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock-vs-novita","amazon-bedrock","novita","Amazon Bedrock vs Novita AI","Amazon Bedrock vs Novita AI: price per model","Amazon Bedrock vs Novita AI: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Novita AI share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.05,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.050","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.15 vs $0.050, 3.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.25,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.25","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.60 vs $0.25, 2.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock-vs-parasail","amazon-bedrock","parasail","Amazon Bedrock vs Parasail","Amazon Bedrock vs Parasail: price per model","Amazon Bedrock vs Parasail: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Parasail share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.1,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.10","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.15 vs $0.10, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.75,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.75","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.60 vs $0.75, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock-vs-sambanova","amazon-bedrock","sambanova","Amazon Bedrock vs SambaNova","Amazon Bedrock vs SambaNova: price per model","Amazon Bedrock vs SambaNova: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and SambaNova share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.14,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.14","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.15 vs $0.14, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.95,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.95","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.60 vs $0.95, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock-vs-siliconflow","amazon-bedrock","siliconflow","Amazon Bedrock vs SiliconFlow","Amazon Bedrock vs SiliconFlow: price per model","Amazon Bedrock vs SiliconFlow: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and SiliconFlow share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.15,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.15","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.6,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.60","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock-vs-together","amazon-bedrock","together","Amazon Bedrock vs Together AI","Amazon Bedrock vs Together AI: price per model","Amazon Bedrock vs Together AI: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Amazon Bedrock and Together AI share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.15,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.15","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.6,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.60","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["cerebras-vs-nebius","cerebras","nebius","Cerebras vs Nebius","Cerebras vs Nebius: price per model","Cerebras vs Nebius: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Cerebras and Nebius share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.35,"bValue":0.15,"unit":"usd","aDisplay":"$0.35","bDisplay":"$0.15","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.35 vs $0.15, 2.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.75,"bValue":0.6,"unit":"usd","aDisplay":"$0.75","bDisplay":"$0.60","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.75 vs $0.60, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["cerebras-vs-novita","cerebras","novita","Cerebras vs Novita AI","Cerebras vs Novita AI: price per model","Cerebras vs Novita AI: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Cerebras and Novita AI share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.35,"bValue":0.05,"unit":"usd","aDisplay":"$0.35","bDisplay":"$0.050","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.35 vs $0.050, 7.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.75,"bValue":0.25,"unit":"usd","aDisplay":"$0.75","bDisplay":"$0.25","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.75 vs $0.25, 3.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["cerebras-vs-sambanova","cerebras","sambanova","Cerebras vs SambaNova","Cerebras vs SambaNova: price per model","Cerebras vs SambaNova: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Cerebras and SambaNova share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.35,"bValue":0.14,"unit":"usd","aDisplay":"$0.35","bDisplay":"$0.14","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.35 vs $0.14, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.75,"bValue":0.95,"unit":"usd","aDisplay":"$0.75","bDisplay":"$0.95","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.75 vs $0.95, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["deepinfra-vs-cerebras","deepinfra","cerebras","DeepInfra vs Cerebras","DeepInfra vs Cerebras: price per model","DeepInfra vs Cerebras: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","DeepInfra and Cerebras share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.037,"bValue":0.35,"unit":"usd","aDisplay":"$0.037","bDisplay":"$0.35","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.037 vs $0.35, 9.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.17,"bValue":0.75,"unit":"usd","aDisplay":"$0.17","bDisplay":"$0.75","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.17 vs $0.75, 4.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["fireworks-vs-nebius","fireworks","nebius","Fireworks AI vs Nebius","Fireworks AI vs Nebius: price per model","Fireworks AI vs Nebius: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Fireworks AI and Nebius share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties; each row says why.",[{"metric":"GLM 5.3: price per million tokens by provider (Input)","aValue":1.4,"bValue":1.4,"unit":"usd","aDisplay":"$1.40","bDisplay":"$1.40","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-glm-5-3","aContext":"GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"GLM 5.3: price per million tokens by provider (Output)","aValue":4.4,"bValue":4.4,"unit":"usd","aDisplay":"$4.40","bDisplay":"$4.40","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-glm-5-3","aContext":"GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["google-vertex-vs-baseten","google-vertex","baseten","Google Vertex AI vs Baseten","Google Vertex AI vs Baseten: price per model","Google Vertex AI vs Baseten: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Baseten share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.09,"bValue":0.1,"unit":"usd","aDisplay":"$0.090","bDisplay":"$0.10","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.090 vs $0.10, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.36,"bValue":0.5,"unit":"usd","aDisplay":"$0.36","bDisplay":"$0.50","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.36 vs $0.50, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["google-vertex-vs-cerebras","google-vertex","cerebras","Google Vertex AI vs Cerebras","Google Vertex AI vs Cerebras: price per model","Google Vertex AI vs Cerebras: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Cerebras share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.09,"bValue":0.35,"unit":"usd","aDisplay":"$0.090","bDisplay":"$0.35","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.090 vs $0.35, 3.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.36,"bValue":0.75,"unit":"usd","aDisplay":"$0.36","bDisplay":"$0.75","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.36 vs $0.75, 2.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["google-vertex-vs-cloudflare","google-vertex","cloudflare","Google Vertex AI vs Cloudflare Workers AI","Google Vertex AI vs Cloudflare Workers AI: price per model","Google Vertex AI vs Cloudflare Workers AI: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Cloudflare Workers AI share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"Llama 3.3 70B Instruct: price per million tokens by provider (Input)","aValue":0.72,"bValue":0.293,"unit":"usd","aDisplay":"$0.72","bDisplay":"$0.29","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.72 vs $0.29, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-llama-3-3-70b-instruct","aContext":"Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"Llama 3.3 70B Instruct: price per million tokens by provider (Output)","aValue":0.72,"bValue":2.253,"unit":"usd","aDisplay":"$0.72","bDisplay":"$2.25","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.72 vs $2.25, 3.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-llama-3-3-70b-instruct","aContext":"Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["google-vertex-vs-nebius","google-vertex","nebius","Google Vertex AI vs Nebius","Google Vertex AI vs Nebius: price per model","Google Vertex AI vs Nebius: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and Nebius share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.09,"bValue":0.15,"unit":"usd","aDisplay":"$0.090","bDisplay":"$0.15","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.090 vs $0.15, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.36,"bValue":0.6,"unit":"usd","aDisplay":"$0.36","bDisplay":"$0.60","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.36 vs $0.60, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["google-vertex-vs-siliconflow","google-vertex","siliconflow","Google Vertex AI vs SiliconFlow","Google Vertex AI vs SiliconFlow: price per model","Google Vertex AI vs SiliconFlow: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Google Vertex AI and SiliconFlow share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.09,"bValue":0.15,"unit":"usd","aDisplay":"$0.090","bDisplay":"$0.15","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.090 vs $0.15, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.36,"bValue":0.6,"unit":"usd","aDisplay":"$0.36","bDisplay":"$0.60","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.36 vs $0.60, 1.7x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["groq-vs-cloudflare","groq","cloudflare","Groq vs Cloudflare Workers AI","Groq vs Cloudflare Workers AI: price per model","Groq vs Cloudflare Workers AI: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and Cloudflare Workers AI share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"Llama 3.3 70B Instruct: price per million tokens by provider (Input)","aValue":0.59,"bValue":0.293,"unit":"usd","aDisplay":"$0.59","bDisplay":"$0.29","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.59 vs $0.29, 2.0x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-llama-3-3-70b-instruct","aContext":"Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"Llama 3.3 70B Instruct: price per million tokens by provider (Output)","aValue":0.79,"bValue":2.253,"unit":"usd","aDisplay":"$0.79","bDisplay":"$2.25","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.79 vs $2.25, 2.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-llama-3-3-70b-instruct","aContext":"Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["groq-vs-nebius","groq","nebius","Groq vs Nebius","Groq vs Nebius: price per model","Groq vs Nebius: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Groq and Nebius share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.15,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.15","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.6,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.60","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["nebius-vs-cloudflare","nebius","cloudflare","Nebius vs Cloudflare Workers AI","Nebius vs Cloudflare Workers AI: price per model","Nebius vs Cloudflare Workers AI: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Nebius and Cloudflare Workers AI share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 ties; each row says why.",[{"metric":"GLM 5.3: price per million tokens by provider (Input)","aValue":1.4,"bValue":1.4,"unit":"usd","aDisplay":"$1.40","bDisplay":"$1.40","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-glm-5-3","aContext":"fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"GLM 5.3: price per million tokens by provider (Output)","aValue":4.4,"bValue":4.4,"unit":"usd","aDisplay":"$4.40","bDisplay":"$4.40","winner":"tie","basis":"Same reported list price.","studySlug":"inference-provider-index","chartId":"provider-prices-glm-5-3","aContext":"fp4 · GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"GLM 5.3 · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["sambanova-vs-baseten","sambanova","baseten","SambaNova vs Baseten","SambaNova vs Baseten: price per model","SambaNova vs Baseten: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","SambaNova and Baseten share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.14,"bValue":0.1,"unit":"usd","aDisplay":"$0.14","bDisplay":"$0.10","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.14 vs $0.10, 1.4x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.95,"bValue":0.5,"unit":"usd","aDisplay":"$0.95","bDisplay":"$0.50","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.95 vs $0.50, 1.9x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["sambanova-vs-cloudflare","sambanova","cloudflare","SambaNova vs Cloudflare Workers AI","SambaNova vs Cloudflare Workers AI: price per model","SambaNova vs Cloudflare Workers AI: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","SambaNova and Cloudflare Workers AI share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"Llama 3.3 70B Instruct: price per million tokens by provider (Input)","aValue":0.45,"bValue":0.293,"unit":"usd","aDisplay":"$0.45","bDisplay":"$0.29","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.45 vs $0.29, 1.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-llama-3-3-70b-instruct","aContext":"Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"Llama 3.3 70B Instruct: price per million tokens by provider (Output)","aValue":0.9,"bValue":2.253,"unit":"usd","aDisplay":"$0.90","bDisplay":"$2.25","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.90 vs $2.25, 2.5x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-llama-3-3-70b-instruct","aContext":"Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["sambanova-vs-nebius","sambanova","nebius","SambaNova vs Nebius","SambaNova vs Nebius: price per model","SambaNova vs Nebius: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","SambaNova and Nebius share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.14,"bValue":0.15,"unit":"usd","aDisplay":"$0.14","bDisplay":"$0.15","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.14 vs $0.15, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.95,"bValue":0.6,"unit":"usd","aDisplay":"$0.95","bDisplay":"$0.60","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.95 vs $0.60, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["sambanova-vs-siliconflow","sambanova","siliconflow","SambaNova vs SiliconFlow","SambaNova vs SiliconFlow: price per model","SambaNova vs SiliconFlow: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","SambaNova and SiliconFlow share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.14,"bValue":0.15,"unit":"usd","aDisplay":"$0.14","bDisplay":"$0.15","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.14 vs $0.15, 1.1x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.95,"bValue":0.6,"unit":"usd","aDisplay":"$0.95","bDisplay":"$0.60","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.95 vs $0.60, 1.6x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["together-vs-cerebras","together","cerebras","Together AI vs Cerebras","Together AI vs Cerebras: price per model","Together AI vs Cerebras: 2 list prices per million tokens for the same models, as reported by OpenRouter’s public API, with the gap stated.","Together AI and Cerebras share 2 reported list prices from one study. No row names a winner: a reported list price has no interval, so each gap is stated, not ranked. Prices change often; the snapshot date is in each row’s context. The rows are 2 unclear; each row says why.",[{"metric":"gpt-oss-120b: price per million tokens by provider (Input)","aValue":0.15,"bValue":0.35,"unit":"usd","aDisplay":"$0.15","bDisplay":"$0.35","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.15 vs $0.35, 2.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"},{"metric":"gpt-oss-120b: price per million tokens by provider (Output)","aValue":0.6,"bValue":0.75,"unit":"usd","aDisplay":"$0.60","bDisplay":"$0.75","winner":"unclear","basis":"A reported list price has no interval, so the gap ($0.60 vs $0.75, 1.3x) is stated, not ranked. Check quantization and context before treating the two as equal products.","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","aContext":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06","bContext":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]]]},"effortComparisons":{"$k":["slug","entity","aEffort","bEffort","title","seoTitle","description","verdict","rows"],"$r":[["claude-sonnet-5-5-low-vs-medium","claude-sonnet-5-5","low","medium","Claude Sonnet 5.5: low vs medium effort","Claude Sonnet 5.5: low vs medium effort, measured","Claude Sonnet 5.5 at low vs medium effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at low effort and Claude Sonnet 5.5 at medium effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at low effort 81% to 100%; Claude Sonnet 5.5 at medium effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",5.82,7.63,"seconds","5.82 s","7.63 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at low effort 2.78 s to 20.0 s; Claude Sonnet 5.5 at medium effort 2.71 s to 24.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","range","minmax",[2.78,19.96],[2.71,24.01],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",667,770,"tokens","667","770","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01219,0.01352,"usd","$0.012","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.012 vs $0.014) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.004317,0.005946,"usd","$0.0043","$0.0059","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012191,0.01352,"usd","$0.012","$0.014","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort medium","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-low-vs-high","claude-sonnet-5-5","low","high","Claude Sonnet 5.5: low vs high effort","Claude Sonnet 5.5: low vs high effort, measured","Claude Sonnet 5.5 at low vs high effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at low effort and Claude Sonnet 5.5 at high effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at low effort 81% to 100%; Claude Sonnet 5.5 at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",5.82,8.81,"seconds","5.82 s","8.81 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at low effort 2.78 s to 20.0 s; Claude Sonnet 5.5 at high effort 2.93 s to 35.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","range","minmax",[2.78,19.96],[2.93,35.81],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",667,1192,"tokens","667","1,192","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01219,0.01671,"usd","$0.012","$0.017","unclear","No interval or range was recorded for either side, so the gap ($0.012 vs $0.017) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.004317,0.009369,"usd","$0.0043","$0.0094","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012191,0.016705,"usd","$0.012","$0.017","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-low-vs-default","claude-sonnet-5-5","low","default","Claude Sonnet 5.5: low vs default effort","Claude Sonnet 5.5: low vs default effort, measured","Claude Sonnet 5.5 at low vs default effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at low effort and Claude Sonnet 5.5 at default effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at low effort 81% to 100%; Claude Sonnet 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",5.82,7.97,"seconds","5.82 s","7.97 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at low effort 2.78 s to 20.0 s; Claude Sonnet 5.5 at default effort 2.26 s to 21.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[2.78,19.96],[2.26,21.61],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",667,1054,"tokens","667","1,054","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01219,0.01398,"usd","$0.012","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.012 vs $0.014) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.004317,0.006299,"usd","$0.0043","$0.0063","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012191,0.013978,"usd","$0.012","$0.014","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-medium-vs-high","claude-sonnet-5-5","medium","high","Claude Sonnet 5.5: medium vs high effort","Claude Sonnet 5.5: medium vs high effort, measured","Claude Sonnet 5.5 at medium vs high effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at medium effort and Claude Sonnet 5.5 at high effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at medium effort 81% to 100%; Claude Sonnet 5.5 at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.63,8.81,"seconds","7.63 s","8.81 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at medium effort 2.71 s to 24.0 s; Claude Sonnet 5.5 at high effort 2.93 s to 35.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","range","minmax",[2.71,24.01],[2.93,35.81],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",770,1192,"tokens","770","1,192","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01352,0.01671,"usd","$0.014","$0.017","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.017) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005946,0.009369,"usd","$0.0059","$0.0094","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.01352,0.016705,"usd","$0.014","$0.017","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-medium-vs-default","claude-sonnet-5-5","medium","default","Claude Sonnet 5.5: medium vs default effort","Claude Sonnet 5.5: medium vs default effort, measured","Claude Sonnet 5.5 at medium vs default effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at medium effort and Claude Sonnet 5.5 at default effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at medium effort 81% to 100%; Claude Sonnet 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.63,7.97,"seconds","7.63 s","7.97 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at medium effort 2.71 s to 24.0 s; Claude Sonnet 5.5 at default effort 2.26 s to 21.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[2.71,24.01],[2.26,21.61],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",770,1054,"tokens","770","1,054","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01352,0.01398,"usd","$0.014","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.014) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005946,0.006299,"usd","$0.0059","$0.0063","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.01352,0.013978,"usd","$0.014","$0.014","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-high-vs-default","claude-sonnet-5-5","high","default","Claude Sonnet 5.5: high vs default effort","Claude Sonnet 5.5: high vs default effort, measured","Claude Sonnet 5.5 at high vs default effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at high effort and Claude Sonnet 5.5 at default effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at high effort 81% to 100%; Claude Sonnet 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",8.81,7.97,"seconds","8.81 s","7.97 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at high effort 2.93 s to 35.8 s; Claude Sonnet 5.5 at default effort 2.26 s to 21.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[2.93,35.81],[2.26,21.61],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",1192,1054,"tokens","1,192","1,054","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01671,0.01398,"usd","$0.017","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.017 vs $0.014) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.009369,0.006299,"usd","$0.0094","$0.0063","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.016705,0.013978,"usd","$0.017","$0.014","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-low-vs-medium","claude-opus-5-5","low","medium","Claude Opus 5.5: low vs medium effort","Claude Opus 5.5: low vs medium effort, measured","Claude Opus 5.5 at low vs medium effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at low effort and Claude Opus 5.5 at medium effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 81% to 100%; Claude Opus 5.5 at medium effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.5,9.72,"seconds","7.50 s","9.72 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 3.34 s to 15.8 s; Claude Opus 5.5 at medium effort 4.78 s to 31.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","range","minmax",[3.34,15.82],[4.78,31.36],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",594,853,"tokens","594","853","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02115,0.02947,"usd","$0.021","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.021 vs $0.029) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005031,0.01344,"usd","$0.0050","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.021152,0.029475,"usd","$0.021","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort medium","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-low-vs-high","claude-opus-5-5","low","high","Claude Opus 5.5: low vs high effort","Claude Opus 5.5: low vs high effort, measured","Claude Opus 5.5 at low vs high effort: 9 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at low effort and Claude Opus 5.5 at high effort share 9 measured metrics and 5 list-price calculations from 3 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 11 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 80% to 100%; Claude Opus 5.5 at high effort 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.83,2.71,"seconds","2.83 s","2.71 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 2.35 s to 6.62 s; Claude Opus 5.5 at high effort 2.45 s to 11.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","range","minmax",[2.35,6.62],[2.45,11.78],"\u0001"],["Time to first useful output",2.39,2.04,"seconds","2.39 s","2.04 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 1.45 s to 4.90 s; Claude Opus 5.5 at high effort 1.40 s to 9.94 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","range","minmax",[1.45,4.9],[1.4,9.94],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1463,1463,"tokens","1,463","1,463","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",618,619,"tokens","618","619","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",64,78,"tokens","64","78","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00688,0.00694,"usd","$0.0069","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort $0.0058 to $0.018; Claude Opus 5.5 at high effort $0.0059 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","range","minmax",[0.00582,0.01793],[0.00592,0.02708],true],["List-price cost per passing answer (calculation)",0.00829,0.01049,"usd","$0.0083","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0083 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 81% to 100%; Claude Opus 5.5 at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.5,10.11,"seconds","7.50 s","10.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 3.34 s to 15.8 s; Claude Opus 5.5 at high effort 3.63 s to 63.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","range","minmax",[3.34,15.82],[3.63,63],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",594,1052,"tokens","594","1,052","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02115,0.03368,"usd","$0.021","$0.034","unclear","No interval or range was recorded for either side, so the gap ($0.021 vs $0.034) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005031,0.018034,"usd","$0.0050","$0.018","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.021152,0.033677,"usd","$0.021","$0.034","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-low-vs-default","claude-opus-5-5","low","default","Claude Opus 5.5: low vs default effort","Claude Opus 5.5: low vs default effort, measured","Claude Opus 5.5 at low vs default effort: 9 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at low effort and Claude Opus 5.5 at default effort share 9 measured metrics and 5 list-price calculations from 3 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 11 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 80% to 100%; Claude Opus 5.5 at default effort 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.83,2.75,"seconds","2.83 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 2.35 s to 6.62 s; Claude Opus 5.5 at default effort 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.35,6.62],[2.47,8.91],"\u0001"],["Time to first useful output",2.39,1.92,"seconds","2.39 s","1.92 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 1.45 s to 4.90 s; Claude Opus 5.5 at default effort 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[1.45,4.9],[1.56,7.23],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1463,1401,"tokens","1,463","1,401","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",618,680,"tokens","618","680","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",64,64,"tokens","64","64","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00688,0.00688,"usd","$0.0069","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort $0.0058 to $0.018; Claude Opus 5.5 at default effort $0.0059 to $0.022); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00582,0.01793],[0.00592,0.02226],true],["List-price cost per passing answer (calculation)",0.00829,0.01009,"usd","$0.0083","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0083 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 81% to 100%; Claude Opus 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.5,9.18,"seconds","7.50 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 3.34 s to 15.8 s; Claude Opus 5.5 at default effort 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[3.34,15.82],[4.24,27.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",594,945,"tokens","594","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02115,0.02893,"usd","$0.021","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.021 vs $0.029) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005031,0.013104,"usd","$0.0050","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.021152,0.028925,"usd","$0.021","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-medium-vs-high","claude-opus-5-5","medium","high","Claude Opus 5.5: medium vs high effort","Claude Opus 5.5: medium vs high effort, measured","Claude Opus 5.5 at medium vs high effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at medium effort and Claude Opus 5.5 at high effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at medium effort 81% to 100%; Claude Opus 5.5 at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",9.72,10.11,"seconds","9.72 s","10.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at medium effort 4.78 s to 31.4 s; Claude Opus 5.5 at high effort 3.63 s to 63.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","range","minmax",[4.78,31.36],[3.63,63],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",853,1052,"tokens","853","1,052","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02947,0.03368,"usd","$0.029","$0.034","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.034) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.01344,0.018034,"usd","$0.013","$0.018","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.029475,0.033677,"usd","$0.029","$0.034","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-medium-vs-default","claude-opus-5-5","medium","default","Claude Opus 5.5: medium vs default effort","Claude Opus 5.5: medium vs default effort, measured","Claude Opus 5.5 at medium vs default effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at medium effort and Claude Opus 5.5 at default effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at medium effort 81% to 100%; Claude Opus 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",9.72,9.18,"seconds","9.72 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at medium effort 4.78 s to 31.4 s; Claude Opus 5.5 at default effort 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[4.78,31.36],[4.24,27.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",853,945,"tokens","853","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02947,0.02893,"usd","$0.029","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.029) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.01344,0.013104,"usd","$0.013","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.029475,0.028925,"usd","$0.029","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-high-vs-default","claude-opus-5-5","high","default","Claude Opus 5.5: high vs default effort","Claude Opus 5.5: high vs default effort, measured","Claude Opus 5.5 at high vs default effort: 14 measured metrics from 3 studies, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at high effort and Claude Opus 5.5 at default effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 4 ties and 22 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 at high effort 80% to 100%; Claude Opus 5.5 at default effort 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.71,2.75,"seconds","2.71 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 2.45 s to 11.8 s; Claude Opus 5.5 at default effort 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.45,11.78],[2.47,8.91],"\u0001"],["Time to first useful output",2.04,1.92,"seconds","2.04 s","1.92 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 1.40 s to 9.94 s; Claude Opus 5.5 at default effort 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[1.4,9.94],[1.56,7.23],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1463,1401,"tokens","1,463","1,401","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",619,680,"tokens","619","680","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",78,64,"tokens","78","64","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00694,0.00688,"usd","$0.0069","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort $0.0059 to $0.027; Claude Opus 5.5 at default effort $0.0059 to $0.022); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00592,0.02708],[0.00592,0.02226],true],["List-price cost per passing answer (calculation)",0.01049,0.01009,"usd","$0.010","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Opus 5.5 at high effort 86% to 100%; Claude Opus 5.5 at default effort 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Opus 5.5 at high effort 86% to 100%; Claude Opus 5.5 at default effort 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",11.03,9.18,"seconds","11.0 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 3.63 s to 63.0 s; Claude Opus 5.5 at default effort 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[3.63,63],[4.24,27.21],"\u0001"],["Time to first useful output on hard tasks",7.13,6.78,"seconds","7.13 s","6.78 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 2.15 s to 56.2 s; Claude Opus 5.5 at default effort 2.39 s to 21.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[2.15,56.23],[2.39,21.77],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1052,945,"tokens","1,052","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.03337,0.02824,"usd","$0.033","$0.028","unclear","No interval or range was recorded for either side, so the gap ($0.033 vs $0.028) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at high effort 81% to 100%; Claude Opus 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",10.11,9.18,"seconds","10.1 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 3.63 s to 63.0 s; Claude Opus 5.5 at default effort 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[3.63,63],[4.24,27.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",1052,945,"tokens","1,052","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.03368,0.02893,"usd","$0.034","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.034 vs $0.029) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",54.43,54.79,"percent","54.4%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code · effort high","Claude Code","range","minmax",[36.14,96.23],[29.92,95.6],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.017969,0.012528,"usd","$0.018","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.00799,0.008003,"usd","$0.0080","$0.0080","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.007407,0.00771,"usd","$0.0074","$0.0077","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.018034,0.013104,"usd","$0.018","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.033677,0.028925,"usd","$0.034","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.43,54.79,"percent","54.4%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code · effort high","Claude Code","range","minmax",[36.14,96.23],[29.92,95.6],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",43.59,0,"percent","43.6%","0%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code · effort high","Claude Code","range","minmax",[0,93.33],[0,93.33],true]]}],["gpt-6-1-sol-codex-cli-low-vs-medium","gpt-6-1-sol-codex-cli","low","medium","GPT-6.1 Sol (Codex CLI): low vs medium effort","GPT-6.1 Sol (Codex CLI): low vs medium effort, measured","GPT-6.1 Sol (Codex CLI) at low vs medium effort: 9 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (Codex CLI) at low effort and GPT-6.1 Sol (Codex CLI) at medium effort share 9 measured metrics and 5 list-price calculations from 3 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 11 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 10 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation","n"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (10/10)","100% (15/15)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at low effort 72% to 100%; GPT-6.1 Sol (Codex CLI) at medium effort 80% to 100%), so this sample cannot separate them.","model-head-to-head","h2h-pass-rate",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.7225,1],[0.7961,1],"\u0001","\u0001"],["Total time per call",6.26,5.65,"seconds","6.26 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 4.65 s to 10.5 s; GPT-6.1 Sol (Codex CLI) at medium effort 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-total-latency",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[4.65,10.47],[4.1,25.46],"\u0001","\u0001"],["Time to first useful output",5.14,5.05,"seconds","5.14 s","5.05 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 4.02 s to 8.50 s; GPT-6.1 Sol (Codex CLI) at medium effort 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-first-useful-latency",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[4.02,8.5],[3.36,17.82],"\u0001","\u0001"],["Input tokens per call: what the CLI sends (Cache read)",8064,5180,"tokens","8,064","5,180","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head","h2h-input-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",4059,6943,"tokens","4,059","6,943","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head","h2h-input-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",42,42,"tokens","42","42","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head","h2h-output-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00769,0.01018,"usd","$0.0077","$0.010","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort $0.0074 to $0.026; GPT-6.1 Sol (Codex CLI) at medium effort $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-list-price-per-call",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.00742,0.02649],[0.0054,0.02686],true,"\u0001"],["List-price cost per passing answer (calculation)",0.00998,0.01564,"usd","$0.010","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.016) is not tested against run-to-run variation.","model-head-to-head","h2h-cost-per-pass",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true,"\u0001"],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at low effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%), so this sample cannot separate them.","effort-ladder","effort-ladder-pass-rate",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001",16],["Total time per call by effort on hard tasks",13.62,13.11,"seconds","13.6 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 7.94 s to 44.3 s; GPT-6.1 Sol (Codex CLI) at medium effort 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder","effort-ladder-total-latency",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","range","minmax",[7.94,44.29],[8.54,61.6],"\u0001",16],["Output tokens per call by effort on hard tasks (Output tokens)",284,335,"tokens","284","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder","effort-ladder-output-tokens",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001",16],["List-price cost per strict pass by effort (calculation)",0.01284,0.02564,"usd","$0.013","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.013 vs $0.026) is not tested against run-to-run variation.","effort-ladder","effort-ladder-cost-per-pass",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true,16],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.001223,0.002273,"usd","$0.0012","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","thinking-bill-by-effort",16,16,"Codex CLI · effort low","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true,16],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012837,0.025637,"usd","$0.013","$0.026","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","thinking-bill-by-effort",16,16,"Codex CLI · effort low","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true,16]]}],["gpt-6-1-sol-codex-cli-low-vs-high","gpt-6-1-sol-codex-cli","low","high","GPT-6.1 Sol (Codex CLI): low vs high effort","GPT-6.1 Sol (Codex CLI): low vs high effort, measured","GPT-6.1 Sol (Codex CLI) at low vs high effort: 14 measured metrics from 3 studies, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (Codex CLI) at low effort and GPT-6.1 Sol (Codex CLI) at high effort share 14 measured metrics and 5 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 16 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation","n"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (10/10)","100% (15/15)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at low effort 72% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 80% to 100%), so this sample cannot separate them.","model-head-to-head","h2h-pass-rate",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","ci95","ci95",[0.7225,1],[0.7961,1],"\u0001","\u0001"],["Total time per call",6.26,5.6,"seconds","6.26 s","5.60 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 4.65 s to 10.5 s; GPT-6.1 Sol (Codex CLI) at high effort 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-total-latency",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[4.65,10.47],[4.05,19.52],"\u0001","\u0001"],["Time to first useful output",5.14,5.32,"seconds","5.14 s","5.32 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 4.02 s to 8.50 s; GPT-6.1 Sol (Codex CLI) at high effort 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-first-useful-latency",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[4.02,8.5],[3.64,16.37],"\u0001","\u0001"],["Input tokens per call: what the CLI sends (Cache read)",8064,6716,"tokens","8,064","6,716","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head","h2h-input-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",4059,5406,"tokens","4,059","5,406","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head","h2h-input-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",42,42,"tokens","42","42","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head","h2h-output-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00769,0.01047,"usd","$0.0077","$0.010","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort $0.0074 to $0.026; GPT-6.1 Sol (Codex CLI) at high effort $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-list-price-per-call",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[0.00742,0.02649],[0.0066,0.02812],true,"\u0001"],["List-price cost per passing answer (calculation)",0.00998,0.01322,"usd","$0.010","$0.013","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.013) is not tested against run-to-run variation.","model-head-to-head","h2h-cost-per-pass",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true,"\u0001"],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at low effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder","effort-ladder-pass-rate",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001",16],["Total time per call by effort on hard tasks",13.62,18.12,"seconds","13.6 s","18.1 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 7.94 s to 44.3 s; GPT-6.1 Sol (Codex CLI) at high effort 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder","effort-ladder-total-latency",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","range","minmax",[7.94,44.29],[11.67,92.21],"\u0001",16],["Output tokens per call by effort on hard tasks (Output tokens)",284,436,"tokens","284","436","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder","effort-ladder-output-tokens",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001",16],["List-price cost per strict pass by effort (calculation)",0.01284,0.01514,"usd","$0.013","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.013 vs $0.015) is not tested against run-to-run variation.","effort-ladder","effort-ladder-cost-per-pass",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true,16],["CLI vs API: time for a one-line answer (Total time)",4.18,4.19,"seconds","4.18 s","4.19 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 3.86 s to 4.53 s; GPT-6.1 Sol (Codex CLI) at high effort 3.81 s to 4.69 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens","cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort low · fixed exact reply, 5 runs","Codex CLI · effort high · fixed exact reply, 5 runs","range","minmax",[3.86,4.53],[3.81,4.69],"\u0001",5],["CLI vs API: time for a one-line answer (First useful output)",3.75,3.79,"seconds","3.75 s","3.79 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 3.44 s to 4.10 s; GPT-6.1 Sol (Codex CLI) at high effort 3.37 s to 4.30 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens","cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort low · fixed exact reply, 5 runs","Codex CLI · effort high · fixed exact reply, 5 runs","range","minmax",[3.44,4.1],[3.37,4.3],"\u0001",5],["CLI vs API: time for a small coding task (Total time)",14.15,17.85,"seconds","14.2 s","17.9 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) at low effort 13.0 s to 14.4 s; GPT-6.1 Sol (Codex CLI) at high effort 17.7 s to 22.4 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens","cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort low · small coding task, 3 runs","Codex CLI · effort high · small coding task, 3 runs","range","minmax",[13.02,14.41],[17.68,22.42],"\u0001",3],["CLI vs API: time for a small coding task (First useful output)",13.6,17.27,"seconds","13.6 s","17.3 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) at low effort 12.5 s to 13.8 s; GPT-6.1 Sol (Codex CLI) at high effort 17.1 s to 21.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens","cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort low · small coding task, 3 runs","Codex CLI · effort high · small coding task, 3 runs","range","minmax",[12.52,13.83],[17.13,21.86],"\u0001",3],["Hidden prompt: input tokens for the same one-line request",19551,19555,"tokens","19,551","19,555","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens","cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort low · short fixed tasks","Codex CLI · effort high · short fixed tasks","\u0001","\u0001","\u0001","\u0001","\u0001",5],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.001223,0.004114,"usd","$0.0012","$0.0041","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","thinking-bill-by-effort",16,16,"Codex CLI · effort low","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true,16],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012837,0.015137,"usd","$0.013","$0.015","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","thinking-bill-by-effort",16,16,"Codex CLI · effort low","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true,16]]}],["gpt-6-1-sol-codex-cli-medium-vs-high","gpt-6-1-sol-codex-cli","medium","high","GPT-6.1 Sol (Codex CLI): medium vs high effort","GPT-6.1 Sol (Codex CLI): medium vs high effort, measured","GPT-6.1 Sol (Codex CLI) at medium vs high effort: 14 measured metrics from 3 studies, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (Codex CLI) at medium effort and GPT-6.1 Sol (Codex CLI) at high effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 5 ties and 21 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 80% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",5.65,5.6,"seconds","5.65 s","5.60 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 4.10 s to 25.5 s; GPT-6.1 Sol (Codex CLI) at high effort 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[4.1,25.46],[4.05,19.52],"\u0001"],["Time to first useful output",5.05,5.32,"seconds","5.05 s","5.32 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 3.36 s to 17.8 s; GPT-6.1 Sol (Codex CLI) at high effort 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[3.36,17.82],[3.64,16.37],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",5180,6716,"tokens","5,180","6,716","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",6943,5406,"tokens","6,943","5,406","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",42,42,"tokens","42","42","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head",15,"h2h-output-tokens",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.01018,0.01047,"usd","$0.010","$0.010","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort $0.0054 to $0.027; GPT-6.1 Sol (Codex CLI) at high effort $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[0.0054,0.02686],[0.0066,0.02812],true],["List-price cost per passing answer (calculation)",0.01564,0.01322,"usd","$0.016","$0.013","unclear","No interval or range was recorded for either side, so the gap ($0.016 vs $0.013) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head",16,"hard-h2h-pass-rate",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head",16,"hard-h2h-pass-rate",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",13.11,18.12,"seconds","13.1 s","18.1 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 8.54 s to 61.6 s; GPT-6.1 Sol (Codex CLI) at high effort 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",16,"hard-h2h-total-latency",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","range","minmax",[8.54,61.6],[11.67,92.21],"\u0001"],["Time to first useful output on hard tasks",10.23,12.69,"seconds","10.2 s","12.7 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 6.09 s to 40.4 s; GPT-6.1 Sol (Codex CLI) at high effort 8.93 s to 75.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",16,"hard-h2h-first-useful-latency",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","range","minmax",[6.09,40.41],[8.93,75.91],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",335,436,"tokens","335","436","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",16,"hard-h2h-output-tokens",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.02564,0.01514,"usd","$0.026","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.026 vs $0.015) is not tested against run-to-run variation.","hard-model-head-to-head",16,"hard-h2h-cost-per-pass",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Codex CLI · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",13.11,18.12,"seconds","13.1 s","18.1 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 8.54 s to 61.6 s; GPT-6.1 Sol (Codex CLI) at high effort 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Codex CLI · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","range","minmax",[8.54,61.6],[11.67,92.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",335,436,"tokens","335","436","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Codex CLI · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02564,0.01514,"usd","$0.026","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.026 vs $0.015) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Codex CLI · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",46.33,57.01,"percent","46.3%","57%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-share",16,16,"Codex CLI · effort medium","Codex CLI · effort high","range","minmax",[11.42,86.85],[29.19,90.8],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.002273,0.004114,"usd","$0.0023","$0.0041","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-cost-per-call",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.003025,0.002905,"usd","$0.0030","$0.0029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-cost-per-call",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.020339,0.008117,"usd","$0.020","$0.0081","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-cost-per-call",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.002273,0.004114,"usd","$0.0023","$0.0041","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.025637,0.015137,"usd","$0.026","$0.015","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",46.33,57.01,"percent","46.3%","57%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-short-vs-hard",16,16,"Codex CLI · effort medium","Codex CLI · effort high","range","minmax",[11.42,86.85],[29.19,90.8],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",41.05,58.06,"percent","41%","58.1%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Codex CLI · effort medium","Codex CLI · effort high","range","minmax",[0,71.43],[0,75.76],true]]}],["gpt-6-1-sol-openai-api-low-vs-high","gpt-6-1-sol-openai-api","low","high","GPT-6.1 Sol (OpenAI API): low vs high effort","GPT-6.1 Sol (OpenAI API): low vs high effort, measured","GPT-6.1 Sol (OpenAI API) at low vs high effort: 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (OpenAI API) at low effort and GPT-6.1 Sol (OpenAI API) at high effort share 5 measured metrics from one study. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 4 unclear; each row says why. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",1.02,1.52,"seconds","1.02 s","1.52 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.96 s to 1.87 s; GPT-6.1 Sol (OpenAI API) at high effort 1.35 s to 2.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort low · fixed exact reply, 5 runs","OpenAI API · effort high · fixed exact reply, 5 runs","range","minmax",[0.96,1.87],[1.35,2.23]],["CLI vs API: time for a one-line answer (First useful output)",0.87,1.34,"seconds","0.87 s","1.34 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.84 s to 1.74 s; GPT-6.1 Sol (OpenAI API) at high effort 1.26 s to 2.12 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort low · fixed exact reply, 5 runs","OpenAI API · effort high · fixed exact reply, 5 runs","range","minmax",[0.84,1.74],[1.26,2.12]],["CLI vs API: time for a small coding task (Total time)",6,9.56,"seconds","6.00 s","9.56 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) at low effort 5.44 s to 6.20 s; GPT-6.1 Sol (OpenAI API) at high effort 9.44 s to 10.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort low · small coding task, 3 runs","OpenAI API · effort high · small coding task, 3 runs","range","minmax",[5.44,6.2],[9.44,10.94]],["CLI vs API: time for a small coding task (First useful output)",1.05,5.31,"seconds","1.05 s","5.31 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.97 s to 1.40 s; GPT-6.1 Sol (OpenAI API) at high effort 4.99 s to 6.42 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort low · small coding task, 3 runs","OpenAI API · effort high · small coding task, 3 runs","range","minmax",[0.97,1.4],[4.99,6.42]],["Hidden prompt: input tokens for the same one-line request",17,17,"tokens","17","17","tie","Same value. More or fewer is not better by itself for this metric.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"OpenAI API · effort low · short fixed tasks","OpenAI API · effort high · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}]]}}