{"schema":"agent-public-bench@1","generatedAt":"2026-10-07T00:00:00.000Z","sources":{"$k":["id","title","kind","date","note","data","url"],"$r":[["agent-provider-h2h","Provider head-to-head: Claude Code models vs Codex efforts","run","2026-10-05","Five short tasks with deterministic validators, declared protocol, every attempt kept.",["/benchmarks/raw/provider-h2h/receipts.json"],"\u0001"],["agent-provider-h2h-hard","Provider head-to-head, hard set: eight hard tasks with strict validators","run","2026-10-06","Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.",["/benchmarks/raw/provider-h2h-hard/receipts.json"],"\u0001"],["agent-effort-ladder","Effort ladder: the hard task set at each effort level","run","2026-10-06","The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.",["/benchmarks/raw/effort-ladder/receipts.json","/benchmarks/raw/provider-h2h-hard/receipts.json"],"\u0001"],["agent-coding-agents","Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks","run","2026-10-06","Six small Node.js repositories with hidden tests; controls before the first session (every base fails, every reference passes). Claude Code 2.1.286 with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol at medium effort, 2 repetitions per task, OS sandboxes without network, protocol declared before the first session, every session kept. Gemini CLI was probed and not run (browser login).",["/benchmarks/raw/coding-agents/sessions.json"],"\u0001"],["agent-caching-consistency","Caching sessions and repeated prompts (Claude Code and Codex CLI)","run","2026-10-06","Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.",["/benchmarks/raw/caching-consistency/caching.json","/benchmarks/raw/caching-consistency/consistency.json"],"\u0001"],["agent-routing","Routing runs: Jev router vs LLM routing","run","2026-10-05","Routing decisions recorded per case and arm.",["/benchmarks/raw/routing/receipts.json"],"\u0001"],["agent-jev-live","Jev live run: 246 timed calls on the 82 routing decisions","run","2026-10-06","Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.",["/benchmarks/raw/jev-live/summary.json","/benchmarks/raw/jev-live/calls.json"],"\u0001"],["price-anthropic","Anthropic list prices (Claude models)","price-list","2026-09-21","Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.","\u0001","https://platform.claude.com/docs/en/about-claude/pricing"],["price-openai","OpenAI list prices","price-list","2026-10-03","Token prices as listed by the vendor on 2026-10-03.","\u0001","https://developers.openai.com/api/docs/pricing"],["price-jev","Jev 1.13 list price","price-list","2026-09-23","Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.","\u0001","https://docs.typesafe.ai/models"],["agent-routing-overhead","Routing overhead runs: policy microbenchmark and CLI start-up","run","2026-10-06","In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.",["/benchmarks/raw/routing-overhead/results.json"],"\u0001"],["calc-repricing","Repricing calculation","calculation","2026-10-05","Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.","\u0001","\u0001"]]},"studies":{"$k":["slug","title","seoTitle","description","question","answer","date","updated","tags","method","caveats","sourceIds","stats","charts","tables","related","hero"],"$r":[["model-head-to-head","Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head","Haiku vs Sonnet vs Opus vs Fable vs Codex: speed and tokens","130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.","On short tasks with strict validators, how do the Claude Code models and efforts compare with Codex on pass rate, speed and tokens?","127 of 130 calls passed (98%, 95% interval 93% to 99%), so on these short tasks pass rate barely separates the configurations (non-passes: Claude Sonnet 5.5 · Claude Code on Multi-step shift arithmetic ×3 (\"Output does not exactly match the expected text\")). Every non-pass ended on the expected answer (2292) but added working lines, which the exact-text validator rejects by design: a format miss, not a wrong answer. Speed separates them more: the fastest configuration was Claude Fable 5.1 · Claude Code at a median 1.9 s per call, the slowest GPT-6.1 Sol (low) · Codex CLI at 6.3 s. Claude Code medians ran from 1.9 s to 4.4 s. Codex CLI medians ran from 5.6 s to 6.3 s. Single calls vary a lot (see the ranges), so neighbouring configurations are not separated. Codex CLI sends a median 12,124 input tokens per call against 2,130 for Claude Code, mostly the CLI’s own context. At list price (a calculation; the calls ran on subscriptions), the cheapest passing answer came from Claude Sonnet 5.5 · Claude Code at $0.0062. On speed, cost per pass and pass rate together, the frontier is Claude Fable 5.1 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (low) · Claude Code; with 10 to 15 calls per configuration, small gaps inside it are within the spread of single calls. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.","2026-10-05","2026-10-05",["head-to-head","claude-haiku","claude-sonnet","claude-opus","claude-fable","codex","latency"],[],[],["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["h2h-pass-rate","Pass rate on five validated tasks","Every call counts; failures and timeouts are non-passes","dot-range","rate","Passed",[{"name":"Pass rate","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Fable 5.1 · Claude Code",1,0.7961,1,15],["Claude Sonnet 5.5 · Claude Code",0.8,0.5481,0.9295,15],["Claude Opus 5.5 (high) · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 · Claude Code",1,0.7961,1,15],["Claude Opus 5.5 (low) · Claude Code",1,0.7961,1,15],["Claude Haiku 4.5 · Claude Code",1,0.7961,1,15],["GPT-6.1 Sol (high) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7961,1,15],["GPT-6.1 Sol (low) · Codex CLI",1,0.7225,1,10]]}}],"Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.",["agent-provider-h2h"]],["h2h-total-latency","Total time per call","Median per configuration; whiskers = fastest and slowest call","dot-range","seconds","Seconds",[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Fable 5.1 · Claude Code",1.94,1.41,9.83,15,true],["Claude Sonnet 5.5 · Claude Code",2.31,2.17,7.73,15,true],["Claude Opus 5.5 (high) · Claude Code",2.71,2.45,11.78,15,true],["Claude Opus 5.5 · Claude Code",2.75,2.47,8.91,15,true],["Claude Opus 5.5 (low) · Claude Code",2.83,2.35,6.62,15,true],["Claude Haiku 4.5 · Claude Code",4.43,3.16,23.57,15,true],["GPT-6.1 Sol (high) · Codex CLI",5.6,4.05,19.52,15,false],["GPT-6.1 Sol (medium) · Codex CLI",5.65,4.1,25.46,15,false],["GPT-6.1 Sol (low) · Codex CLI",6.26,4.65,10.47,10,false]]}}],"One host, one network, one day. Whiskers are a range, not a confidence interval.",["agent-provider-h2h"]],["h2h-cost-per-pass","List-price cost per passing answer (calculation)","All calls in a configuration, failures included, divided by its passes","bar","usd","USD per passing answer",[{"name":"Cost per pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.00624,15,true],["Claude Opus 5.5 (low) · Claude Code",0.00829,15,true],["Claude Haiku 4.5 · Claude Code",0.00836,15,false],["GPT-6.1 Sol (low) · Codex CLI",0.00998,10,false],["Claude Opus 5.5 · Claude Code",0.01009,15,true],["Claude Opus 5.5 (high) · Claude Code",0.01049,15,true],["GPT-6.1 Sol (high) · Codex CLI",0.01322,15,false],["GPT-6.1 Sol (medium) · Codex CLI",0.01564,15,false],["Claude Fable 5.1 · Claude Code",0.02054,15,true]]}}],"Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.",["agent-provider-h2h","calc-repricing","price-anthropic","price-openai"]]]},[],[],"\u0001"],["hard-model-head-to-head","Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks","Hard tasks: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol","152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.","On 8 hard tasks with deterministic validators, does pass rate separate the Claude Code models and GPT-6.1 Sol through the Codex CLI, and what do speed, tokens and cost per pass add?","139 of 152 calls that reached a model passed strictly (91%). 6 of 7 configurations passed every call: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Fable 5.1 · Claude Code (24/24 each, 95% interval 86% to 100%) and GPT-6.1 Sol (medium) · Codex CLI, GPT-6.1 Sol (high) · Codex CLI (16/16 each, 95% interval 81% to 100%), so the hard set still has a ceiling for these models and pass rate does not separate them. Claude Haiku 4.5 · Claude Code passed 11/24 strictly (46%, 95% interval 28% to 65%). 5 more replies had the right answer in the wrong format (for example inside a code fence), so 16/24 on a lenient reading (47% to 82%); 8 replies were wrong. It passed none of the tasks “Predict JavaScript event-loop output order”, “Solve a multi-constraint room schedule” and “Write a SQLite reporting query (fan-out, ties, boundaries)”. Median total time per call was Sonnet 7.7 s, Opus 9.2 s, Opus (high) 11.0 s, GPT-6.1 Sol (medium) 13.1 s, Fable 16.1 s, GPT-6.1 Sol (high) 18.1 s, Haiku 39.0 s. The fastest and slowest single calls of every configuration overlap with every other, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on a subscription), the lowest cost per strict pass was Claude Sonnet 5.5 · Claude Code at $0.0143; the quality-vs-cost frontier is Claude Sonnet 5.5 · Claude Code. 30 earlier attempts were blocked before any model call (Codex CLI: the CLI reported no signed-in account); they are reported, not scored, and that route ran in a later batch.","2026-10-06","2026-10-06",["head-to-head","hard-tasks","claude-haiku","claude-sonnet","claude-opus","claude-fable","gpt-6-1-sol","codex-cli","format-misses","latency"],[],[],["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["hard-h2h-pass-rate","Pass rate on eight hard tasks","Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","dot-range","rate","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.4583,0.2789,0.6493,24]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.6667,0.4671,0.8203,24]]}}],"Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.",["agent-provider-h2h-hard"]],["hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Median per configuration; whiskers = fastest and slowest call","dot-range","seconds","Seconds",[{"name":"Total time per call on hard tasks (separate batches)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",7.75,2.26,34.79,24,true],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,24,true],["Claude Opus 5.5 (high) · Claude Code",11.03,3.63,63,24,true],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16,true],["Claude Fable 5.1 · Claude Code",16.13,4.46,90,24,true],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16,true],["Claude Haiku 4.5 · Claude Code",39.01,15.27,75.13,24,false]]}}],"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.",["agent-provider-h2h-hard"]],["hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)","All calls in a configuration, failures and format misses included, divided by its strict passes","bar","usd","USD per strict pass",[{"name":"Cost per strict pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.01435,24,true],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16,false],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16,false],["Claude Opus 5.5 · Claude Code",0.02824,24,false],["Claude Opus 5.5 (high) · Claude Code",0.03337,24,false],["Claude Haiku 4.5 · Claude Code",0.0672,24,false],["Claude Fable 5.1 · Claude Code",0.09331,24,false]]}}],"Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.",["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"]]]},[],[],"\u0001"],["coding-agents-head-to-head","Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks","Claude Code vs Codex CLI: 6 coding tasks, hidden tests","36 graded sessions: Claude Code with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol. All passed every hidden test; time, tool calls and diffs differ.","On small real repository tasks graded by hidden tests, how do coding-agent CLIs compare when they run with their normal file and shell tools?","All 3 agents passed every hidden check in every session (12/12, 12/12, 12/12; 95% Wilson 76–100% each), so this task set cannot separate them on quality. Sonnet 5.5 in Claude Code was fastest (median 23.1 s; its sessions took 18.7 s to 44.5 s), Opus 5.5 in Claude Code took a median 56.9 s, GPT-6.1 Sol in Codex CLI took a median 113.4 s; the run ranges of Sonnet 5.5 in Claude Code and GPT-6.1 Sol in Codex CLI do not overlap. Codex made more tool calls (median 12.5 vs 7.5 and 7.5) and larger diffs (median 94 lines vs 45 and 77.5, mostly added tests), and every Codex session also followed the tester’s global AGENTS.md (10 of 12 wrote a work log nobody asked for), so its time and diff include extra work. Opus used 1.9× the output tokens of Sonnet in the same CLI (medians). Gemini CLI was not run: it needed a browser login.","2026-10-06","2026-10-06",["claude-code","codex-cli","coding-agents","hidden-tests","head-to-head"],[],[],["agent-coding-agents","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","subtitle","kind","unit","whisker","yLabel","viz","series","note","sourceIds"],"$r":[["coding-agents-pass-rate","Coding sessions that passed every hidden check","A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals","dot-range","rate","ci95","Passed","IntervalDotPlot",[{"name":"Passed every hidden check","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.7575,1,12],["Claude Opus 5.5 · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",1,0.7575,1,12]]}}],"6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.",["agent-coding-agents"]],["coding-agents-wall-time","Time per coding session","Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)","dot-range","seconds","minmax","Seconds","LatencyLanes",[{"name":"Wall time per session","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",23.1,18.7,44.5,12,true],["Claude Opus 5.5 · Claude Code",56.9,29.8,185.8,12,"\u0001"],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",113.4,78.5,221.9,12,"\u0001"]]}}],"CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.",["agent-coding-agents"]],["coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)","Reported tokens of all 12 sessions × list price, divided by the passes","bar","usd","\u0001","USD per pass","CostBars",[{"name":"List-price cost per pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.085,12],["Claude Opus 5.5 · Claude Code",0.2229,12],["GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",0.0978,12]]}}],"Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.",["agent-coding-agents","calc-repricing","price-anthropic","price-openai"]]]},[],[],{"statIds":["coding-agents-time-ratio"]}],["effort-ladder","Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks","Effort ladder: does more AI effort buy quality?","176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.","On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate for Sonnet, Opus and GPT-6.1 Sol, and what does it cost in time, tokens and list price per pass?","Not on this set. Every one of the 11 configurations (Sonnet at low, medium, high and default, Opus at low, medium, high and default and GPT-6.1 Sol through Codex CLI at low, medium and high) passed all 16 calls strictly (95% interval 81% to 100% each), so pass rate does not separate any effort level. The hard set has a ceiling for these models: with 16 calls per cell it cannot rule out a difference of up to about 19 points. What more effort did change is output and time. Median total time per call by effort: Sonnet: low 5.8 s, medium 7.6 s, high 8.8 s, default 8.0 s (median output tokens 667 / 770 / 1,192 / 1,054); Opus: low 7.5 s, medium 9.7 s, high 10.1 s, default 9.2 s (median output tokens 594 / 853 / 1,052 / 945); GPT-6.1 Sol (Codex CLI): low 13.6 s, medium 13.1 s, high 18.1 s (median output tokens 284 / 335 / 436). Within each model, the fastest and slowest calls of every effort overlap, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on subscriptions), cost per strict pass went from $0.0122 at low to $0.0167 at high for Sonnet, $0.0212 at low to $0.0337 at high for Opus and $0.0128 at low to $0.0151 at high for GPT-6.1 Sol (Codex CLI).","2026-10-06","2026-10-06",["effort","reasoning-effort","hard-tasks","claude-sonnet","claude-opus","gpt-6-1-sol","claude-code","codex-cli","latency","tokens"],[],[],["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"],[],{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","whisker","sourceIds"],"$r":[["effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks","Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals","dot-range","rate","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 · Claude Code",1,0.8064,1,16],["GPT-6.1 Sol (low) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16]]}}],"Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.","ci95",["agent-effort-ladder"]],["effort-ladder-total-latency","Total time per call by effort on hard tasks","Median per configuration; whiskers = fastest and slowest call","dot-range","seconds","Seconds",[{"name":"Total time per call by effort on hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",5.82,2.78,19.96,16],["Claude Sonnet 5.5 (medium) · Claude Code",7.63,2.71,24.01,16],["Claude Sonnet 5.5 (high) · Claude Code",8.81,2.93,35.81,16],["Claude Sonnet 5.5 · Claude Code",7.97,2.26,21.61,16],["Claude Opus 5.5 (low) · Claude Code",7.5,3.34,15.82,16],["Claude Opus 5.5 (medium) · Claude Code",9.72,4.78,31.36,16],["Claude Opus 5.5 (high) · Claude Code",10.11,3.63,63,16],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,16],["GPT-6.1 Sol (low) · Codex CLI",13.62,7.94,44.29,16],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16]]}}],"Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.","minmax",["agent-effort-ladder"]],["effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)","All calls in a configuration divided by its strict passes","bar","usd","USD per strict pass",[{"name":"Cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.01219,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.01352,16],["Claude Sonnet 5.5 (high) · Claude Code",0.01671,16],["Claude Sonnet 5.5 · Claude Code",0.01398,16],["Claude Opus 5.5 (low) · Claude Code",0.02115,16],["Claude Opus 5.5 (medium) · Claude Code",0.02947,16],["Claude Opus 5.5 (high) · Claude Code",0.03368,16],["Claude Opus 5.5 · Claude Code",0.02893,16],["GPT-6.1 Sol (low) · Codex CLI",0.01284,16],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16]]}}],"Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.","\u0001",["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"]]]},[],[],"\u0001"],["caching-consistency","Prompt caching and run-to-run consistency in Claude Code and Codex CLI","Prompt caching savings and LLM consistency, measured","135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.","When a CLI session reuses a fixed context, how much input comes from the cache, what does that save at list price, and does it change latency? When the same prompt runs 10 times, how much do the pass rate, the answer and the time vary?","Caching: inside one Claude Code session, turns 2-5 read 97% of their input from the cache on average (turn 1: 19%, the CLI's own prefix). At list price, a calculation, all recorded turns cost Sonnet $0.1350 vs $0.2698 (50% less) and Opus $0.2551 vs $0.5442 (53% less) without the cache. Turn 1 costs more with the cache, because a 1-hour cache write costs twice the input price. A new session did not reuse the cache of an earlier one: on turn 1, all 4 later sessions wrote the ledger to the cache again. The cause was not tested. The cache showed no clear speed effect: median turn time was Sonnet 1.6 s on turn 1 vs 1.6 s on turns 2-5 and Opus 1.9 s on turn 1 vs 2.4 s on turns 2-5, and the fastest-to-slowest ranges overlap. Codex CLI (GPT-6.1 Sol) read 99% of later-turn input from its cache on a larger context; its app-server reports no cache writes, so no cost is calculated for it. Consistency: 7 of 9 model-and-prompt cells passed all 10 repetitions (95% interval 72% to 100%). Haiku passed 0/10 on the exact-number prompt; Haiku passed 1/10 on the JSON prompt (9 more were correct but in the wrong format). Haiku gave the same wrong answer every time (289; expected 282): consistent is not the same as correct. The code-fix prompt gave 6 different code bodies for Haiku, 3 different code bodies for Sonnet and 6 different code bodies for GPT-6.1 Sol (medium).","2026-10-06","2026-10-06",["prompt-caching","consistency","variance","claude-code","codex-cli","claude-haiku","claude-sonnet","claude-opus","gpt-6-1-sol","calculation"],[],[],["agent-caching-consistency","calc-cache-pricing","price-anthropic"],[],[{"id":"consistency-pass-rate","title":"Same prompt, 10 times: strict pass rate","subtitle":"One series per prompt; whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","yLabel":"Passed","series":{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0,0,0.2775,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["JSON object",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0.1,0.0179,0.4042,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["Code fix",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,0.7225,1,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}]]},"note":"Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.","whisker":"ci95","sourceIds":["agent-caching-consistency"]},{"id":"consistency-latency-spread","title":"Same prompt, 10 times: time per call","subtitle":"Median; whiskers = fastest and slowest of 10 calls","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",5.06,4.42,6.2,10],["Claude Sonnet 5.5 · Claude Code",6.89,5.81,7.81,10],["GPT-6.1 Sol (medium) · Codex CLI",13.38,12.29,17.97,10]]}],["JSON object",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",7.03,5.28,12.27,10],["Claude Sonnet 5.5 · Claude Code",2.89,2.68,5.3,10],["GPT-6.1 Sol (medium) · Codex CLI",6.42,5.25,8.26,10]]}],["Code fix",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",5.95,4.89,7.33,10],["Claude Sonnet 5.5 · Claude Code",2.67,2.32,4.34,10],["GPT-6.1 Sol (medium) · Codex CLI",11.29,9.08,14.85,10]]}]]},"note":"Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.","whisker":"minmax","sourceIds":["agent-caching-consistency"]}],[],[],"\u0001"],["routing-jev-vs-llm","Jev vs Claude as a router: accuracy and cost","Jev router vs Claude Haiku and Sonnet: routing and cost","Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.","Should a small dedicated router or a general LLM make the platform’s typed routing decisions?","Exact decisions: Jev 1.13 (TypeSafe) 221 of 246 live calls (90%; 74, 73 and 74 of 82 per repeat; case-level interval 82% to 95%); Claude Haiku 4.5 73 of 82 (89%, 80% to 94%); Claude Sonnet 5.5 77 of 82 (94%, 87% to 97%). The intervals overlap, so accuracy does not separate the routers here. Cost does: Jev costs $0.0337 per 1,000 decisions (a calculation from its reported input tokens) against Claude Haiku 4.5 $8.92, Claude Sonnet 5.5 $5.00. Case by case against Jev’s recorded production run (one pass), the exact McNemar test finds no difference (Claude Haiku 4.5 p = 1, Claude Sonnet 5.5 p = 0.375). Per decision, Jev is about 265x cheaper than Claude Haiku 4.5 and about 148x cheaper than Claude Sonnet 5.5. Median model time per decision through the Claude Code CLI: Claude Haiku 4.5 10.7 s, Claude Sonnet 5.5 1.6 s. Jev, called directly over HTTPS from one Mac, took a median 137 ms per call (p95 196 ms, 246 calls, wall time with the network inside it). That is a different route from the CLI, so it is not a model-against-model compute comparison. The case sets were tuned against Jev answers, which gives Jev a home advantage. Thought experiment (a calculation on 2,362 recorded calls, not a run): the same tokens cost $108.54 all on Sonnet 5.5, $161.62 under the platform policy mix (1.49x) and $105.53 with Haiku on side jobs only (2.8% less), because the main coding stages hold most of the spend.","2026-10-05","2026-10-06",["routing","jev","claude-haiku","claude-sonnet","model-routing","thought-experiment"],[],[],["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"],[],{"$k":["id","title","subtitle","kind","unit","yLabel","whisker","series","note","sourceIds"],"$r":[["routing-exact-decisions","Typed routing decisions answered exactly right","Share of asked cases where every scored question was acceptable","dot-range","rate","Exact","ci95",[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8984,0.8191,0.9497,82,true],["Claude Haiku 4.5",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5",0.939,0.8651,0.9737,82,false]]}}],"Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.",["agent-routing","agent-jev-live"]],["routing-cost-per-1000","Cost per 1,000 routing decisions","List price × reported tokens per decision","bar","usd","USD per 1,000 decisions","\u0001",[{"name":"Cost","points":{"$k":["label","value","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.0337,246,true],["Claude Haiku 4.5",8.924,82,false],["Claude Sonnet 5.5",4.996,82,false]]}}],"List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.",["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"]],["routing-decision-latency","Time per routing decision","Median wall time, whisker to the 95th percentile","dot-range","ms","Time per decision","p50-p95",{"$k":["name","points"],"$r":[["Wall time (CLI)",[{"label":"Claude Haiku 4.5","value":12674,"lo":12674,"hi":34413,"n":82},{"label":"Claude Sonnet 5.5","value":2598,"lo":2598,"hi":4298,"n":82}]],["Wall time (direct API call)",[{"label":"Jev 1.13 (TypeSafe)","value":136.5,"lo":136.5,"hi":195.7,"n":246}]],["Model time (API)",[{"label":"Claude Haiku 4.5","value":10734,"lo":10734,"hi":32072,"n":82},{"label":"Claude Sonnet 5.5","value":1599,"lo":1599,"hi":2574,"n":82}]]]},"Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.",["agent-routing","agent-jev-live"]]]},[],[],"\u0001"],["routing-overhead","Routing overhead: deterministic policy vs LLM routers vs Jev","Routing overhead: rules vs LLM routers vs Jev","How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.","What delay and what cost does each kind of router add before the real work of a call starts?","The deterministic routing policy decided in a median 1.42 µs (p95 2.33 µs, 20,000 decisions, $0). The fastest LLM router, Claude Sonnet 5.5 through the Claude Code CLI, took a median 2.60 s per decision (p95 4.30 s, n = 82), about 1.8 million times longer; 973 ms of that median was CLI time, not model time. Haiku 4.5 with its default thinking took 12.54 s (p95 34.48 s). Jev 1.13, called directly over HTTPS from the same Mac, took a median 137 ms per decision (p95 196 ms, n = 246 calls, client wall time with the network inside it) and cost $0.0337 per 1,000 decisions (a calculation from its reported input tokens). That is a different route from the CLI routers, so it is not a model-against-model comparison. As a calculation over 48 recorded tasks (median 49.5 model calls each), routing every call would add $1.67 per 1,000 tasks and up to 7 s of waiting per task with Jev, $247.30 with Sonnet (8.2% of the work cost) and up to 129 s of waiting per task with Sonnet; routing only the 7 System One decisions cuts Sonnet to $34.97 and 18 s. CLI start-up alone, for a one-word answer: Claude Code (Haiku 4.5) took 2.53 s and sent 6,761 input tokens; Codex CLI took 6.00 s and sent 17,051 input tokens, 13,184 of them read from the cache (5 runs each, different models).","2026-10-06","2026-10-06",["routing","latency","overhead","jev","llm-router","cli","calculation"],[],[],["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"],[],[{"id":"router-overhead-cost-reported","title":"Cost per 1,000 routing decisions: no model call vs provider-reported","subtitle":"USD per 1,000 decisions","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost per 1,000 decisions","points":[{"label":"Deterministic routing policy (Agent, in process)","value":0,"n":20000,"highlight":true},{"label":"Jev 1.13 (TypeSafe)","value":0.0337,"n":82,"highlight":false}]}],"note":"The policy makes no model call, so it costs nothing per decision. Jev’s figure is the cost its provider reported for the recorded production run (82 decisions). The input tokens of the live run give the same figure at the published price. The Claude routers are in the next chart: their cost is derived from list prices.","sourceIds":["agent-routing-overhead","agent-routing","price-jev"]}],[],[],"\u0001"]]},"entities":{"$k":["slug","name","vendor","kind","description","aliases","facts"],"$r":[["claude-sonnet-5-5","Claude Sonnet 5.5","Anthropic","model","Anthropic’s mid-tier Claude model. Agent’s default coding model; measured here through Claude Code at low, medium, high and default effort, and as an LLM router.",["Claude Sonnet 5.5","Sonnet 5.5"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation"],"$r":[["model-head-to-head","h2h-pass-rate","Pass rate","Claude Sonnet 5.5 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",0.8,"rate","80% (12/15)",15,[0.5481,0.9295],"ci95","Claude Code · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Sonnet 5.5 · Claude Code","h2h-total-latency","Total time per call",2.31,"seconds","2.31 s",15,"\u0001","minmax","Claude Code · five short validated tasks",[2.17,7.73],"\u0001"],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Sonnet 5.5 · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.00624,"usd","$0.0062",15,"\u0001","\u0001","Claude Code · five short validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Sonnet 5.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Sonnet 5.5 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Sonnet 5.5 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",7.75,"seconds","7.75 s",24,"\u0001","minmax","Claude Code · eight hard validated tasks",[2.26,34.79],"\u0001"],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.01435,"usd","$0.014",24,"\u0001","\u0001","Claude Code · eight hard validated tasks","\u0001",true],["coding-agents-head-to-head","coding-agents-pass-rate","Passed every hidden check","Claude Sonnet 5.5 · Claude Code","coding-agents-pass-rate","Coding sessions that passed every hidden check",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","Claude Code · six small repository tasks with hidden tests","\u0001","\u0001"],["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","Claude Sonnet 5.5 · Claude Code","coding-agents-wall-time","Time per coding session",23.1,"seconds","23.1 s",12,"\u0001","minmax","Claude Code · six small repository tasks with hidden tests",[18.7,44.5],"\u0001"],["coding-agents-head-to-head","coding-agents-cost-per-pass","List-price cost per pass","Claude Sonnet 5.5 · Claude Code","coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)",0.085,"usd","$0.085",12,"\u0001","\u0001","Claude Code · six small repository tasks with hidden tests","\u0001",true],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Sonnet 5.5 (low) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · effort low · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Sonnet 5.5 (medium) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Sonnet 5.5 (high) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Sonnet 5.5 · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Sonnet 5.5 (low) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",5.82,"seconds","5.82 s",16,"\u0001","minmax","Claude Code · effort low · eight hard validated tasks, effort ladder",[2.78,19.96],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Sonnet 5.5 (medium) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",7.63,"seconds","7.63 s",16,"\u0001","minmax","Claude Code · effort medium · eight hard validated tasks, effort ladder",[2.71,24.01],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Sonnet 5.5 (high) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",8.81,"seconds","8.81 s",16,"\u0001","minmax","Claude Code · effort high · eight hard validated tasks, effort ladder",[2.93,35.81],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Sonnet 5.5 · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",7.97,"seconds","7.97 s",16,"\u0001","minmax","Claude Code · eight hard validated tasks, effort ladder",[2.26,21.61],"\u0001"],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 (low) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01219,"usd","$0.012",16,"\u0001","\u0001","Claude Code · effort low · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 (medium) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01352,"usd","$0.014",16,"\u0001","\u0001","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 (high) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01671,"usd","$0.017",16,"\u0001","\u0001","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01398,"usd","$0.014",16,"\u0001","\u0001","Claude Code · eight hard validated tasks, effort ladder","\u0001",true],["caching-consistency","consistency-pass-rate","Exact number","Claude Sonnet 5.5 · Claude Code","consistency-pass-rate/Exact number","Same prompt, 10 times: strict pass rate (Exact number)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","JSON object","Claude Sonnet 5.5 · Claude Code","consistency-pass-rate/JSON object","Same prompt, 10 times: strict pass rate (JSON object)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Code fix","Claude Sonnet 5.5 · Claude Code","consistency-pass-rate/Code fix","Same prompt, 10 times: strict pass rate (Code fix)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",6.89,"seconds","6.89 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[5.81,7.81],"\u0001"],["caching-consistency","consistency-latency-spread","JSON object","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/JSON object","Same prompt, 10 times: time per call (JSON object)",2.89,"seconds","2.89 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[2.68,5.3],"\u0001"],["caching-consistency","consistency-latency-spread","Code fix","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",2.67,"seconds","2.67 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[2.32,4.34],"\u0001"],["routing-jev-vs-llm","routing-exact-decisions","Exact rate","Claude Sonnet 5.5","routing-exact-decisions","Typed routing decisions answered exactly right",0.939,"rate","94% (77/82)",82,[0.8651,0.9737],"ci95","typed routing decisions · via Claude Code","\u0001","\u0001"],["routing-jev-vs-llm","routing-cost-per-1000","Cost","Claude Sonnet 5.5","routing-cost-per-1000","Cost per 1,000 routing decisions",4.996,"usd","$5.00",82,"\u0001","\u0001","typed routing decisions · via Claude Code","\u0001",true],["routing-jev-vs-llm","routing-decision-latency","Wall time (CLI)","Claude Sonnet 5.5","routing-decision-latency/Wall time (CLI)","Time per routing decision (Wall time (CLI))",2598,"ms","2,598 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code",[2598,4298],"\u0001"],["routing-jev-vs-llm","routing-decision-latency","Model time (API)","Claude Sonnet 5.5","routing-decision-latency/Model time (API)","Time per routing decision (Model time (API))",1599,"ms","1,599 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code",[1599,2574],"\u0001"]]}],["claude-opus-5-5","Claude Opus 5.5","Anthropic","model","Anthropic’s large Claude model, measured through Claude Code at low, medium, high and default effort.",["Claude Opus 5.5","Opus 5.5"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation"],"$r":[["model-head-to-head","h2h-pass-rate","Pass rate","Claude Opus 5.5 (high) · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Code · effort high · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","Claude Opus 5.5 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Code · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","Claude Opus 5.5 (low) · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Code · effort low · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Opus 5.5 (high) · Claude Code","h2h-total-latency","Total time per call",2.71,"seconds","2.71 s",15,"\u0001","minmax","Claude Code · effort high · five short validated tasks",[2.45,11.78],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Opus 5.5 · Claude Code","h2h-total-latency","Total time per call",2.75,"seconds","2.75 s",15,"\u0001","minmax","Claude Code · five short validated tasks",[2.47,8.91],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Opus 5.5 (low) · Claude Code","h2h-total-latency","Total time per call",2.83,"seconds","2.83 s",15,"\u0001","minmax","Claude Code · effort low · five short validated tasks",[2.35,6.62],"\u0001"],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Opus 5.5 (low) · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.00829,"usd","$0.0083",15,"\u0001","\u0001","Claude Code · effort low · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Opus 5.5 · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.01009,"usd","$0.010",15,"\u0001","\u0001","Claude Code · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Opus 5.5 (high) · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.01049,"usd","$0.010",15,"\u0001","\u0001","Claude Code · effort high · five short validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Opus 5.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Opus 5.5 (high) · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · effort high · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Opus 5.5 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Opus 5.5 (high) · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · effort high · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Opus 5.5 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",9.18,"seconds","9.18 s",24,"\u0001","minmax","Claude Code · eight hard validated tasks",[4.24,27.21],"\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Opus 5.5 (high) · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",11.03,"seconds","11.0 s",24,"\u0001","minmax","Claude Code · effort high · eight hard validated tasks",[3.63,63],"\u0001"],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Opus 5.5 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.02824,"usd","$0.028",24,"\u0001","\u0001","Claude Code · eight hard validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Opus 5.5 (high) · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.03337,"usd","$0.033",24,"\u0001","\u0001","Claude Code · effort high · eight hard validated tasks","\u0001",true],["coding-agents-head-to-head","coding-agents-pass-rate","Passed every hidden check","Claude Opus 5.5 · Claude Code","coding-agents-pass-rate","Coding sessions that passed every hidden check",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","Claude Code · six small repository tasks with hidden tests","\u0001","\u0001"],["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","Claude Opus 5.5 · Claude Code","coding-agents-wall-time","Time per coding session",56.9,"seconds","56.9 s",12,"\u0001","minmax","Claude Code · six small repository tasks with hidden tests",[29.8,185.8],"\u0001"],["coding-agents-head-to-head","coding-agents-cost-per-pass","List-price cost per pass","Claude Opus 5.5 · Claude Code","coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)",0.2229,"usd","$0.22",12,"\u0001","\u0001","Claude Code · six small repository tasks with hidden tests","\u0001",true],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Opus 5.5 (low) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · effort low · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Opus 5.5 (medium) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Opus 5.5 (high) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Opus 5.5 · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Opus 5.5 (low) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",7.5,"seconds","7.50 s",16,"\u0001","minmax","Claude Code · effort low · eight hard validated tasks, effort ladder",[3.34,15.82],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Opus 5.5 (medium) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",9.72,"seconds","9.72 s",16,"\u0001","minmax","Claude Code · effort medium · eight hard validated tasks, effort ladder",[4.78,31.36],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Opus 5.5 (high) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",10.11,"seconds","10.1 s",16,"\u0001","minmax","Claude Code · effort high · eight hard validated tasks, effort ladder",[3.63,63],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Opus 5.5 · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",9.18,"seconds","9.18 s",16,"\u0001","minmax","Claude Code · eight hard validated tasks, effort ladder",[4.24,27.21],"\u0001"],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Opus 5.5 (low) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02115,"usd","$0.021",16,"\u0001","\u0001","Claude Code · effort low · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Opus 5.5 (medium) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02947,"usd","$0.029",16,"\u0001","\u0001","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Opus 5.5 (high) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.03368,"usd","$0.034",16,"\u0001","\u0001","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Opus 5.5 · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02893,"usd","$0.029",16,"\u0001","\u0001","Claude Code · eight hard validated tasks, effort ladder","\u0001",true]]}],["claude-haiku-4-5","Claude Haiku 4.5","Anthropic","model","Anthropic’s small, low-price Claude model, measured through Claude Code, as an LLM router and in the public SWE-bench panel.",["Claude Haiku 4.5","Haiku 4.5","Claude 4.5 Haiku"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation"],"$r":[["model-head-to-head","h2h-pass-rate","Pass rate","Claude Haiku 4.5 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Code · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Haiku 4.5 · Claude Code","h2h-total-latency","Total time per call",4.43,"seconds","4.43 s",15,"\u0001","minmax","Claude Code · five short validated tasks",[3.16,23.57],"\u0001"],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Haiku 4.5 · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.00836,"usd","$0.0084",15,"\u0001","\u0001","Claude Code · five short validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Haiku 4.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",0.4583,"rate","46% (11/24)",24,[0.2789,0.6493],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Haiku 4.5 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,"rate","67% (16/24)",24,[0.4671,0.8203],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Haiku 4.5 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",39.01,"seconds","39.0 s",24,"\u0001","minmax","Claude Code · eight hard validated tasks",[15.27,75.13],"\u0001"],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Haiku 4.5 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.0672,"usd","$0.067",24,"\u0001","\u0001","Claude Code · eight hard validated tasks","\u0001",true],["caching-consistency","consistency-pass-rate","Exact number","Claude Haiku 4.5 · Claude Code","consistency-pass-rate/Exact number","Same prompt, 10 times: strict pass rate (Exact number)",0,"rate","0% (0/10)",10,[0,0.2775],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","JSON object","Claude Haiku 4.5 · Claude Code","consistency-pass-rate/JSON object","Same prompt, 10 times: strict pass rate (JSON object)",0.1,"rate","10% (1/10)",10,[0.0179,0.4042],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Code fix","Claude Haiku 4.5 · Claude Code","consistency-pass-rate/Code fix","Same prompt, 10 times: strict pass rate (Code fix)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Code · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",5.06,"seconds","5.06 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[4.42,6.2],"\u0001"],["caching-consistency","consistency-latency-spread","JSON object","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/JSON object","Same prompt, 10 times: time per call (JSON object)",7.03,"seconds","7.03 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[5.28,12.27],"\u0001"],["caching-consistency","consistency-latency-spread","Code fix","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",5.95,"seconds","5.95 s",10,"\u0001","minmax","Claude Code · same prompt repeated 10 times",[4.89,7.33],"\u0001"],["routing-jev-vs-llm","routing-exact-decisions","Exact rate","Claude Haiku 4.5","routing-exact-decisions","Typed routing decisions answered exactly right",0.8902,"rate","89% (73/82)",82,[0.8044,0.9412],"ci95","typed routing decisions · via Claude Code","\u0001","\u0001"],["routing-jev-vs-llm","routing-cost-per-1000","Cost","Claude Haiku 4.5","routing-cost-per-1000","Cost per 1,000 routing decisions",8.924,"usd","$8.92",82,"\u0001","\u0001","typed routing decisions · via Claude Code","\u0001",true],["routing-jev-vs-llm","routing-decision-latency","Wall time (CLI)","Claude Haiku 4.5","routing-decision-latency/Wall time (CLI)","Time per routing decision (Wall time (CLI))",12674,"ms","12,674 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code",[12674,34413],"\u0001"],["routing-jev-vs-llm","routing-decision-latency","Model time (API)","Claude Haiku 4.5","routing-decision-latency/Model time (API)","Time per routing decision (Model time (API))",10734,"ms","10,734 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code",[10734,32072],"\u0001"]]}],["claude-fable-5-1","Claude Fable 5.1","Anthropic","model","Anthropic’s highest-priced Claude model in these studies, measured through Claude Code.",["Claude Fable 5.1","Fable 5.1"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation"],"$r":[["model-head-to-head","h2h-pass-rate","Pass rate","Claude Fable 5.1 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Code · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Fable 5.1 · Claude Code","h2h-total-latency","Total time per call",1.94,"seconds","1.94 s",15,"\u0001","minmax","Claude Code · five short validated tasks",[1.41,9.83],"\u0001"],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Fable 5.1 · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.02054,"usd","$0.021",15,"\u0001","\u0001","Claude Code · five short validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Fable 5.1 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Fable 5.1 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Fable 5.1 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",16.13,"seconds","16.1 s",24,"\u0001","minmax","Claude Code · eight hard validated tasks",[4.46,90],"\u0001"],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Fable 5.1 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.09331,"usd","$0.093",24,"\u0001","\u0001","Claude Code · eight hard validated tasks","\u0001",true]]}],["gpt-6-1-sol-codex-cli","GPT-6.1 Sol (Codex CLI)","OpenAI","model","OpenAI’s GPT-6.1 Sol model run through the Codex CLI at low, medium and high effort.",["GPT-6.1 Sol · Codex CLI"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation"],"$r":[["model-head-to-head","h2h-pass-rate","Pass rate","GPT-6.1 Sol (high) · Codex CLI","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Codex CLI · effort high · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","GPT-6.1 Sol (medium) · Codex CLI","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","GPT-6.1 Sol (low) · Codex CLI","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Codex CLI · effort low · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","GPT-6.1 Sol (high) · Codex CLI","h2h-total-latency","Total time per call",5.6,"seconds","5.60 s",15,"\u0001","minmax","Codex CLI · effort high · five short validated tasks",[4.05,19.52],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","GPT-6.1 Sol (medium) · Codex CLI","h2h-total-latency","Total time per call",5.65,"seconds","5.65 s",15,"\u0001","minmax","Codex CLI · effort medium · five short validated tasks",[4.1,25.46],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","GPT-6.1 Sol (low) · Codex CLI","h2h-total-latency","Total time per call",6.26,"seconds","6.26 s",10,"\u0001","minmax","Codex CLI · effort low · five short validated tasks",[4.65,10.47],"\u0001"],["model-head-to-head","h2h-cost-per-pass","Cost per pass","GPT-6.1 Sol (low) · Codex CLI","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.00998,"usd","$0.010",10,"\u0001","\u0001","Codex CLI · effort low · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","GPT-6.1 Sol (high) · Codex CLI","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.01322,"usd","$0.013",15,"\u0001","\u0001","Codex CLI · effort high · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","GPT-6.1 Sol (medium) · Codex CLI","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.01564,"usd","$0.016",15,"\u0001","\u0001","Codex CLI · effort medium · five short validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","GPT-6.1 Sol (high) · Codex CLI","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","GPT-6.1 Sol (high) · Codex CLI","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",13.11,"seconds","13.1 s",16,"\u0001","minmax","Codex CLI · effort medium · eight hard validated tasks",[8.54,61.6],"\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","GPT-6.1 Sol (high) · Codex CLI","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",18.12,"seconds","18.1 s",16,"\u0001","minmax","Codex CLI · effort high · eight hard validated tasks",[11.67,92.21],"\u0001"],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (high) · Codex CLI","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.01514,"usd","$0.015",16,"\u0001","\u0001","Codex CLI · effort high · eight hard validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.02564,"usd","$0.026",16,"\u0001","\u0001","Codex CLI · effort medium · eight hard validated tasks","\u0001",true],["coding-agents-head-to-head","coding-agents-pass-rate","Passed every hidden check","GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI","coding-agents-pass-rate","Coding sessions that passed every hidden check",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001"],["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI","coding-agents-wall-time","Time per coding session",113.4,"seconds","113.4 s",12,"\u0001","minmax","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests",[78.5,221.9],"\u0001"],["coding-agents-head-to-head","coding-agents-cost-per-pass","List-price cost per pass","GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI","coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)",0.0978,"usd","$0.098",12,"\u0001","\u0001","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001",true],["effort-ladder","effort-ladder-pass-rate","Strict pass","GPT-6.1 Sol (low) · Codex CLI","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort low · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","GPT-6.1 Sol (high) · Codex CLI","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","GPT-6.1 Sol (low) · Codex CLI","effort-ladder-total-latency","Total time per call by effort on hard tasks",13.62,"seconds","13.6 s",16,"\u0001","minmax","Codex CLI · effort low · eight hard validated tasks, effort ladder",[7.94,44.29],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-total-latency","Total time per call by effort on hard tasks",13.11,"seconds","13.1 s",16,"\u0001","minmax","Codex CLI · effort medium · eight hard validated tasks, effort ladder",[8.54,61.6],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","GPT-6.1 Sol (high) · Codex CLI","effort-ladder-total-latency","Total time per call by effort on hard tasks",18.12,"seconds","18.1 s",16,"\u0001","minmax","Codex CLI · effort high · eight hard validated tasks, effort ladder",[11.67,92.21],"\u0001"],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (low) · Codex CLI","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01284,"usd","$0.013",16,"\u0001","\u0001","Codex CLI · effort low · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02564,"usd","$0.026",16,"\u0001","\u0001","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (high) · Codex CLI","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01514,"usd","$0.015",16,"\u0001","\u0001","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001",true],["caching-consistency","consistency-pass-rate","Exact number","GPT-6.1 Sol (medium) · Codex CLI","consistency-pass-rate/Exact number","Same prompt, 10 times: strict pass rate (Exact number)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","JSON object","GPT-6.1 Sol (medium) · Codex CLI","consistency-pass-rate/JSON object","Same prompt, 10 times: strict pass rate (JSON object)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Code fix","GPT-6.1 Sol (medium) · Codex CLI","consistency-pass-rate/Code fix","Same prompt, 10 times: strict pass rate (Code fix)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Codex CLI · effort medium · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",13.38,"seconds","13.4 s",10,"\u0001","minmax","Codex CLI · effort medium · same prompt repeated 10 times",[12.29,17.97],"\u0001"],["caching-consistency","consistency-latency-spread","JSON object","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/JSON object","Same prompt, 10 times: time per call (JSON object)",6.42,"seconds","6.42 s",10,"\u0001","minmax","Codex CLI · effort medium · same prompt repeated 10 times",[5.25,8.26],"\u0001"],["caching-consistency","consistency-latency-spread","Code fix","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",11.29,"seconds","11.3 s",10,"\u0001","minmax","Codex CLI · effort medium · same prompt repeated 10 times",[9.08,14.85],"\u0001"]]}],["claude-code-cli","Claude Code","Anthropic","cli","Anthropic’s coding CLI. Each measurement pairs it with one Claude model; the context names the model.",["Claude Code","Claude Code CLI"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation"],"$r":[["model-head-to-head","h2h-pass-rate","Pass rate","Claude Fable 5.1 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Fable 5.1 · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","Claude Sonnet 5.5 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",0.8,"rate","80% (12/15)",15,[0.5481,0.9295],"ci95","Claude Sonnet 5.5 · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","Claude Opus 5.5 (high) · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Opus 5.5 · effort high · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","Claude Opus 5.5 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Opus 5.5 · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","Claude Opus 5.5 (low) · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Opus 5.5 · effort low · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","Claude Haiku 4.5 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Haiku 4.5 · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Fable 5.1 · Claude Code","h2h-total-latency","Total time per call",1.94,"seconds","1.94 s",15,"\u0001","minmax","Claude Fable 5.1 · five short validated tasks",[1.41,9.83],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Sonnet 5.5 · Claude Code","h2h-total-latency","Total time per call",2.31,"seconds","2.31 s",15,"\u0001","minmax","Claude Sonnet 5.5 · five short validated tasks",[2.17,7.73],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Opus 5.5 (high) · Claude Code","h2h-total-latency","Total time per call",2.71,"seconds","2.71 s",15,"\u0001","minmax","Claude Opus 5.5 · effort high · five short validated tasks",[2.45,11.78],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Opus 5.5 · Claude Code","h2h-total-latency","Total time per call",2.75,"seconds","2.75 s",15,"\u0001","minmax","Claude Opus 5.5 · five short validated tasks",[2.47,8.91],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Opus 5.5 (low) · Claude Code","h2h-total-latency","Total time per call",2.83,"seconds","2.83 s",15,"\u0001","minmax","Claude Opus 5.5 · effort low · five short validated tasks",[2.35,6.62],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","Claude Haiku 4.5 · Claude Code","h2h-total-latency","Total time per call",4.43,"seconds","4.43 s",15,"\u0001","minmax","Claude Haiku 4.5 · five short validated tasks",[3.16,23.57],"\u0001"],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Sonnet 5.5 · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.00624,"usd","$0.0062",15,"\u0001","\u0001","Claude Sonnet 5.5 · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Opus 5.5 (low) · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.00829,"usd","$0.0083",15,"\u0001","\u0001","Claude Opus 5.5 · effort low · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Haiku 4.5 · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.00836,"usd","$0.0084",15,"\u0001","\u0001","Claude Haiku 4.5 · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Opus 5.5 · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.01009,"usd","$0.010",15,"\u0001","\u0001","Claude Opus 5.5 · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Opus 5.5 (high) · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.01049,"usd","$0.010",15,"\u0001","\u0001","Claude Opus 5.5 · effort high · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","Claude Fable 5.1 · Claude Code","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.02054,"usd","$0.021",15,"\u0001","\u0001","Claude Fable 5.1 · five short validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Sonnet 5.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Sonnet 5.5 · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Opus 5.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Opus 5.5 · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Opus 5.5 (high) · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Opus 5.5 · effort high · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Fable 5.1 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Fable 5.1 · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Haiku 4.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",0.4583,"rate","46% (11/24)",24,[0.2789,0.6493],"ci95","Claude Haiku 4.5 · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Sonnet 5.5 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Sonnet 5.5 · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Opus 5.5 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Opus 5.5 · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Opus 5.5 (high) · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Opus 5.5 · effort high · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Fable 5.1 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Fable 5.1 · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","Claude Haiku 4.5 · Claude Code","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,"rate","67% (16/24)",24,[0.4671,0.8203],"ci95","Claude Haiku 4.5 · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Sonnet 5.5 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",7.75,"seconds","7.75 s",24,"\u0001","minmax","Claude Sonnet 5.5 · eight hard validated tasks",[2.26,34.79],"\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Opus 5.5 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",9.18,"seconds","9.18 s",24,"\u0001","minmax","Claude Opus 5.5 · eight hard validated tasks",[4.24,27.21],"\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Opus 5.5 (high) · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",11.03,"seconds","11.0 s",24,"\u0001","minmax","Claude Opus 5.5 · effort high · eight hard validated tasks",[3.63,63],"\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Fable 5.1 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",16.13,"seconds","16.1 s",24,"\u0001","minmax","Claude Fable 5.1 · eight hard validated tasks",[4.46,90],"\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Haiku 4.5 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",39.01,"seconds","39.0 s",24,"\u0001","minmax","Claude Haiku 4.5 · eight hard validated tasks",[15.27,75.13],"\u0001"],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.01435,"usd","$0.014",24,"\u0001","\u0001","Claude Sonnet 5.5 · eight hard validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Opus 5.5 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.02824,"usd","$0.028",24,"\u0001","\u0001","Claude Opus 5.5 · eight hard validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Opus 5.5 (high) · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.03337,"usd","$0.033",24,"\u0001","\u0001","Claude Opus 5.5 · effort high · eight hard validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Haiku 4.5 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.0672,"usd","$0.067",24,"\u0001","\u0001","Claude Haiku 4.5 · eight hard validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Fable 5.1 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.09331,"usd","$0.093",24,"\u0001","\u0001","Claude Fable 5.1 · eight hard validated tasks","\u0001",true],["coding-agents-head-to-head","coding-agents-pass-rate","Passed every hidden check","Claude Sonnet 5.5 · Claude Code","coding-agents-pass-rate","Coding sessions that passed every hidden check",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","Claude Sonnet 5.5 · six small repository tasks with hidden tests","\u0001","\u0001"],["coding-agents-head-to-head","coding-agents-pass-rate","Passed every hidden check","Claude Opus 5.5 · Claude Code","coding-agents-pass-rate","Coding sessions that passed every hidden check",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","Claude Opus 5.5 · six small repository tasks with hidden tests","\u0001","\u0001"],["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","Claude Sonnet 5.5 · Claude Code","coding-agents-wall-time","Time per coding session",23.1,"seconds","23.1 s",12,"\u0001","minmax","Claude Sonnet 5.5 · six small repository tasks with hidden tests",[18.7,44.5],"\u0001"],["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","Claude Opus 5.5 · Claude Code","coding-agents-wall-time","Time per coding session",56.9,"seconds","56.9 s",12,"\u0001","minmax","Claude Opus 5.5 · six small repository tasks with hidden tests",[29.8,185.8],"\u0001"],["coding-agents-head-to-head","coding-agents-cost-per-pass","List-price cost per pass","Claude Sonnet 5.5 · Claude Code","coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)",0.085,"usd","$0.085",12,"\u0001","\u0001","Claude Sonnet 5.5 · six small repository tasks with hidden tests","\u0001",true],["coding-agents-head-to-head","coding-agents-cost-per-pass","List-price cost per pass","Claude Opus 5.5 · Claude Code","coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)",0.2229,"usd","$0.22",12,"\u0001","\u0001","Claude Opus 5.5 · six small repository tasks with hidden tests","\u0001",true],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Sonnet 5.5 (low) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Sonnet 5.5 · effort low · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Sonnet 5.5 (medium) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Sonnet 5.5 (high) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Sonnet 5.5 · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Sonnet 5.5 · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Sonnet 5.5 · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Opus 5.5 (low) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Opus 5.5 · effort low · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Opus 5.5 (medium) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Opus 5.5 · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Opus 5.5 (high) · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Opus 5.5 · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Opus 5.5 · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Opus 5.5 · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Sonnet 5.5 (low) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",5.82,"seconds","5.82 s",16,"\u0001","minmax","Claude Sonnet 5.5 · effort low · eight hard validated tasks, effort ladder",[2.78,19.96],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Sonnet 5.5 (medium) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",7.63,"seconds","7.63 s",16,"\u0001","minmax","Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder",[2.71,24.01],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Sonnet 5.5 (high) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",8.81,"seconds","8.81 s",16,"\u0001","minmax","Claude Sonnet 5.5 · effort high · eight hard validated tasks, effort ladder",[2.93,35.81],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Sonnet 5.5 · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",7.97,"seconds","7.97 s",16,"\u0001","minmax","Claude Sonnet 5.5 · eight hard validated tasks, effort ladder",[2.26,21.61],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Opus 5.5 (low) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",7.5,"seconds","7.50 s",16,"\u0001","minmax","Claude Opus 5.5 · effort low · eight hard validated tasks, effort ladder",[3.34,15.82],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Opus 5.5 (medium) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",9.72,"seconds","9.72 s",16,"\u0001","minmax","Claude Opus 5.5 · effort medium · eight hard validated tasks, effort ladder",[4.78,31.36],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Opus 5.5 (high) · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",10.11,"seconds","10.1 s",16,"\u0001","minmax","Claude Opus 5.5 · effort high · eight hard validated tasks, effort ladder",[3.63,63],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Opus 5.5 · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",9.18,"seconds","9.18 s",16,"\u0001","minmax","Claude Opus 5.5 · eight hard validated tasks, effort ladder",[4.24,27.21],"\u0001"],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 (low) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01219,"usd","$0.012",16,"\u0001","\u0001","Claude Sonnet 5.5 · effort low · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 (medium) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01352,"usd","$0.014",16,"\u0001","\u0001","Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 (high) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01671,"usd","$0.017",16,"\u0001","\u0001","Claude Sonnet 5.5 · effort high · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Sonnet 5.5 · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01398,"usd","$0.014",16,"\u0001","\u0001","Claude Sonnet 5.5 · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Opus 5.5 (low) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02115,"usd","$0.021",16,"\u0001","\u0001","Claude Opus 5.5 · effort low · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Opus 5.5 (medium) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02947,"usd","$0.029",16,"\u0001","\u0001","Claude Opus 5.5 · effort medium · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Opus 5.5 (high) · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.03368,"usd","$0.034",16,"\u0001","\u0001","Claude Opus 5.5 · effort high · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","Claude Opus 5.5 · Claude Code","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02893,"usd","$0.029",16,"\u0001","\u0001","Claude Opus 5.5 · eight hard validated tasks, effort ladder","\u0001",true],["caching-consistency","consistency-pass-rate","Exact number","Claude Haiku 4.5 · Claude Code","consistency-pass-rate/Exact number","Same prompt, 10 times: strict pass rate (Exact number)",0,"rate","0% (0/10)",10,[0,0.2775],"ci95","Claude Haiku 4.5 · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Exact number","Claude Sonnet 5.5 · Claude Code","consistency-pass-rate/Exact number","Same prompt, 10 times: strict pass rate (Exact number)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Sonnet 5.5 · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","JSON object","Claude Haiku 4.5 · Claude Code","consistency-pass-rate/JSON object","Same prompt, 10 times: strict pass rate (JSON object)",0.1,"rate","10% (1/10)",10,[0.0179,0.4042],"ci95","Claude Haiku 4.5 · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","JSON object","Claude Sonnet 5.5 · Claude Code","consistency-pass-rate/JSON object","Same prompt, 10 times: strict pass rate (JSON object)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Sonnet 5.5 · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Code fix","Claude Haiku 4.5 · Claude Code","consistency-pass-rate/Code fix","Same prompt, 10 times: strict pass rate (Code fix)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Haiku 4.5 · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Code fix","Claude Sonnet 5.5 · Claude Code","consistency-pass-rate/Code fix","Same prompt, 10 times: strict pass rate (Code fix)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","Claude Sonnet 5.5 · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",5.06,"seconds","5.06 s",10,"\u0001","minmax","Claude Haiku 4.5 · same prompt repeated 10 times",[4.42,6.2],"\u0001"],["caching-consistency","consistency-latency-spread","Exact number","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",6.89,"seconds","6.89 s",10,"\u0001","minmax","Claude Sonnet 5.5 · same prompt repeated 10 times",[5.81,7.81],"\u0001"],["caching-consistency","consistency-latency-spread","JSON object","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/JSON object","Same prompt, 10 times: time per call (JSON object)",7.03,"seconds","7.03 s",10,"\u0001","minmax","Claude Haiku 4.5 · same prompt repeated 10 times",[5.28,12.27],"\u0001"],["caching-consistency","consistency-latency-spread","JSON object","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/JSON object","Same prompt, 10 times: time per call (JSON object)",2.89,"seconds","2.89 s",10,"\u0001","minmax","Claude Sonnet 5.5 · same prompt repeated 10 times",[2.68,5.3],"\u0001"],["caching-consistency","consistency-latency-spread","Code fix","Claude Haiku 4.5 · Claude Code","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",5.95,"seconds","5.95 s",10,"\u0001","minmax","Claude Haiku 4.5 · same prompt repeated 10 times",[4.89,7.33],"\u0001"],["caching-consistency","consistency-latency-spread","Code fix","Claude Sonnet 5.5 · Claude Code","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",2.67,"seconds","2.67 s",10,"\u0001","minmax","Claude Sonnet 5.5 · same prompt repeated 10 times",[2.32,4.34],"\u0001"]]}],["codex-cli","Codex CLI","OpenAI","cli","OpenAI’s coding CLI. Each measurement pairs it with one GPT model and effort; the context names them.",["Codex CLI"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation"],"$r":[["model-head-to-head","h2h-pass-rate","Pass rate","GPT-6.1 Sol (high) · Codex CLI","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","GPT-6.1 Sol · effort high · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","GPT-6.1 Sol (medium) · Codex CLI","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-pass-rate","Pass rate","GPT-6.1 Sol (low) · Codex CLI","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","GPT-6.1 Sol · effort low · five short validated tasks","\u0001","\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","GPT-6.1 Sol (high) · Codex CLI","h2h-total-latency","Total time per call",5.6,"seconds","5.60 s",15,"\u0001","minmax","GPT-6.1 Sol · effort high · five short validated tasks",[4.05,19.52],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","GPT-6.1 Sol (medium) · Codex CLI","h2h-total-latency","Total time per call",5.65,"seconds","5.65 s",15,"\u0001","minmax","GPT-6.1 Sol · effort medium · five short validated tasks",[4.1,25.46],"\u0001"],["model-head-to-head","h2h-total-latency","Total time per call","GPT-6.1 Sol (low) · Codex CLI","h2h-total-latency","Total time per call",6.26,"seconds","6.26 s",10,"\u0001","minmax","GPT-6.1 Sol · effort low · five short validated tasks",[4.65,10.47],"\u0001"],["model-head-to-head","h2h-cost-per-pass","Cost per pass","GPT-6.1 Sol (low) · Codex CLI","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.00998,"usd","$0.010",10,"\u0001","\u0001","GPT-6.1 Sol · effort low · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","GPT-6.1 Sol (high) · Codex CLI","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.01322,"usd","$0.013",15,"\u0001","\u0001","GPT-6.1 Sol · effort high · five short validated tasks","\u0001",true],["model-head-to-head","h2h-cost-per-pass","Cost per pass","GPT-6.1 Sol (medium) · Codex CLI","h2h-cost-per-pass","List-price cost per passing answer (calculation)",0.01564,"usd","$0.016",15,"\u0001","\u0001","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","GPT-6.1 Sol (high) · Codex CLI","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","GPT-6.1 Sol · effort high · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-pass-rate","Lenient (format misses counted)","GPT-6.1 Sol (high) · Codex CLI","hard-h2h-pass-rate/Lenient (format misses counted)","Pass rate on eight hard tasks (Lenient (format misses counted))",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","GPT-6.1 Sol · effort high · eight hard validated tasks","\u0001","\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",13.11,"seconds","13.1 s",16,"\u0001","minmax","GPT-6.1 Sol · effort medium · eight hard validated tasks",[8.54,61.6],"\u0001"],["hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","GPT-6.1 Sol (high) · Codex CLI","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",18.12,"seconds","18.1 s",16,"\u0001","minmax","GPT-6.1 Sol · effort high · eight hard validated tasks",[11.67,92.21],"\u0001"],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (high) · Codex CLI","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.01514,"usd","$0.015",16,"\u0001","\u0001","GPT-6.1 Sol · effort high · eight hard validated tasks","\u0001",true],["hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.02564,"usd","$0.026",16,"\u0001","\u0001","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001",true],["coding-agents-head-to-head","coding-agents-pass-rate","Passed every hidden check","GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI","coding-agents-pass-rate","Coding sessions that passed every hidden check",1,"rate","100% (12/12)",12,[0.7575,1],"ci95","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001"],["coding-agents-head-to-head","coding-agents-wall-time","Wall time per session","GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI","coding-agents-wall-time","Time per coding session",113.4,"seconds","113.4 s",12,"\u0001","minmax","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests",[78.5,221.9],"\u0001"],["coding-agents-head-to-head","coding-agents-cost-per-pass","List-price cost per pass","GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI","coding-agents-cost-per-pass","List-price cost per passing coding session (calculation)",0.0978,"usd","$0.098",12,"\u0001","\u0001","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001",true],["effort-ladder","effort-ladder-pass-rate","Strict pass","GPT-6.1 Sol (low) · Codex CLI","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","GPT-6.1 Sol · effort low · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-pass-rate","Strict pass","GPT-6.1 Sol (high) · Codex CLI","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","GPT-6.1 Sol · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","GPT-6.1 Sol (low) · Codex CLI","effort-ladder-total-latency","Total time per call by effort on hard tasks",13.62,"seconds","13.6 s",16,"\u0001","minmax","GPT-6.1 Sol · effort low · eight hard validated tasks, effort ladder",[7.94,44.29],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-total-latency","Total time per call by effort on hard tasks",13.11,"seconds","13.1 s",16,"\u0001","minmax","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder",[8.54,61.6],"\u0001"],["effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","GPT-6.1 Sol (high) · Codex CLI","effort-ladder-total-latency","Total time per call by effort on hard tasks",18.12,"seconds","18.1 s",16,"\u0001","minmax","GPT-6.1 Sol · effort high · eight hard validated tasks, effort ladder",[11.67,92.21],"\u0001"],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (low) · Codex CLI","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01284,"usd","$0.013",16,"\u0001","\u0001","GPT-6.1 Sol · effort low · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02564,"usd","$0.026",16,"\u0001","\u0001","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001",true],["effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (high) · Codex CLI","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.01514,"usd","$0.015",16,"\u0001","\u0001","GPT-6.1 Sol · effort high · eight hard validated tasks, effort ladder","\u0001",true],["caching-consistency","consistency-pass-rate","Exact number","GPT-6.1 Sol (medium) · Codex CLI","consistency-pass-rate/Exact number","Same prompt, 10 times: strict pass rate (Exact number)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","JSON object","GPT-6.1 Sol (medium) · Codex CLI","consistency-pass-rate/JSON object","Same prompt, 10 times: strict pass rate (JSON object)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-pass-rate","Code fix","GPT-6.1 Sol (medium) · Codex CLI","consistency-pass-rate/Code fix","Same prompt, 10 times: strict pass rate (Code fix)",1,"rate","100% (10/10)",10,[0.7225,1],"ci95","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","\u0001","\u0001"],["caching-consistency","consistency-latency-spread","Exact number","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/Exact number","Same prompt, 10 times: time per call (Exact number)",13.38,"seconds","13.4 s",10,"\u0001","minmax","GPT-6.1 Sol · effort medium · same prompt repeated 10 times",[12.29,17.97],"\u0001"],["caching-consistency","consistency-latency-spread","JSON object","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/JSON object","Same prompt, 10 times: time per call (JSON object)",6.42,"seconds","6.42 s",10,"\u0001","minmax","GPT-6.1 Sol · effort medium · same prompt repeated 10 times",[5.25,8.26],"\u0001"],["caching-consistency","consistency-latency-spread","Code fix","GPT-6.1 Sol (medium) · Codex CLI","consistency-latency-spread/Code fix","Same prompt, 10 times: time per call (Code fix)",11.29,"seconds","11.3 s",10,"\u0001","minmax","GPT-6.1 Sol · effort medium · same prompt repeated 10 times",[9.08,14.85],"\u0001"]]}],["jev-1-13","Jev 1.13","TypeSafe","router","A small routing model from TypeSafe that answers typed routing decisions over an API. Priced on input tokens only. Measured here as a router (accuracy, live time per call and cost per 1,000 decisions) and in the System One arena.",["Jev 1.13"],{"$k":["studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","calculation","range"],"$r":[["routing-jev-vs-llm","routing-exact-decisions","Exact rate","Jev 1.13 (TypeSafe)","routing-exact-decisions","Typed routing decisions answered exactly right",0.8984,"rate","90%",82,[0.8191,0.9497],"ci95","typed routing decisions · TypeSafe API","\u0001","\u0001"],["routing-jev-vs-llm","routing-cost-per-1000","Cost","Jev 1.13 (TypeSafe)","routing-cost-per-1000","Cost per 1,000 routing decisions",0.0337,"usd","$0.034",246,"\u0001","\u0001","typed routing decisions · TypeSafe API",true,"\u0001"],["routing-jev-vs-llm","routing-decision-latency","Wall time (direct API call)","Jev 1.13 (TypeSafe)","routing-decision-latency/Wall time (direct API call)","Time per routing decision (Wall time (direct API call))",136.5,"ms","137 ms",246,"\u0001","p50-p95","typed routing decisions · TypeSafe API","\u0001",[136.5,195.7]],["routing-overhead","router-overhead-cost-reported","Cost per 1,000 decisions","Jev 1.13 (TypeSafe)","router-overhead-cost-reported","Cost per 1,000 routing decisions: no model call vs provider-reported",0.0337,"usd","$0.034",82,"\u0001","\u0001","routing overhead per decision · TypeSafe API","\u0001","\u0001"]]}],["deterministic-routing-policy","Deterministic routing policy","Agent","router","Agent’s rule-based routing: an in-process policy picks the model and effort for each call from the task stage and signals. No model call, so no token cost.",["Deterministic routing policy"],[{"studySlug":"routing-overhead","chartId":"router-overhead-cost-reported","series":"Cost per 1,000 decisions","point":"Deterministic routing policy (Agent, in process)","metric":"router-overhead-cost-reported","label":"Cost per 1,000 routing decisions: no model call vs provider-reported","value":0,"unit":"usd","display":"$0.00","n":20000,"context":"Agent · in process · routing overhead per decision"}]]]},"comparisons":{"$k":["slug","a","b","title","seoTitle","description","verdict","rows"],"$r":[["claude-sonnet-5-5-vs-claude-opus-5-5","claude-sonnet-5-5","claude-opus-5-5","Claude Sonnet 5.5 vs Claude Opus 5.5","Claude Sonnet 5.5 vs Claude Opus 5.5: measured benchmarks","Claude Sonnet 5.5 vs Claude Opus 5.5: 35 measured metrics from 8 studies, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 and Claude Opus 5.5 share 35 measured metrics and 31 list-price calculations from 10 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 16 ties and 50 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,2.75,"seconds","2.31 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.17,7.73],[2.47,8.91],"\u0001"],["List-price cost per passing answer (calculation)",0.00624,0.01009,"usd","$0.0062","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Opus 5.5 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Opus 5.5 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,9.18,"seconds","7.75 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[2.26,34.79],[4.24,27.21],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.02824,"usd","$0.014","$0.028","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.028) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; Claude Opus 5.5 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",23.1,56.9,"seconds","23.1 s","56.9 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; Claude Opus 5.5 29.8 s to 185.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","range","minmax",[18.7,44.5],[29.8,185.8],"\u0001"],["List-price cost per passing coding session (calculation)",0.085,0.2229,"usd","$0.085","$0.22","unclear","No interval or range was recorded for either side, so the gap ($0.085 vs $0.22, 2.6x) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Code · six small repository tasks with hidden tests","Claude Code · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 81% to 100%; Claude Opus 5.5 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.97,9.18,"seconds","7.97 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 21.6 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[2.26,21.61],[4.24,27.21],"\u0001"],["List-price cost per strict pass by effort (calculation)",0.01398,0.02893,"usd","$0.014","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.029, 2.1x) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-haiku-4-5-vs-claude-sonnet-5-5","claude-haiku-4-5","claude-sonnet-5-5","Claude Haiku 4.5 vs Claude Sonnet 5.5","Claude Haiku 4.5 vs Claude Sonnet 5.5: measured benchmarks","Claude Haiku 4.5 vs Claude Sonnet 5.5: 113 measured metrics from 12 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Haiku 4.5 and Claude Sonnet 5.5 share 113 measured metrics and 40 list-price calculations from 14 studies. Claude Sonnet 5.5 leads on 21 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); and 18 more. On those rows the 95% intervals, run ranges and p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 58 ties and 74 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 2 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,0.8,"rate","100% (15/15)","80% (12/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.5481,0.9295],"\u0001"],["Total time per call",4.43,2.31,"seconds","4.43 s","2.31 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Sonnet 5.5 2.17 s to 7.73 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[3.16,23.57],[2.17,7.73],"\u0001"],["List-price cost per passing answer (calculation)",0.00836,0.00624,"usd","$0.0084","$0.0062","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.0062) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Sonnet 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,7.75,"seconds","39.0 s","7.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Sonnet 5.5 2.26 s to 34.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[15.27,75.13],[2.26,34.79],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.01435,"usd","$0.067","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.014, 4.7x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",0,1,"rate","0% (0/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","ci95","ci95",[0,0.2775],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",0.1,1,"rate","10% (1/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","ci95","ci95",[0.0179,0.4042],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: time per call (Exact number)",5.06,6.89,"seconds","5.06 s","6.89 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 4.42 s to 6.20 s; Claude Sonnet 5.5 5.81 s to 7.81 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","range","minmax",[4.42,6.2],[5.81,7.81],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",7.03,2.89,"seconds","7.03 s","2.89 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.28 s to 12.3 s; Claude Sonnet 5.5 2.68 s to 5.30 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","range","minmax",[5.28,12.27],[2.68,5.3],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",5.95,2.67,"seconds","5.95 s","2.67 s","b","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; Claude Sonnet 5.5 2.32 s to 4.34 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Claude Code · same prompt repeated 10 times","range","minmax",[4.89,7.33],[2.32,4.34],"\u0001"],["Typed routing decisions answered exactly right",0.8902,0.939,"rate","89% (73/82)","94% (77/82)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 94%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","routing-jev-vs-llm",82,"routing-exact-decisions",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","ci95","ci95",[0.8044,0.9412],[0.8651,0.9737],"\u0001"],["Cost per 1,000 routing decisions",8.924,4.996,"usd","$8.92","$5.00","unclear","No interval or range was recorded for either side, so the gap ($8.92 vs $5.00) is not tested against run-to-run variation.","routing-jev-vs-llm",82,"routing-cost-per-1000",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Time per routing decision (Wall time (CLI))",12674,2598,"ms","12,674 ms","2,598 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,674 ms to 34,413 ms; Claude Sonnet 5.5 2,598 ms to 4,298 ms); not a confidence interval.","routing-jev-vs-llm",82,"routing-decision-latency",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","range","p50-p95",[12674,34413],[2598,4298],"\u0001"],["Time per routing decision (Model time (API))",10734,1599,"ms","10,734 ms","1,599 ms","b","Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,734 ms to 32,072 ms; Claude Sonnet 5.5 1,599 ms to 2,574 ms); not a confidence interval.","routing-jev-vs-llm",82,"routing-decision-latency",82,82,"typed routing decisions · via Claude Code","typed routing decisions · via Claude Code","range","p50-p95",[10734,32072],[1599,2574],"\u0001"]]}],["claude-opus-5-5-vs-claude-fable-5-1","claude-opus-5-5","claude-fable-5-1","Claude Opus 5.5 vs Claude Fable 5.1","Claude Opus 5.5 vs Claude Fable 5.1: measured benchmarks","Claude Opus 5.5 vs Claude Fable 5.1: 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Opus 5.5 and Claude Fable 5.1 share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 5 ties and 18 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 4 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 80% to 100%; Claude Fable 5.1 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.75,1.94,"seconds","2.75 s","1.94 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.47 s to 8.91 s; Claude Fable 5.1 1.41 s to 9.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.47,8.91],[1.41,9.83],"\u0001"],["List-price cost per passing answer (calculation)",0.01009,0.02054,"usd","$0.010","$0.021","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.021, 2.0x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; Claude Fable 5.1 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; Claude Fable 5.1 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",9.18,16.13,"seconds","9.18 s","16.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 4.24 s to 27.2 s; Claude Fable 5.1 4.46 s to 90.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[4.24,27.21],[4.46,90],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.02824,0.09331,"usd","$0.028","$0.093","unclear","No interval or range was recorded for either side, so the gap ($0.028 vs $0.093, 3.3x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-vs-claude-fable-5-1","claude-sonnet-5-5","claude-fable-5-1","Claude Sonnet 5.5 vs Claude Fable 5.1","Claude Sonnet 5.5 vs Claude Fable 5.1: measured benchmarks","Claude Sonnet 5.5 vs Claude Fable 5.1: 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Sonnet 5.5 and Claude Fable 5.1 share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 4 ties and 19 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 4 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; Claude Fable 5.1 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,1.94,"seconds","2.31 s","1.94 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; Claude Fable 5.1 1.41 s to 9.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.17,7.73],[1.41,9.83],"\u0001"],["List-price cost per passing answer (calculation)",0.00624,0.02054,"usd","$0.0062","$0.021","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.021, 3.3x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Fable 5.1 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Fable 5.1 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,16.13,"seconds","7.75 s","16.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; Claude Fable 5.1 4.46 s to 90.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[2.26,34.79],[4.46,90],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.09331,"usd","$0.014","$0.093","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.093, 6.5x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-haiku-4-5-vs-claude-opus-5-5","claude-haiku-4-5","claude-opus-5-5","Claude Haiku 4.5 vs Claude Opus 5.5","Claude Haiku 4.5 vs Claude Opus 5.5: measured benchmarks","Claude Haiku 4.5 vs Claude Opus 5.5: 25 measured metrics from 4 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Haiku 4.5 and Claude Opus 5.5 share 25 measured metrics and 14 list-price calculations from 5 studies. Claude Opus 5.5 leads on 3 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Pass rate on 4 harder tasks (Lenient (format misses counted)), 50% (6/12) vs 0% (0/12). On those rows the 95% intervals do not overlap. The other rows are 7 ties and 29 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",4.43,2.75,"seconds","4.43 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[3.16,23.57],[2.47,8.91],"\u0001"],["List-price cost per passing answer (calculation)",0.00836,0.01009,"usd","$0.0084","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Opus 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Opus 5.5 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,9.18,"seconds","39.0 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[15.27,75.13],[4.24,27.21],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.02824,"usd","$0.067","$0.028","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.028, 2.4x) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-code-cli-vs-codex-cli","claude-code-cli","codex-cli","Claude Code vs Codex CLI","Claude Code vs Codex CLI: measured benchmarks","Claude Code vs Codex CLI: 66 measured metrics from 11 studies (Pass rate on five validated tasks; Total time per call; more), with sample sizes and intervals.","Claude Code and Codex CLI share 66 measured metrics and 22 list-price calculations from 12 studies. Claude Code leads on 7 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 4 more. Codex CLI leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 31 ties and 49 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Each run pairs a CLI with a model, so these rows cannot separate the CLI from the model; the contexts name both. Some rows rest on small samples (n = 2 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Code 55% to 93%; Codex CLI 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,5.65,"seconds","2.31 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.17 s to 7.73 s; Codex CLI 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","range","minmax",[2.17,7.73],[4.1,25.46],"\u0001"],["List-price cost per passing answer (calculation)",0.00624,0.01564,"usd","$0.0062","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Sonnet 5.5 · five short validated tasks","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,13.11,"seconds","7.75 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","range","minmax",[2.26,34.79],[8.54,61.6],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Sonnet 5.5 · eight hard validated tasks","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",23.1,113.4,"seconds","23.1 s","113.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[18.7,44.5],[78.5,221.9],"\u0001"],["List-price cost per passing coding session (calculation)",0.085,0.0978,"usd","$0.085","$0.098","unclear","No interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Sonnet 5.5 · six small repository tasks with hidden tests","GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Code 81% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.63,13.11,"seconds","7.63 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.71 s to 24.0 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","range","minmax",[2.71,24.01],[8.54,61.6],"\u0001"],["List-price cost per strict pass by effort (calculation)",0.01352,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: time per call (Exact number)",6.89,13.38,"seconds","6.89 s","13.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","range","minmax",[5.81,7.81],[12.29,17.97],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",2.89,6.42,"seconds","2.89 s","6.42 s","unclear","The run ranges (fastest to slowest) overlap (Claude Code 2.68 s to 5.30 s; Codex CLI 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","range","minmax",[2.68,5.3],[5.25,8.26],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",2.67,11.29,"seconds","2.67 s","11.3 s","a","The run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Sonnet 5.5 · same prompt repeated 10 times","GPT-6.1 Sol · effort medium · same prompt repeated 10 times","range","minmax",[2.32,4.34],[9.08,14.85],"\u0001"]]}],["claude-sonnet-5-5-vs-gpt-6-1-sol-codex-cli","claude-sonnet-5-5","gpt-6-1-sol-codex-cli","Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)","Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): benchmarks","Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): 49 measured metrics from 9 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) share 49 measured metrics and 21 list-price calculations from 10 studies. Claude Sonnet 5.5 leads on 4 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 1 more. GPT-6.1 Sol (Codex CLI) leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 22 ties and 43 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",0.8,1,"rate","80% (12/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.5481,0.9295],[0.7961,1],"\u0001"],["Total time per call",2.31,5.65,"seconds","2.31 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[2.17,7.73],[4.1,25.46],"\u0001"],["List-price cost per passing answer (calculation)",0.00624,0.01564,"usd","$0.0062","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",7.75,13.11,"seconds","7.75 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[2.26,34.79],[8.54,61.6],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.01435,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",23.1,113.4,"seconds","23.1 s","113.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s). A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[18.7,44.5],[78.5,221.9],"\u0001"],["List-price cost per passing coding session (calculation)",0.085,0.0978,"usd","$0.085","$0.098","unclear","No interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.63,13.11,"seconds","7.63 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.71 s to 24.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","range","minmax",[2.71,24.01],[8.54,61.6],"\u0001"],["List-price cost per strict pass by effort (calculation)",0.01352,0.02564,"usd","$0.014","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: time per call (Exact number)",6.89,13.38,"seconds","6.89 s","13.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 5.81 s to 7.81 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[5.81,7.81],[12.29,17.97],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",2.89,6.42,"seconds","2.89 s","6.42 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.68 s to 5.30 s; GPT-6.1 Sol (Codex CLI) 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[2.68,5.3],[5.25,8.26],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",2.67,11.29,"seconds","2.67 s","11.3 s","a","The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.32 s to 4.34 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[2.32,4.34],[9.08,14.85],"\u0001"]]}],["jev-1-13-vs-claude-haiku-4-5","jev-1-13","claude-haiku-4-5","Jev 1.13 vs Claude Haiku 4.5","Jev 1.13 vs Claude Haiku 4.5: measured benchmarks","Jev 1.13 vs Claude Haiku 4.5: 17 measured metrics from 3 studies (Typed routing decisions answered exactly right; more), with sample sizes and intervals.","Jev 1.13 and Claude Haiku 4.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 12,543 ms; Time per routing decision, by route (Wall time), 0.14 s vs 9.44 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).",[{"metric":"Typed routing decisions answered exactly right","aValue":0.8984,"bValue":0.8902,"unit":"rate","aDisplay":"90%","bDisplay":"89% (73/82)","winner":"tie","basis":"The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them.","studySlug":"routing-jev-vs-llm","n":82,"chartId":"routing-exact-decisions","aN":82,"bN":82,"aContext":"typed routing decisions · TypeSafe API","bContext":"typed routing decisions · via Claude Code","rangeKind":"ci95","spanKind":"ci95","aRange":[0.8191,0.9497],"bRange":[0.8044,0.9412]},{"metric":"Cost per 1,000 routing decisions","aValue":0.0337,"bValue":8.924,"unit":"usd","aDisplay":"$0.034","bDisplay":"$8.92","winner":"unclear","basis":"No interval or range was recorded for either side, so the gap ($0.034 vs $8.92, 265x) is not tested against run-to-run variation.","studySlug":"routing-jev-vs-llm","chartId":"routing-cost-per-1000","aN":246,"bN":82,"aContext":"typed routing decisions · TypeSafe API","bContext":"typed routing decisions · via Claude Code","calculation":true}]],["jev-1-13-vs-claude-sonnet-5-5","jev-1-13","claude-sonnet-5-5","Jev 1.13 vs Claude Sonnet 5.5","Jev 1.13 vs Claude Sonnet 5.5: measured benchmarks","Jev 1.13 vs Claude Sonnet 5.5: 17 measured metrics from 3 studies (Typed routing decisions answered exactly right; more), with sample sizes and intervals.","Jev 1.13 and Claude Sonnet 5.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 2,597 ms; Time per routing decision, by route (Wall time), 0.14 s vs 2.36 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).",[{"metric":"Typed routing decisions answered exactly right","aValue":0.8984,"bValue":0.939,"unit":"rate","aDisplay":"90%","bDisplay":"94% (77/82)","winner":"tie","basis":"The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","studySlug":"routing-jev-vs-llm","n":82,"chartId":"routing-exact-decisions","aN":82,"bN":82,"aContext":"typed routing decisions · TypeSafe API","bContext":"typed routing decisions · via Claude Code","rangeKind":"ci95","spanKind":"ci95","aRange":[0.8191,0.9497],"bRange":[0.8651,0.9737]},{"metric":"Cost per 1,000 routing decisions","aValue":0.0337,"bValue":4.996,"unit":"usd","aDisplay":"$0.034","bDisplay":"$5.00","winner":"unclear","basis":"No interval or range was recorded for either side, so the gap ($0.034 vs $5.00, 148x) is not tested against run-to-run variation.","studySlug":"routing-jev-vs-llm","chartId":"routing-cost-per-1000","aN":246,"bN":82,"aContext":"typed routing decisions · TypeSafe API","bContext":"typed routing decisions · via Claude Code","calculation":true}]],["claude-opus-5-5-vs-gpt-6-1-sol-codex-cli","claude-opus-5-5","gpt-6-1-sol-codex-cli","Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)","Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): benchmarks","Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): 31 measured metrics from 6 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) share 31 measured metrics and 19 list-price calculations from 7 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 12 ties and 38 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.71,5.6,"seconds","2.71 s","5.60 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.45 s to 11.8 s; GPT-6.1 Sol (Codex CLI) 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[2.45,11.78],[4.05,19.52],"\u0001"],["List-price cost per passing answer (calculation)",0.01049,0.01322,"usd","$0.010","$0.013","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.013) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · effort high · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",11.03,18.12,"seconds","11.0 s","18.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.63 s to 63.0 s; GPT-6.1 Sol (Codex CLI) 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","range","minmax",[3.63,63],[11.67,92.21],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.03337,0.01514,"usd","$0.033","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.033 vs $0.015, 2.2x) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · effort high · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Coding sessions that passed every hidden check",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Claude Opus 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.","coding-agents-head-to-head",12,"coding-agents-pass-rate",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Time per coding session",56.9,113.4,"seconds","56.9 s","113.4 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 29.8 s to 185.8 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","coding-agents-head-to-head",12,"coding-agents-wall-time",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","range","minmax",[29.8,185.8],[78.5,221.9],"\u0001"],["List-price cost per passing coding session (calculation)",0.2229,0.0978,"usd","$0.22","$0.098","unclear","No interval or range was recorded for either side, so the gap ($0.22 vs $0.098, 2.3x) is not tested against run-to-run variation.","coding-agents-head-to-head",12,"coding-agents-cost-per-pass",12,12,"Claude Code · six small repository tasks with hidden tests","Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",9.72,13.11,"seconds","9.72 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 4.78 s to 31.4 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","range","minmax",[4.78,31.36],[8.54,61.6],"\u0001"],["List-price cost per strict pass by effort (calculation)",0.02947,0.02564,"usd","$0.029","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.026) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-haiku-4-5-vs-gpt-6-1-sol-codex-cli","claude-haiku-4-5","gpt-6-1-sol-codex-cli","Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)","Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI): benchmarks","Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI): 40 measured metrics from 6 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) share 40 measured metrics and 16 list-price calculations from 7 studies. Claude Haiku 4.5 leads on 2 rows: Same prompt, 10 times: time per call (Exact number), 5.06 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 5.95 s vs 11.3 s. GPT-6.1 Sol (Codex CLI) leads on 6 rows: Pass rate on eight hard tasks (Strict pass), 100% (16/16) vs 46% (11/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); Same prompt, 10 times: strict pass rate (JSON object), 100% (10/10) vs 10% (1/10); and 3 more. On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 13 ties and 35 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",4.43,5.65,"seconds","4.43 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[3.16,23.57],[4.1,25.46],"\u0001"],["List-price cost per passing answer (calculation)",0.00836,0.01564,"usd","$0.0084","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.016) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (16/16)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; GPT-6.1 Sol (Codex CLI) 81% to 100%).","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Haiku 4.5 47% to 82%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,13.11,"seconds","39.0 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[15.27,75.13],[8.54,61.6],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.02564,"usd","$0.067","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.026, 2.6x) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Same prompt, 10 times: strict pass rate (Exact number)",0,1,"rate","0% (0/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; GPT-6.1 Sol (Codex CLI) 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0,0.2775],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (JSON object)",0.1,1,"rate","10% (1/10)","100% (10/10)","b","The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; GPT-6.1 Sol (Codex CLI) 72% to 100%).","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.0179,0.4042],[0.7225,1],"\u0001"],["Same prompt, 10 times: strict pass rate (Code fix)",1,1,"rate","100% (10/10)","100% (10/10)","tie","The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.","caching-consistency",10,"consistency-pass-rate",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","ci95","ci95",[0.7225,1],[0.7225,1],"\u0001"],["Same prompt, 10 times: time per call (Exact number)",5.06,13.38,"seconds","5.06 s","13.4 s","a","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.42 s to 6.20 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[4.42,6.2],[12.29,17.97],"\u0001"],["Same prompt, 10 times: time per call (JSON object)",7.03,6.42,"seconds","7.03 s","6.42 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.28 s to 12.3 s; GPT-6.1 Sol (Codex CLI) 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[5.28,12.27],[5.25,8.26],"\u0001"],["Same prompt, 10 times: time per call (Code fix)",5.95,11.29,"seconds","5.95 s","11.3 s","a","The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.","caching-consistency",10,"consistency-latency-spread",10,10,"Claude Code · same prompt repeated 10 times","Codex CLI · effort medium · same prompt repeated 10 times","range","minmax",[4.89,7.33],[9.08,14.85],"\u0001"]]}],["claude-fable-5-1-vs-gpt-6-1-sol-codex-cli","claude-fable-5-1","gpt-6-1-sol-codex-cli","Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI)","Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI): benchmarks","Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI): 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 4 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Fable 5.1 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",1.94,5.65,"seconds","1.94 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 1.41 s to 9.83 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[1.41,9.83],[4.1,25.46],"\u0001"],["List-price cost per passing answer (calculation)",0.02054,0.01564,"usd","$0.021","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.021 vs $0.016) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",16.13,13.11,"seconds","16.1 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 4.46 s to 90.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[4.46,90],[8.54,61.6],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.09331,0.02564,"usd","$0.093","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.093 vs $0.026, 3.6x) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true]]}],["jev-1-13-vs-deterministic-routing-policy","jev-1-13","deterministic-routing-policy","Jev 1.13 vs Deterministic routing policy","Jev 1.13 vs Deterministic routing policy: benchmarks","Jev 1.13 vs Deterministic routing policy: 3 measured metrics from one study (Time to make one routing decision; more), with sample sizes and intervals.","Jev 1.13 and Deterministic routing policy share 3 measured metrics and 4 list-price calculations from one study. Deterministic routing policy leads on 1 row: Time to make one routing decision, 1.42 µs vs 137 ms. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",[{"metric":"Cost per 1,000 routing decisions: no model call vs provider-reported","aValue":0.0337,"bValue":0,"unit":"usd","aDisplay":"$0.034","bDisplay":"$0.00","winner":"unclear","basis":"No interval or range was recorded for either side, so the gap ($0.034 vs $0.00) is not tested against run-to-run variation.","studySlug":"routing-overhead","chartId":"router-overhead-cost-reported","aN":82,"bN":20000,"aContext":"routing overhead per decision · TypeSafe API","bContext":"Agent · in process · routing overhead per decision"}]],["claude-haiku-4-5-vs-claude-fable-5-1","claude-haiku-4-5","claude-fable-5-1","Claude Haiku 4.5 vs Claude Fable 5.1","Claude Haiku 4.5 vs Claude Fable 5.1: measured benchmarks","Claude Haiku 4.5 vs Claude Fable 5.1: 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","Claude Haiku 4.5 and Claude Fable 5.1 share 12 measured metrics and 11 list-price calculations from 4 studies. Claude Fable 5.1 leads on 2 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24). On those rows the 95% intervals do not overlap. The other rows are 1 tie and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Fable 5.1 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",4.43,1.94,"seconds","4.43 s","1.94 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Fable 5.1 1.41 s to 9.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[3.16,23.57],[1.41,9.83],"\u0001"],["List-price cost per passing answer (calculation)",0.00836,0.02054,"usd","$0.0084","$0.021","unclear","No interval or range was recorded for either side, so the gap ($0.0084 vs $0.021, 2.5x) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",0.4583,1,"rate","46% (11/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Fable 5.1 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.2789,0.6493],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",0.6667,1,"rate","67% (16/24)","100% (24/24)","b","The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Fable 5.1 86% to 100%).","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.4671,0.8203],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",39.01,16.13,"seconds","39.0 s","16.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Fable 5.1 4.46 s to 90.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[15.27,75.13],[4.46,90],"\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.0672,0.09331,"usd","$0.067","$0.093","unclear","No interval or range was recorded for either side, so the gap ($0.067 vs $0.093) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true]]}]]}}