{"note":"Per entity, the best-supported metrics across studies: a 95% interval first, then a run range, then the larger sample. Each keeps its study, n, interval or range and configuration. There is no composite score: the studies differ in task, route, effort and sample size, so one weighted number would hide what each value measured.","perCategory":3,"entries":{"$k":["entity","name","kind","vendor","measuredIn","studies","factCount","best"],"$r":[["claude-sonnet-5-5","Claude Sonnet 5.5","model","Anthropic",16,["swe-bench-opus-vs-sonnet","model-head-to-head","hard-model-head-to-head","coding-agents-head-to-head","effort-ladder","caching-consistency","agent-memory","routing-jev-vs-llm","routing-overhead","cli-model-latency-tokens","single-call-vs-agent-loop","haiku-thinking-on-off","json-schema-vs-instructions","prompt-cache-break-even","routing-holdout","thinking-token-bill","llm-speed-anatomy","haiku-retry-or-escalate","harder-tasks-head-to-head"],229,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","polarity","range","calculation"],"$r":[["quality","routing-jev-vs-llm","routing-key-accuracy","Key accuracy","Claude Sonnet 5.5","routing-key-accuracy","Per-question accuracy",0.9742,"rate","97% (189/194)",194,[0.9411,0.9889],"ci95","typed routing decisions · via Claude Code","\u0001","\u0001","\u0001"],["quality","haiku-thinking-on-off","haiku-thinking-router-exact","Per-question accuracy","Claude Sonnet 5.5 (low) · Claude Code","haiku-thinking-router-exact/Per-question accuracy","Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)",0.9742,"rate","97% (189/194)",194,[0.9411,0.9889],"ci95","Claude Code · effort low · typed routing decisions, thinking on vs off","higher","\u0001","\u0001"],["quality","routing-holdout","routing-holdout-key-accuracy","Key accuracy","Claude Sonnet 5.5 (low) · Claude Code","routing-holdout-key-accuracy","Per-question accuracy on unseen decisions",0.92,"rate","92% (115/125)",125,[0.859,0.956],"ci95","Claude Code · effort low","higher","\u0001","\u0001"],["speed","routing-jev-vs-llm","routing-decision-latency","Wall time (CLI)","Claude Sonnet 5.5","routing-decision-latency/Wall time (CLI)","Time per routing decision (Wall time (CLI))",2598,"ms","2,598 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code","\u0001",[2598,4298],"\u0001"],["speed","routing-overhead","router-overhead-decision-latency","Decision time","Claude Sonnet 5.5 (effort low, via Claude Code)","router-overhead-decision-latency","Time to make one routing decision",2597,"ms","2,597 ms",82,"\u0001","p50-p95","effort low · via Claude Code · routing overhead per decision","\u0001",[2597,4298],"\u0001"],["speed","haiku-thinking-on-off","haiku-thinking-router-latency","Wall time (CLI)","Claude Sonnet 5.5 (low) · Claude Code","haiku-thinking-router-latency/Wall time (CLI)","Haiku thinking study: time per routing decision (Wall time (CLI))",2.6,"seconds","2.60 s",82,"\u0001","p50-p95","Claude Code · effort low · typed routing decisions, thinking on vs off","\u0001",[2.6,4.3],"\u0001"],["cost","model-head-to-head","h2h-list-price-per-call","Cost per call","Claude Sonnet 5.5 · Claude Code","h2h-list-price-per-call","List-price cost per call (calculation)",0.0036,"usd","$0.0036",15,"\u0001","minmax","Claude Code · five short validated tasks","\u0001",[0.00342,0.01021],true],["cost","routing-jev-vs-llm","routing-cost-per-1000","Cost","Claude Sonnet 5.5","routing-cost-per-1000","Cost per 1,000 routing decisions",4.996,"usd","$5.00",82,"\u0001","\u0001","typed routing decisions · via Claude Code","\u0001","\u0001",true],["cost","haiku-thinking-on-off","haiku-thinking-router-cost","Cost per 1,000 decisions","Claude Sonnet 5.5 (low) · Claude Code","haiku-thinking-router-cost","Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)",7.324,"usd","$7.32",82,"\u0001","\u0001","Claude Code · effort low · typed routing decisions, thinking on vs off","\u0001","\u0001",true]]}],["claude-code-cli","Claude Code","cli","Anthropic",14,["model-head-to-head","hard-model-head-to-head","coding-agents-head-to-head","effort-ladder","caching-consistency","agent-memory","routing-overhead","cli-model-latency-tokens","single-call-vs-agent-loop","haiku-thinking-on-off","json-schema-vs-instructions","routing-holdout","thinking-token-bill","llm-speed-anatomy","haiku-retry-or-escalate","harder-tasks-head-to-head"],412,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","polarity","context","range","calculation"],"$r":[["quality","haiku-thinking-on-off","haiku-thinking-router-exact","Per-question accuracy","Claude Sonnet 5.5 (low) · Claude Code","haiku-thinking-router-exact/Per-question accuracy","Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)",0.9742,"rate","97% (189/194)",194,[0.9411,0.9889],"ci95","higher","Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off","\u0001","\u0001"],["quality","routing-holdout","routing-holdout-key-accuracy","Key accuracy","Claude Sonnet 5.5 (low) · Claude Code","routing-holdout-key-accuracy","Per-question accuracy on unseen decisions",0.92,"rate","92% (115/125)",125,[0.859,0.956],"ci95","higher","Claude Sonnet 5.5 · effort low","\u0001","\u0001"],["quality","hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Sonnet 5.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","\u0001","Claude Sonnet 5.5 · eight hard validated tasks","\u0001","\u0001"],["speed","haiku-thinking-on-off","haiku-thinking-router-latency","Wall time (CLI)","Claude Sonnet 5.5 (low) · Claude Code","haiku-thinking-router-latency/Wall time (CLI)","Haiku thinking study: time per routing decision (Wall time (CLI))",2.6,"seconds","2.60 s",82,"\u0001","p50-p95","\u0001","Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off",[2.6,4.3],"\u0001"],["speed","routing-holdout","routing-holdout-latency","Wall time","Claude Sonnet 5.5 (low) · Claude Code","routing-holdout-latency/Wall time","Time per routing decision, by route (Wall time)",2.359,"seconds","2.36 s",56,"\u0001","p50-p95","\u0001","Claude Sonnet 5.5 · effort low",[2.359,3.657],"\u0001"],["speed","hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Sonnet 5.5 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",7.75,"seconds","7.75 s",24,"\u0001","minmax","\u0001","Claude Sonnet 5.5 · eight hard validated tasks",[2.26,34.79],"\u0001"],["cost","model-head-to-head","h2h-list-price-per-call","Cost per call","Claude Sonnet 5.5 · Claude Code","h2h-list-price-per-call","List-price cost per call (calculation)",0.0036,"usd","$0.0036",15,"\u0001","minmax","\u0001","Claude Sonnet 5.5 · five short validated tasks",[0.00342,0.01021],true],["cost","haiku-thinking-on-off","haiku-thinking-router-cost","Cost per 1,000 decisions","Claude Sonnet 5.5 (low) · Claude Code","haiku-thinking-router-cost","Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)",7.324,"usd","$7.32",82,"\u0001","\u0001","\u0001","Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off","\u0001",true],["cost","routing-holdout","routing-holdout-cost-per-1000","Cost","Claude Sonnet 5.5 (low) · Claude Code","routing-holdout-cost-per-1000","Cost per 1,000 unseen routing decisions",7.244,"usd","$7.24",56,"\u0001","\u0001","\u0001","Claude Sonnet 5.5 · effort low","\u0001",true]]}],["claude-haiku-4-5","Claude Haiku 4.5","model","Anthropic",14,["swe-bench-verified","model-head-to-head","hard-model-head-to-head","caching-consistency","agent-memory","routing-jev-vs-llm","routing-overhead","cost-thought-experiments","single-call-vs-agent-loop","haiku-thinking-on-off","json-schema-vs-instructions","routing-holdout","thinking-token-bill","llm-speed-anatomy","haiku-retry-or-escalate","harder-tasks-head-to-head"],198,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","polarity","range","calculation"],"$r":[["quality","routing-jev-vs-llm","routing-key-accuracy","Key accuracy","Claude Haiku 4.5","routing-key-accuracy","Per-question accuracy",0.9433,"rate","94% (183/194)",194,[0.9013,0.968],"ci95","typed routing decisions · via Claude Code","\u0001","\u0001","\u0001"],["quality","haiku-thinking-on-off","haiku-thinking-router-exact","Per-question accuracy","Claude Haiku 4.5 (thinking off) · Claude Code","haiku-thinking-router-exact/Per-question accuracy","Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)",0.9124,"rate","91% (177/194)",194,[0.8642,0.9446],"ci95","Claude Code · thinking off · typed routing decisions, thinking on vs off","higher","\u0001","\u0001"],["quality","routing-holdout","routing-holdout-key-accuracy","Key accuracy","Claude Haiku 4.5 · Claude Code","routing-holdout-key-accuracy","Per-question accuracy on unseen decisions",0.816,"rate","82% (102/125)",125,[0.739,0.8741],"ci95","Claude Code","higher","\u0001","\u0001"],["speed","routing-jev-vs-llm","routing-decision-latency","Wall time (CLI)","Claude Haiku 4.5","routing-decision-latency/Wall time (CLI)","Time per routing decision (Wall time (CLI))",12674,"ms","12,674 ms",82,"\u0001","p50-p95","typed routing decisions · via Claude Code","\u0001",[12674,34413],"\u0001"],["speed","routing-overhead","router-overhead-decision-latency","Decision time","Claude Haiku 4.5 (thinking on, via Claude Code)","router-overhead-decision-latency","Time to make one routing decision",12543,"ms","12,543 ms",82,"\u0001","p50-p95","thinking on · via Claude Code · routing overhead per decision","\u0001",[12543,34481],"\u0001"],["speed","haiku-thinking-on-off","haiku-thinking-router-latency","Wall time (CLI)","Claude Haiku 4.5 (thinking off) · Claude Code","haiku-thinking-router-latency/Wall time (CLI)","Haiku thinking study: time per routing decision (Wall time (CLI))",4.66,"seconds","4.66 s",82,"\u0001","p50-p95","Claude Code · thinking off · typed routing decisions, thinking on vs off","\u0001",[4.66,8.18],"\u0001"],["cost","model-head-to-head","h2h-list-price-per-call","Cost per call","Claude Haiku 4.5 · Claude Code","h2h-list-price-per-call","List-price cost per call (calculation)",0.00566,"usd","$0.0057",15,"\u0001","minmax","Claude Code · five short validated tasks","\u0001",[0.00513,0.01804],true],["cost","routing-jev-vs-llm","routing-cost-per-1000","Cost","Claude Haiku 4.5","routing-cost-per-1000","Cost per 1,000 routing decisions",8.924,"usd","$8.92",82,"\u0001","\u0001","typed routing decisions · via Claude Code","\u0001","\u0001",true],["cost","haiku-thinking-on-off","haiku-thinking-router-cost","Cost per 1,000 decisions","Claude Haiku 4.5 (thinking off) · Claude Code","haiku-thinking-router-cost","Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)",3.364,"usd","$3.36",82,"\u0001","\u0001","Claude Code · thinking off · typed routing decisions, thinking on vs off","\u0001","\u0001",true]]}],["codex-cli","Codex CLI","cli","OpenAI",11,["model-head-to-head","hard-model-head-to-head","coding-agents-head-to-head","effort-ladder","caching-consistency","routing-overhead","cli-model-latency-tokens","single-call-vs-agent-loop","json-schema-vs-instructions","thinking-token-bill","llm-speed-anatomy","harder-tasks-head-to-head"],167,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","polarity","range","calculation"],"$r":[["quality","hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001"],["quality","effort-ladder","effort-ladder-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001"],["quality","single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","GPT-6 Luna (single call) · Codex CLI","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.625,"rate","63% (10/16)",16,[0.3864,0.8152],"ci95","GPT-6 Luna · single call","higher","\u0001","\u0001"],["speed","hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",13.11,"seconds","13.1 s",16,"\u0001","minmax","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001",[8.54,61.6],"\u0001"],["speed","effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-total-latency","Total time per call by effort on hard tasks",13.11,"seconds","13.1 s",16,"\u0001","minmax","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001",[8.54,61.6],"\u0001"],["speed","single-call-vs-agent-loop","agent-loop-total-time","Total time per attempt","GPT-6 Luna (single call) · Codex CLI","agent-loop-total-time","Total time per attempt: single call vs agent loop",5.16,"seconds","5.16 s",16,"\u0001","minmax","GPT-6 Luna · single call","\u0001",[3.59,11.32],"\u0001"],["cost","model-head-to-head","h2h-list-price-per-call","Cost per call","GPT-6.1 Sol (medium) · Codex CLI","h2h-list-price-per-call","List-price cost per call (calculation)",0.01018,"usd","$0.010",15,"\u0001","minmax","GPT-6.1 Sol · effort medium · five short validated tasks","\u0001",[0.0054,0.02686],true],["cost","hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.02564,"usd","$0.026",16,"\u0001","\u0001","GPT-6.1 Sol · effort medium · eight hard validated tasks","\u0001","\u0001",true],["cost","effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02564,"usd","$0.026",16,"\u0001","\u0001","GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001",true]]}],["gpt-6-1-sol-codex-cli","GPT-6.1 Sol (Codex CLI)","model","OpenAI",9,["model-head-to-head","hard-model-head-to-head","coding-agents-head-to-head","effort-ladder","caching-consistency","cli-model-latency-tokens","json-schema-vs-instructions","thinking-token-bill","llm-speed-anatomy","harder-tasks-head-to-head"],128,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","polarity","range","calculation"],"$r":[["quality","hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001"],["quality","effort-ladder","effort-ladder-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001"],["quality","harder-tasks-head-to-head","harder-h2h-pass-rate","Strict pass","GPT-6.1 Sol (medium) · Codex CLI","harder-h2h-pass-rate/Strict pass","Pass rate on 4 harder tasks (Strict pass)",0.6875,"rate","69% (11/16)",16,[0.444,0.8584],"ci95","Codex CLI · effort medium","higher","\u0001","\u0001"],["speed","hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",13.11,"seconds","13.1 s",16,"\u0001","minmax","Codex CLI · effort medium · eight hard validated tasks","\u0001",[8.54,61.6],"\u0001"],["speed","effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-total-latency","Total time per call by effort on hard tasks",13.11,"seconds","13.1 s",16,"\u0001","minmax","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001",[8.54,61.6],"\u0001"],["speed","model-head-to-head","h2h-total-latency","Total time per call","GPT-6.1 Sol (medium) · Codex CLI","h2h-total-latency","Total time per call",5.65,"seconds","5.65 s",15,"\u0001","minmax","Codex CLI · effort medium · five short validated tasks","\u0001",[4.1,25.46],"\u0001"],["cost","model-head-to-head","h2h-list-price-per-call","Cost per call","GPT-6.1 Sol (medium) · Codex CLI","h2h-list-price-per-call","List-price cost per call (calculation)",0.01018,"usd","$0.010",15,"\u0001","minmax","Codex CLI · effort medium · five short validated tasks","\u0001",[0.0054,0.02686],true],["cost","hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (medium) · Codex CLI","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.02564,"usd","$0.026",16,"\u0001","\u0001","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001",true],["cost","effort-ladder","effort-ladder-cost-per-pass","Cost per strict pass","GPT-6.1 Sol (medium) · Codex CLI","effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)",0.02564,"usd","$0.026",16,"\u0001","\u0001","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001",true]]}],["claude-opus-5-5","Claude Opus 5.5","model","Anthropic",8,["swe-bench-opus-vs-sonnet","model-head-to-head","hard-model-head-to-head","coding-agents-head-to-head","effort-ladder","caching-consistency","prompt-cache-break-even","thinking-token-bill","llm-speed-anatomy","harder-tasks-head-to-head"],127,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation","polarity"],"$r":[["quality","hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Opus 5.5 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001"],["quality","effort-ladder","effort-ladder-pass-rate","Strict pass","Claude Opus 5.5 · Claude Code","effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"ci95","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001"],["quality","model-head-to-head","h2h-pass-rate","Pass rate","Claude Opus 5.5 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Code · five short validated tasks","\u0001","\u0001","\u0001"],["speed","hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Opus 5.5 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",9.18,"seconds","9.18 s",24,"\u0001","minmax","Claude Code · eight hard validated tasks",[4.24,27.21],"\u0001","\u0001"],["speed","effort-ladder","effort-ladder-total-latency","Total time per call by effort on hard tasks","Claude Opus 5.5 · Claude Code","effort-ladder-total-latency","Total time per call by effort on hard tasks",9.18,"seconds","9.18 s",16,"\u0001","minmax","Claude Code · eight hard validated tasks, effort ladder",[4.24,27.21],"\u0001","\u0001"],["speed","model-head-to-head","h2h-total-latency","Total time per call","Claude Opus 5.5 · Claude Code","h2h-total-latency","Total time per call",2.75,"seconds","2.75 s",15,"\u0001","minmax","Claude Code · five short validated tasks",[2.47,8.91],"\u0001","\u0001"],["cost","model-head-to-head","h2h-list-price-per-call","Cost per call","Claude Opus 5.5 · Claude Code","h2h-list-price-per-call","List-price cost per call (calculation)",0.00688,"usd","$0.0069",15,"\u0001","minmax","Claude Code · five short validated tasks",[0.00592,0.02226],true,"\u0001"],["cost","hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Opus 5.5 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.02824,"usd","$0.028",24,"\u0001","\u0001","Claude Code · eight hard validated tasks","\u0001",true,"\u0001"],["cost","thinking-token-bill","thinking-bill-cost-per-call","Reasoning (output tokens)","Claude Opus 5.5 · Claude Code","thinking-bill-cost-per-call/Reasoning (output tokens)","List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.012528,"usd","$0.013",24,"\u0001","\u0001","Claude Code","\u0001",true,"none"]]}],["jev-1-13","Jev 1.13","router","TypeSafe",4,["system-one-arena","routing-jev-vs-llm","routing-overhead","routing-holdout"],57,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation"],"$r":[["quality","system-one-arena","arena-accuracy","Accuracy","Jev 1.13","arena-accuracy","Who decides right? Accuracy on 1,000+ checkable decisions",0.7681,"rate","77% (805/1048)",1048,[0.7416,0.7927],"ci95","","\u0001","\u0001"],["quality","routing-overhead","router-overhead-completed","Completed","Jev 1.13 (TypeSafe)","router-overhead-completed","Routing calls that returned a decision",1,"rate","100% (246/246)",246,[0.9846,1],"ci95","routing overhead per decision · TypeSafe API","\u0001","\u0001"],["quality","routing-jev-vs-llm","routing-key-accuracy","Key accuracy","Jev 1.13 (TypeSafe)","routing-key-accuracy","Per-question accuracy",0.9485,"rate","95% (184/194)",194,[0.9077,0.9718],"ci95","typed routing decisions · TypeSafe API","\u0001","\u0001"],["speed","routing-jev-vs-llm","routing-decision-latency","Wall time (direct API call)","Jev 1.13 (TypeSafe)","routing-decision-latency/Wall time (direct API call)","Time per routing decision (Wall time (direct API call))",136.5,"ms","137 ms",246,"\u0001","p50-p95","typed routing decisions · TypeSafe API",[136.5,195.7],"\u0001"],["speed","routing-overhead","router-overhead-decision-latency","Decision time","Jev 1.13 (TypeSafe)","router-overhead-decision-latency","Time to make one routing decision",136.5,"ms","137 ms",246,"\u0001","p50-p95","routing overhead per decision · TypeSafe API",[136.5,195.7],"\u0001"],["speed","routing-holdout","routing-holdout-latency","Wall time","Jev 1.13 (TypeSafe)","routing-holdout-latency/Wall time","Time per routing decision, by route (Wall time)",0.139,"seconds","0.14 s",168,"\u0001","p50-p95","",[0.139,0.192],"\u0001"],["cost","routing-jev-vs-llm","routing-cost-per-1000","Cost","Jev 1.13 (TypeSafe)","routing-cost-per-1000","Cost per 1,000 routing decisions",0.0337,"usd","$0.034",246,"\u0001","\u0001","typed routing decisions · TypeSafe API","\u0001",true],["cost","routing-holdout","routing-holdout-cost-per-1000","Cost","Jev 1.13 (TypeSafe)","routing-holdout-cost-per-1000","Cost per 1,000 unseen routing decisions",0.03065,"usd","$0.031",168,"\u0001","\u0001","","\u0001",true],["cost","routing-overhead","router-overhead-cost-reported","Cost per 1,000 decisions","Jev 1.13 (TypeSafe)","router-overhead-cost-reported","Cost per 1,000 routing decisions: no model call vs provider-reported",0.0337,"usd","$0.034",82,"\u0001","\u0001","routing overhead per decision · TypeSafe API","\u0001","\u0001"]]}],["gpt-6-luna-codex-cli","GPT-6 Luna (Codex CLI)","model","OpenAI",3,["cli-model-latency-tokens","single-call-vs-agent-loop","llm-speed-anatomy"],34,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","polarity","context","range","calculation"],"$r":[["quality","single-call-vs-agent-loop","agent-loop-pass-rate","Strict pass","GPT-6 Luna (single call) · Codex CLI","agent-loop-pass-rate","Strict pass rate: single call vs agent loop on eight hard tasks",0.625,"rate","63% (10/16)",16,[0.3864,0.8152],"ci95","higher","Codex CLI · single call","\u0001","\u0001"],["speed","single-call-vs-agent-loop","agent-loop-total-time","Total time per attempt","GPT-6 Luna (single call) · Codex CLI","agent-loop-total-time","Total time per attempt: single call vs agent loop",5.16,"seconds","5.16 s",16,"\u0001","minmax","\u0001","Codex CLI · single call",[3.59,11.32],"\u0001"],["speed","cli-model-latency-tokens","cli-vs-api-exact-reply-latency","Total time","Codex CLI · GPT-6 Luna · none","cli-vs-api-exact-reply-latency/Total time","CLI vs API: time for a one-line answer (Total time)",3.19,"seconds","3.19 s",5,"\u0001","minmax","\u0001","Codex CLI · effort none · fixed exact reply, 5 runs",[2.88,3.83],"\u0001"],["speed","llm-speed-anatomy","speed-anatomy-first-text","Time to first text","GPT-6 Luna (low) · Codex CLI","speed-anatomy-first-text","Time to first text: a 250-line answer, six models",3.3,"seconds","3.30 s",4,"\u0001","minmax","\u0001","Codex CLI · effort low",[3.19,3.47],"\u0001"],["cost","single-call-vs-agent-loop","agent-loop-cost-per-pass","Cost per strict pass","GPT-6 Luna (single call) · Codex CLI","agent-loop-cost-per-pass","List-price cost per strict pass: single call vs agent loop (calculation)",0.00116,"usd","$0.0012",16,"\u0001","\u0001","\u0001","Codex CLI · single call","\u0001",true]]}],["claude-fable-5-1","Claude Fable 5.1","model","Anthropic",3,["model-head-to-head","hard-model-head-to-head","thinking-token-bill","llm-speed-anatomy"],23,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range","calculation","polarity"],"$r":[["quality","hard-model-head-to-head","hard-h2h-pass-rate","Strict pass","Claude Fable 5.1 · Claude Code","hard-h2h-pass-rate/Strict pass","Pass rate on eight hard tasks (Strict pass)",1,"rate","100% (24/24)",24,[0.862,1],"ci95","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001"],["quality","model-head-to-head","h2h-pass-rate","Pass rate","Claude Fable 5.1 · Claude Code","h2h-pass-rate","Pass rate on five validated tasks",1,"rate","100% (15/15)",15,[0.7961,1],"ci95","Claude Code · five short validated tasks","\u0001","\u0001","\u0001"],["quality","thinking-token-bill","thinking-bill-share","Median call","Claude Fable 5.1 · Claude Code","thinking-bill-share","Reasoning share of output tokens per call on hard tasks (calculation)",64.24,"percent","64.2%",24,"\u0001","minmax","Claude Code",[23.44,97.19],true,"none"],["speed","hard-model-head-to-head","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Claude Fable 5.1 · Claude Code","hard-h2h-total-latency","Total time per call on hard tasks (separate batches)",16.13,"seconds","16.1 s",24,"\u0001","minmax","Claude Code · eight hard validated tasks",[4.46,90],"\u0001","\u0001"],["speed","model-head-to-head","h2h-total-latency","Total time per call","Claude Fable 5.1 · Claude Code","h2h-total-latency","Total time per call",1.94,"seconds","1.94 s",15,"\u0001","minmax","Claude Code · five short validated tasks",[1.41,9.83],"\u0001","\u0001"],["speed","llm-speed-anatomy","speed-anatomy-first-text","Time to first text","Claude Fable 5.1 · Claude Code","speed-anatomy-first-text","Time to first text: a 250-line answer, six models",4.43,"seconds","4.43 s",4,"\u0001","minmax","Claude Code",[2.27,4.64],"\u0001","\u0001"],["cost","model-head-to-head","h2h-list-price-per-call","Cost per call","Claude Fable 5.1 · Claude Code","h2h-list-price-per-call","List-price cost per call (calculation)",0.00987,"usd","$0.0099",15,"\u0001","minmax","Claude Code · five short validated tasks",[0.0049,0.05843],true,"\u0001"],["cost","hard-model-head-to-head","hard-h2h-cost-per-pass","Cost per strict pass","Claude Fable 5.1 · Claude Code","hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)",0.09331,"usd","$0.093",24,"\u0001","\u0001","Claude Code · eight hard validated tasks","\u0001",true,"\u0001"],["cost","thinking-token-bill","thinking-bill-cost-per-call","Reasoning (output tokens)","Claude Fable 5.1 · Claude Code","thinking-bill-cost-per-call/Reasoning (output tokens)","List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.053696,"usd","$0.054",24,"\u0001","\u0001","Claude Code","\u0001",true,"none"]]}],["agent-harness","Agent","harness","Agent",3,["swe-bench-verified","blind-review-head-to-head","cost-thought-experiments","coding-calibration"],23,{"$k":["category","studySlug","statId","metric","label","value","unit","display","n","ci","spanKind","context","chartId","series","point","calculation"],"$r":[["quality","blind-review-head-to-head","verdicts-ai","stat:verdicts-ai","Single critic verdicts that preferred the AI change",0.6894,"rate","69% (91/132)",132,[0.606,0.762],"ci95","blind panel: Agent change vs merged human change · blind review panel","\u0001","\u0001","\u0001","\u0001"],["quality","swe-bench-verified","\u0001","swebench-same-instance-leaderboard","Resolved rate on the same 33 SWE-bench Verified instances",0.7576,"rate","76% (25/33)",33,[0.5898,0.8717],"ci95","full pipeline on Claude Sonnet 5.5","swebench-same-instance-leaderboard","Resolved rate","Agent (Sonnet 5.5, full pipeline)","\u0001"],["speed","swe-bench-verified","median-minutes","stat:median-minutes","Median worker time per attempt",9.6,"minutes","9.6 min",33,"\u0001","\u0001","full pipeline on Claude Sonnet 5.5","\u0001","\u0001","\u0001","\u0001"],["cost","swe-bench-verified","cost-per-attempt","stat:cost-per-attempt","Agent model cost per attempt (notional)",2.81,"usd","$2.81",33,"\u0001","\u0001","full pipeline on Claude Sonnet 5.5","\u0001","\u0001","\u0001",true],["cost","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel","Recorded cost per resolved instance: Agent vs the public panel",3.706,"usd","$3.71",25,"\u0001","\u0001","full pipeline on Claude Sonnet 5.5 · notional","cost-per-resolved-agent-vs-panel","Cost per resolved instance","Agent (notional)",true],["cost","coding-calibration","cost-latest","stat:cost-latest","Notional cost, latest build, all 3 tasks",11.06,"usd","$11.06",3,"\u0001","\u0001","three real tasks, platform builds compared","\u0001","\u0001","\u0001",true]]}],["gpt-5-2","GPT 5.2","model","OpenAI",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"GPT 5.2 (high)","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.8485,"unit":"rate","display":"85% (28/33)","n":33,"ci":[0.6908,0.9335],"spanKind":"ci95","context":"effort high · public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"GPT 5.2 (high)","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":0.628,"unit":"usd","display":"$0.63","n":28,"context":"effort high · public mini-SWE-agent v2 run, same instances"}]],["gemini-3-flash","Gemini 3 Flash","model","Google",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"Gemini 3 Flash (high)","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.8182,"unit":"rate","display":"82% (27/33)","n":33,"ci":[0.6561,0.9139],"spanKind":"ci95","context":"effort high · public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"Gemini 3 Flash (high)","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":0.436,"unit":"usd","display":"$0.44","n":27,"context":"effort high · public mini-SWE-agent v2 run, same instances"}]],["glm-5","GLM 5","model","Z.ai",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"GLM 5 (high)","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.7879,"unit":"rate","display":"79% (26/33)","n":33,"ci":[0.6225,0.8932],"spanKind":"ci95","context":"effort high · public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"GLM 5 (high)","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":0.667,"unit":"usd","display":"$0.67","n":26,"context":"effort high · public mini-SWE-agent v2 run, same instances"}]],["claude-sonnet-4-5","Claude Sonnet 4.5","model","Anthropic",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"Claude 4.5 Sonnet (high)","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.7576,"unit":"rate","display":"76% (25/33)","n":33,"ci":[0.5898,0.8717],"spanKind":"ci95","context":"effort high · public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"Claude 4.5 Sonnet (high)","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":0.913,"unit":"usd","display":"$0.91","n":25,"context":"effort high · public mini-SWE-agent v2 run, same instances"}]],["claude-opus-4-5","Claude Opus 4.5","model","Anthropic",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"Claude 4.5 Opus (high)","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.7273,"unit":"rate","display":"73% (24/33)","n":33,"ci":[0.5578,0.8493],"spanKind":"ci95","context":"effort high · public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"Claude 4.5 Opus (high)","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":1.184,"unit":"usd","display":"$1.18","n":24,"context":"effort high · public mini-SWE-agent v2 run, same instances"}]],["claude-opus-4-6","Claude Opus 4.6","model","Anthropic",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"Claude 4.6 Opus","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.697,"unit":"rate","display":"70% (23/33)","n":33,"ci":[0.5266,0.8262],"spanKind":"ci95","context":"public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"Claude 4.6 Opus","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":0.875,"unit":"usd","display":"$0.88","n":23,"context":"public mini-SWE-agent v2 run, same instances"}]],["deepseek-v3-2","DeepSeek V3.2","model","DeepSeek",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"DeepSeek V3.2 (high)","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.7273,"unit":"rate","display":"73% (24/33)","n":33,"ci":[0.5578,0.8493],"spanKind":"ci95","context":"effort high · public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"DeepSeek V3.2 (high)","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":0.637,"unit":"usd","display":"$0.64","n":24,"context":"effort high · public mini-SWE-agent v2 run, same instances"}]],["minimax-m2-5","MiniMax M2.5","model","MiniMax",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"MiniMax M2.5 (high)","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.697,"unit":"rate","display":"70% (23/33)","n":33,"ci":[0.5266,0.8262],"spanKind":"ci95","context":"effort high · public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"MiniMax M2.5 (high)","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":0.107,"unit":"usd","display":"$0.11","n":23,"context":"effort high · public mini-SWE-agent v2 run, same instances"}]],["kimi-k2-5","Kimi K2.5","model","Moonshot AI",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"Kimi K2.5 (high)","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.697,"unit":"rate","display":"70% (23/33)","n":33,"ci":[0.5266,0.8262],"spanKind":"ci95","context":"effort high · public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"Kimi K2.5 (high)","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":0.256,"unit":"usd","display":"$0.26","n":23,"context":"effort high · public mini-SWE-agent v2 run, same instances"}]],["gpt-5-mini","GPT 5 mini","model","OpenAI",2,["swe-bench-verified","cost-thought-experiments"],3,[{"category":"quality","studySlug":"swe-bench-verified","chartId":"swebench-same-instance-leaderboard","series":"Resolved rate","point":"GPT 5 mini","metric":"swebench-same-instance-leaderboard","label":"Resolved rate on the same 33 SWE-bench Verified instances","value":0.6364,"unit":"rate","display":"64% (21/33)","n":33,"ci":[0.4662,0.7781],"spanKind":"ci95","context":"public mini-SWE-agent v2 run, same instances"},{"category":"cost","studySlug":"cost-thought-experiments","chartId":"cost-per-resolved-agent-vs-panel","series":"Cost per resolved instance","point":"GPT 5 mini","metric":"cost-per-resolved-agent-vs-panel","label":"Recorded cost per resolved instance: Agent vs the public panel","value":0.08,"unit":"usd","display":"$0.080","n":21,"context":"public mini-SWE-agent v2 run, same instances"}]],["google-vertex","Google Vertex AI","provider","Google",1,["inference-provider-index"],37,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-claude-haiku-4-5","series":"Input","point":"Google Vertex","metric":"provider-prices-claude-haiku-4-5/Input","label":"Claude Haiku 4.5: price per million tokens by provider (Input)","value":1,"unit":"usd","display":"$1.00","context":"Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["anthropic","Anthropic","provider","Anthropic",1,["inference-provider-index"],35,[{"category":"cost","studySlug":"inference-provider-index","chartId":"gateway-vs-direct-claude-haiku-4-5","series":"Input","point":"Anthropic (first-party list price)","metric":"gateway-vs-direct-claude-haiku-4-5/Input","label":"Claude Haiku 4.5: OpenRouter vs Anthropic list price (Input)","value":1,"unit":"usd","display":"$1.00","context":"first-party list price · Claude Haiku 4.5 · list price, snapshot 2026-10-06"}]],["azure","Azure","provider","Microsoft",1,["inference-provider-index"],33,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-claude-haiku-4-5","series":"Input","point":"Azure","metric":"provider-prices-claude-haiku-4-5/Input","label":"Claude Haiku 4.5: price per million tokens by provider (Input)","value":1,"unit":"usd","display":"$1.00","context":"Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["amazon-bedrock","Amazon Bedrock","provider","Amazon",1,["inference-provider-index"],23,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-claude-haiku-4-5","series":"Input","point":"Amazon Bedrock","metric":"provider-prices-claude-haiku-4-5/Input","label":"Claude Haiku 4.5: price per million tokens by provider (Input)","value":1,"unit":"usd","display":"$1.00","context":"Claude Haiku 4.5 · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["openrouter","OpenRouter","provider","OpenRouter",1,["inference-provider-index"],20,[{"category":"cost","studySlug":"inference-provider-index","chartId":"gateway-vs-direct-claude-haiku-4-5","series":"Input","point":"OpenRouter","metric":"gateway-vs-direct-claude-haiku-4-5/Input","label":"Claude Haiku 4.5: OpenRouter vs Anthropic list price (Input)","value":1,"unit":"usd","display":"$1.00","context":"Claude Haiku 4.5 · list price, snapshot 2026-10-06"}]],["parasail","Parasail","provider","Parasail",1,["inference-provider-index"],20,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"Parasail (fp4)","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.1,"unit":"usd","display":"$0.10","context":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["google-ai-studio","Google AI Studio","provider","Google",1,["inference-provider-index"],16,[{"category":"cost","studySlug":"inference-provider-index","chartId":"gateway-vs-direct-gemini-3-8-flash","series":"Input","point":"Google AI Studio (first-party list price)","metric":"gateway-vs-direct-gemini-3-8-flash/Input","label":"Gemini 3.8 Flash: OpenRouter vs Google list price (Input)","value":0.75,"unit":"usd","display":"$0.75","context":"first-party list price · Gemini 3.8 Flash · list price, snapshot 2026-10-06"}]],["deepinfra","DeepInfra","provider","DeepInfra",1,["inference-provider-index"],16,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"DeepInfra (bf16)","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.037,"unit":"usd","display":"$0.037","context":"bf16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["claude-platform-on-aws","Claude Platform on AWS","provider","Anthropic",1,["inference-provider-index"],15,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-claude-sonnet-5","series":"Input","point":"Claude Platform on AWS","metric":"provider-prices-claude-sonnet-5/Input","label":"Claude Sonnet 5: price per million tokens by provider (Input)","value":2,"unit":"usd","display":"$2.00","context":"Claude Sonnet 5 · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["novita","Novita AI","provider","Novita AI",1,["inference-provider-index"],15,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"Novita (fp4)","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.05,"unit":"usd","display":"$0.050","context":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["openai","OpenAI","provider","OpenAI",1,["inference-provider-index"],14,[{"category":"cost","studySlug":"inference-provider-index","chartId":"gateway-vs-direct-gpt-6-luna","series":"Input","point":"OpenAI (first-party list price)","metric":"gateway-vs-direct-gpt-6-luna/Input","label":"GPT-6 Luna: OpenRouter vs OpenAI list price (Input)","value":0.1,"unit":"usd","display":"$0.10","context":"first-party list price · GPT-6 Luna · list price, snapshot 2026-10-06"}]],["gpt-6-1-sol-openai-api","GPT-6.1 Sol (OpenAI API)","model","OpenAI",1,["cli-model-latency-tokens"],13,[{"category":"speed","studySlug":"cli-model-latency-tokens","chartId":"cli-vs-api-exact-reply-latency","series":"Total time","point":"OpenAI API · GPT-6.1 Sol · high","metric":"cli-vs-api-exact-reply-latency/Total time","label":"CLI vs API: time for a one-line answer (Total time)","value":1.52,"unit":"seconds","display":"1.52 s","n":5,"range":[1.35,2.23],"spanKind":"minmax","context":"OpenAI API · effort high · fixed exact reply, 5 runs"}]],["siliconflow","SiliconFlow","provider","SiliconFlow",1,["inference-provider-index"],12,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"SiliconFlow (fp8)","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.15,"unit":"usd","display":"$0.15","context":"fp8 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["cloudflare","Cloudflare Workers AI","provider","Cloudflare",1,["inference-provider-index"],11,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-llama-3-3-70b-instruct","series":"Input","point":"Cloudflare (fp8)","metric":"provider-prices-llama-3-3-70b-instruct/Input","label":"Llama 3.3 70B Instruct: price per million tokens by provider (Input)","value":0.293,"unit":"usd","display":"$0.29","context":"fp8 · Llama 3.3 70B Instruct · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["together","Together AI","provider","Together AI",1,["inference-provider-index"],10,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"Together","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.15,"unit":"usd","display":"$0.15","context":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["baseten","Baseten","provider","Baseten",1,["inference-provider-index"],9,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"BaseTen (fp4)","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.1,"unit":"usd","display":"$0.10","context":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["deterministic-routing-policy","Deterministic routing policy","router","Agent",1,["routing-overhead"],7,{"$k":["category","studySlug","chartId","series","point","metric","label","value","unit","display","n","ci","spanKind","context","range"],"$r":[["quality","routing-overhead","router-overhead-completed","Completed","Deterministic routing policy (Agent, in process)","router-overhead-completed","Routing calls that returned a decision",1,"rate","100% (20000/20000)",20000,[0.9998,1],"ci95","Agent · in process · routing overhead per decision","\u0001"],["speed","routing-overhead","router-overhead-decision-latency","Decision time","Deterministic routing policy (Agent, in process)","router-overhead-decision-latency","Time to make one routing decision",0.00142,"ms","1.42 µs",20000,"\u0001","p50-p95","Agent · in process · routing overhead per decision",[0.00142,0.00233]],["cost","routing-overhead","router-overhead-cost-reported","Cost per 1,000 decisions","Deterministic routing policy (Agent, in process)","router-overhead-cost-reported","Cost per 1,000 routing decisions: no model call vs provider-reported",0,"usd","$0.00",20000,"\u0001","\u0001","Agent · in process · routing overhead per decision","\u0001"]]}],["groq","Groq","provider","Groq",1,["inference-provider-index"],6,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"Groq","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.15,"unit":"usd","display":"$0.15","context":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["fireworks","Fireworks AI","provider","Fireworks AI",1,["inference-provider-index"],6,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-kimi-k3","series":"Input","point":"Fireworks","metric":"provider-prices-kimi-k3/Input","label":"Kimi K3: price per million tokens by provider (Input)","value":3,"unit":"usd","display":"$3.00","context":"Kimi K3 · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["gpt-6-luna-openai-api","GPT-6 Luna (OpenAI API)","model","OpenAI",1,["cli-model-latency-tokens"],5,[{"category":"speed","studySlug":"cli-model-latency-tokens","chartId":"cli-vs-api-exact-reply-latency","series":"Total time","point":"OpenAI API · GPT-6 Luna · none","metric":"cli-vs-api-exact-reply-latency/Total time","label":"CLI vs API: time for a one-line answer (Total time)","value":0.97,"unit":"seconds","display":"0.97 s","n":5,"range":[0.65,1.5],"spanKind":"minmax","context":"OpenAI API · effort none · fixed exact reply, 5 runs"}]],["sambanova","SambaNova","provider","SambaNova",1,["inference-provider-index"],4,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"SambaNova","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.14,"unit":"usd","display":"$0.14","context":"gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["nebius","Nebius","provider","Nebius",1,["inference-provider-index"],4,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"Nebius (fp4)","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.15,"unit":"usd","display":"$0.15","context":"fp4 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["cerebras","Cerebras","provider","Cerebras",1,["inference-provider-index"],3,[{"category":"cost","studySlug":"inference-provider-index","chartId":"provider-prices-gpt-oss-120b","series":"Input","point":"Cerebras (fp16)","metric":"provider-prices-gpt-oss-120b/Input","label":"gpt-oss-120b: price per million tokens by provider (Input)","value":0.35,"unit":"usd","display":"$0.35","context":"fp16 · gpt-oss-120b · reported by OpenRouter’s public API, snapshot 2026-10-06"}]],["claude-opus-5","Claude Opus 5","model","Anthropic",1,["blind-review-head-to-head"],1,[{"category":"quality","studySlug":"blind-review-head-to-head","chartId":"blind-review-critic-agreement","series":"Critic model","point":"Claude Opus 5 (Anthropic)","metric":"blind-review-critic-agreement","label":"Does the judge’s model family matter?","value":0.7,"unit":"rate","display":"70% (28/40)","n":40,"ci":[0.5457,0.8193],"spanKind":"ci95","context":"blind review panel"}]],["claude-fable-5","Claude Fable 5","model","Anthropic",1,["blind-review-head-to-head"],1,[{"category":"quality","studySlug":"blind-review-head-to-head","chartId":"blind-review-critic-agreement","series":"Critic model","point":"Claude Fable 5 (Anthropic)","metric":"blind-review-critic-agreement","label":"Does the judge’s model family matter?","value":0.65,"unit":"rate","display":"65% (26/40)","n":40,"ci":[0.4951,0.7787],"spanKind":"ci95","context":"blind review panel"}]],["gpt-5-5","GPT 5.5","model","OpenAI",1,["blind-review-head-to-head"],1,[{"category":"quality","studySlug":"blind-review-head-to-head","chartId":"blind-review-critic-agreement","series":"Critic model","point":"GPT 5.5 (OpenAI)","metric":"blind-review-critic-agreement","label":"Does the judge’s model family matter?","value":1,"unit":"rate","display":"100% (2/2)","n":2,"ci":[0.3424,1],"spanKind":"ci95","context":"blind review panel"}]]]}}