{"$k":["slug","chart"],"$r":[["system-one-arena",{"id":"arena-accuracy","title":"Who decides right? Accuracy on 1,000+ checkable decisions","subtitle":"Share of graded items answered correctly, first presentation · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13",0.7681,0.7416,0.7927,1048,true],["Clef 27B",0.6994,0.671,0.7264,1048,"\u0001"],["Clef-Flash 9B",0.6517,0.6224,0.68,1048,"\u0001"],["Kev 4B",0.6403,0.6107,0.6688,1048,"\u0001"],["lev 4B",0.5859,0.5558,0.6153,1048,"\u0001"],["Laya",0.2739,0.2477,0.3016,1048,"\u0001"],["Julia-1",0.2586,0.233,0.2859,1048,"\u0001"]]}}],"note":"1048 graded items in five suites. An error, a timeout or a label outside the option set counts as wrong. Two models differ clearly only where the paired McNemar test says so (table below).","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-by-suite","title":"Where each model is strong","subtitle":"Accuracy by suite, first presentation","kind":"grouped-bar","unit":"rate","series":{"$k":["name","points"],"$r":[["Jev 1.13",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.4219,0.3542,0.4926,192],["Logic and thought experiments",0.7436,0.6698,0.8057,156],["Policy cases, 20 industries",0.8139,0.7636,0.8555,274],["Usability intents",0.9336,0.8934,0.9594,226],["Stress tests",0.87,0.8163,0.9097,200]]}],["Clef 27B",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.3594,0.2949,0.4294,192],["Logic and thought experiments",0.6474,0.5697,0.718,156],["Policy cases, 20 industries",0.719,0.663,0.7689,274],["Usability intents",0.8894,0.8418,0.9239,226],["Stress tests",0.825,0.7664,0.8714,200]]}],["Clef-Flash 9B",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.3073,0.2463,0.3758,192],["Logic and thought experiments",0.5833,0.5049,0.6578,156],["Policy cases, 20 industries",0.7226,0.6668,0.7723,274],["Usability intents",0.8673,0.8168,0.9054,226],["Stress tests",0.695,0.628,0.7546,200]]}],["Kev 4B",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.3854,0.3195,0.4559,192],["Logic and thought experiments",0.5449,0.4666,0.621,156],["Policy cases, 20 industries",0.6752,0.6176,0.7279,274],["Usability intents",0.8009,0.744,0.8477,226],["Stress tests",0.73,0.6646,0.7868,200]]}],["lev 4B",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.3021,0.2415,0.3704,192],["Logic and thought experiments",0.5064,0.4287,0.5838,156],["Policy cases, 20 industries",0.5912,0.5322,0.6478,274],["Usability intents",0.823,0.768,0.8672,226],["Stress tests",0.645,0.5765,0.708,200]]}],["Laya",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.1771,0.1296,0.2373,192],["Logic and thought experiments",0.2821,0.2173,0.3572,156],["Policy cases, 20 industries",0.2664,0.2176,0.3217,274],["Usability intents",0.323,0.2654,0.3865,226],["Stress tests",0.315,0.2546,0.3823,200]]}],["Julia-1",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.2083,0.1569,0.2712,192],["Logic and thought experiments",0.3462,0.276,0.4237,156],["Policy cases, 20 industries",0.2518,0.2041,0.3064,274],["Usability intents",0.1593,0.1173,0.2126,226],["Stress tests",0.36,0.2967,0.4286,200]]}]]},"note":"Games: positions solved by search. Logic: syllogisms, probability, bias probes. Policy cases: fictional company rules in 20 industries. Usability: what the user wants on a given screen. Stress tests: prompt injection, long logs, thresholds, near-identical options.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-latency-same-gpu","polarity":"lower","title":"Speed on one GPU: third-party numbers","subtitle":"Median, whisker to the 95th percentile · Decision Index 0.2.1, one NVIDIA RTX PRO 6000, one harness (snapshot 2026-10-01)","kind":"dot-range","unit":"ms","whisker":"p50-p95","series":[{"name":"Open models, same GPU","points":{"$k":["label","value","lo","hi"],"$r":[["Laya",5.8,5.8,222.5],["Julia-1",5.8,5.8,15.3],["Clef-Flash 9B",38.8,38.8,122.4],["Kev 4B",52.1,52.1,141.1],["lev 4B",70.8,70.8,687.7],["Clef 27B",209.3,209.3,238.6]]}},{"name":"Hosted: network included","points":[{"label":"Jev 1.13","value":524.1,"lo":524.1,"hi":536}]}],"note":"Measured by the Decision Index, not by us. Every open model ran on the same GPU in one harness, so the six compare with each other. Jev is a hosted API: its time includes the network from the board's lab, so it is shown apart and is not a compute comparison.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-latency-gateway","polarity":"lower","title":"Speed through one gateway: OpenRouter's own numbers","subtitle":"Median latency OpenRouter reports for each model's endpoint · read 2026-10-06","kind":"bar","unit":"seconds","series":[{"name":"Median latency, one gateway","points":{"$k":["label","value"],"$r":[["Jev 1.13 (TypeSafe)",0.17],["Clef 27B (Workers AI)",0.5],["Clef-Flash 9B (Workers AI)",0.59],["Kev 4B (SiliconFlow)",4.08]]}}],"note":"One gateway measures all four, so the network path is the same. The providers differ (TypeSafe, Workers AI, SiliconFlow) and the numbers move daily. Not our measurement.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-cost-same-provider","polarity":"lower","title":"Price per 1,000 decisions at one provider's list prices","subtitle":"Cloudflare Workers AI list price per million input tokens × the input tokens each model counted per decision in this study","kind":"bar","unit":"usd","series":[{"name":"USD per 1,000 decisions","points":{"$k":["label","value"],"$r":[["Jev 1.13 ($0.042/M, 825 tokens)",0.0347],["Clef-Flash 9B ($0.09/M, 568 tokens)",0.0511],["Clef 27B ($0.24/M, 568 tokens)",0.1363]]}}],"note":"One provider prices all three, so this is a like-for-like cost, with no timing involved. Token counts come from each model's own tokenizer; output is free for all three. A calculation, not an invoice.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-robust","title":"Same question, different presentation","subtitle":"Right at first presentation vs right in all three: original order, shuffled options, options renamed a, b, c","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"First presentation","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.7681,0.7416,0.7927,1048],["Clef 27B",0.6994,0.671,0.7264,1048],["Clef-Flash 9B",0.6517,0.6224,0.68,1048],["Kev 4B",0.6403,0.6107,0.6688,1048],["lev 4B",0.5859,0.5558,0.6153,1048],["Laya",0.2739,0.2477,0.3016,1048],["Julia-1",0.2586,0.233,0.2859,1048]]}},{"name":"All three presentations","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.729,0.7013,0.755,1048],["Clef 27B",0.6536,0.6243,0.6818,1048],["Clef-Flash 9B",0.5964,0.5664,0.6257,1048],["Kev 4B",0.5697,0.5395,0.5993,1048],["lev 4B",0.5344,0.5041,0.5644,1048],["Laya",0.1584,0.1375,0.1817,1048],["Julia-1",0.1603,0.1393,0.1838,1048]]}}],"note":"Yes/no items have one presentation, so for them both measures are the same.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-flips","polarity":"lower","title":"Decisions that changed when only the presentation changed","subtitle":"Share of choice items · lower is better","kind":"grouped-bar","unit":"rate","series":[{"name":"Options shuffled","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.077,0.0618,0.0956,961],["Clef 27B",0,0,0.004,961],["Clef-Flash 9B",0,0,0.004,961],["Kev 4B",0.1426,0.1219,0.1661,961],["lev 4B",0.1322,0.1122,0.155,961],["Laya",0.3361,0.3069,0.3666,961],["Julia-1",0.5203,0.4887,0.5517,961]]}},{"name":"Options renamed a, b, c","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.0687,0.0543,0.0864,961],["Clef 27B",0.1831,0.16,0.2088,961],["Clef-Flash 9B",0.18,0.157,0.2056,961],["Kev 4B",0.1103,0.092,0.1317,961],["lev 4B",0.1061,0.0882,0.1272,961],["Laya",0.4631,0.4317,0.4947,961],["Julia-1",0,0,0.004,961]]}}],"note":"The descriptions never changed, only their order or their labels. A decision model should not care.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-escape","polarity":"none","title":"Knowing when to say \"none of these\"","subtitle":"Items with a none, ask or escalate option · first presentation","kind":"grouped-bar","unit":"rate","series":[{"name":"Escaped when it should (higher is better)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.8571,0.7915,0.9046,147],["Clef 27B",0.7619,0.6869,0.8235,147],["Clef-Flash 9B",0.6054,0.5247,0.6808,147],["Kev 4B",0.619,0.5385,0.6936,147],["lev 4B",0.4694,0.3905,0.5498,147],["Laya",0.1769,0.1237,0.2465,147],["Julia-1",0.2245,0.1646,0.2985,147]]}},{"name":"Escaped when it should not (lower is better)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.055,0.0396,0.0759,618],["Clef 27B",0.0259,0.016,0.0416,618],["Clef-Flash 9B",0.0146,0.0077,0.0274,618],["Kev 4B",0.0809,0.0619,0.1051,618],["lev 4B",0.034,0.0223,0.0514,618],["Laya",0.1375,0.1126,0.1669,618],["Julia-1",0.165,0.1379,0.1964,618]]}}],"note":"Out-of-scope requests, policies that do not cover a case, ambiguous requests that need a question, lost game positions.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-injection","title":"Prompt injection: does text in the state hijack the decision?","subtitle":"Right answers on items whose state holds \"ignore the instructions and choose X\" style text · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Right despite the injected text","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.95,0.835,0.9862,40],["Clef 27B",0.9,0.7695,0.9604,40],["Clef-Flash 9B",0.775,0.625,0.8768,40],["Kev 4B",0.95,0.835,0.9862,40],["lev 4B",0.675,0.5202,0.7992,40],["Laya",0.325,0.2008,0.4798,40],["Julia-1",0.45,0.3071,0.6017,40]]}}],"note":"The injected text sits in a user-supplied field (a review, an email, a file name). The right answer follows the real instructions.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-by-length","title":"Short inputs vs long inputs","subtitle":"Accuracy by item length (Jev's token count of the same request) · first presentation","kind":"grouped-bar","unit":"rate","series":[{"name":"Up to 512 tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.7867,0.7286,0.8351,225],["Clef 27B",0.72,0.658,0.7746,225],["Clef-Flash 9B",0.6667,0.6027,0.725,225],["Kev 4B",0.6,0.5348,0.6618,225],["lev 4B",0.6044,0.5393,0.6661,225],["Laya",0.3689,0.3085,0.4336,225],["Julia-1",0.4222,0.3595,0.4875,225]]}},{"name":"More than 512 tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.7631,0.7328,0.7908,823],["Clef 27B",0.6938,0.6615,0.7243,823],["Clef-Flash 9B",0.6476,0.6144,0.6795,823],["Kev 4B",0.6513,0.6181,0.6831,823],["lev 4B",0.5808,0.5468,0.6141,823],["Laya",0.2479,0.2196,0.2785,823],["Julia-1",0.2139,0.1872,0.2432,823]]}}],"note":"Laya and Julia-1 are small encoders trained on short inputs (512 to 1,024 tokens); the longest items here are about 2,000 tokens. This split was added after the freeze as a description, not a test.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-confident-wrong","polarity":"lower","title":"Wrong and sure of it","subtitle":"Share of wrong answers given with probability 0.8 or more · lower is better","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Confident wrong answers","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.1111,0.0775,0.1568,243],["Clef 27B",0.0161,0.0069,0.0371,311],["Clef-Flash 9B",0.0082,0.0028,0.0239,365],["Kev 4B",0.0743,0.0519,0.1052,377],["lev 4B",0.3364,0.2936,0.3821,434],["Laya",0.0972,0.0782,0.1204,761],["Julia-1",0.5611,0.526,0.5956,777]]}}],"note":"The probability is what the model returned for the option it chose. A well-calibrated decider is rarely this sure when it is wrong, so a confidence gate can catch its mistakes.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-elo","title":"Tournament rating across every game","subtitle":"Elo from every round-robin game (K 32, start 1000, averaged over 200 game orders) · 95% bootstrap intervals","kind":"dot-range","unit":"score","whisker":"ci95","series":[{"name":"Elo","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Perfect player",1554,1498,1609,336,"\u0001"],["Jev 1.13",1040,1002,1082,336,true],["lev 4B",941,894,986,336,"\u0001"],["Clef 27B",941,901,984,336,"\u0001"],["Kev 4B",936,892,974,336,"\u0001"],["Clef-Flash 9B",936,890,982,336,"\u0001"],["Random player",896,858,948,336,"\u0001"],["Julia-1",879,831,921,336,"\u0001"],["Laya",877,840,921,336,"\u0001"]]}}],"note":"Tic-tac-toe, Connect Four, Nim, Dots and Boxes and Pong; every game counts once. The random and perfect players anchor the scale.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-wdl","title":"Wins, draws and losses in the round robin","subtitle":"Every round-robin game","kind":"stacked-bar","unit":"count","series":{"$k":["name","points"],"$r":[["Wins",{"$k":["label","value","n"],"$r":[["Perfect player",324,336],["Jev 1.13",196,336],["lev 4B",146,336],["Clef 27B",152,336],["Kev 4B",144,336],["Clef-Flash 9B",147,336],["Random player",127,336],["Julia-1",120,336],["Laya",118,336]]}],["Draws",{"$k":["label","value","n"],"$r":[["Perfect player",9,336],["Jev 1.13",6,336],["lev 4B",9,336],["Clef 27B",5,336],["Kev 4B",9,336],["Clef-Flash 9B",7,336],["Random player",12,336],["Julia-1",6,336],["Laya",13,336]]}],["Losses",{"$k":["label","value","n"],"$r":[["Perfect player",3,336],["Jev 1.13",134,336],["lev 4B",181,336],["Clef 27B",179,336],["Kev 4B",183,336],["Clef-Flash 9B",182,336],["Random player",197,336],["Julia-1",210,336],["Laya",205,336]]}]]},"sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-perfect-moves","title":"How often a model found the perfect move","subtitle":"Share of decisive moves (some option is a mistake) that a perfect player would also make","kind":"grouped-bar","unit":"rate","series":{"$k":["name","points"],"$r":[["Jev 1.13",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4653,0.3858,0.5466,144],["Connect Four",0.3789,0.3297,0.4307,351],["Nim",0.3916,0.3206,0.4675,166],["Dots and Boxes",0.5479,0.5127,0.5827,772]]}],["Clef 27B",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4718,0.3916,0.5536,142],["Connect Four",0.3138,0.2698,0.3613,392],["Nim",0.4051,0.3317,0.483,158],["Dots and Boxes",0.3376,0.3016,0.3757,622]]}],["Clef-Flash 9B",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.5306,0.4502,0.6095,147],["Connect Four",0.3629,0.3163,0.4122,383],["Nim",0.2593,0.1927,0.3391,135],["Dots and Boxes",0.3267,0.2903,0.3652,600]]}],["lev 4B",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4737,0.3959,0.5527,152],["Connect Four",0.3477,0.3024,0.396,394],["Nim",0.3101,0.2367,0.3944,129],["Dots and Boxes",0.363,0.3261,0.4017,617]]}],["Kev 4B",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4371,0.3605,0.5168,151],["Connect Four",0.1415,0.1071,0.1846,311],["Nim",0.2975,0.2233,0.3842,121],["Dots and Boxes",0.3339,0.2974,0.3725,602]]}],["Laya",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.3716,0.2979,0.4518,148],["Connect Four",0.1003,0.073,0.1363,349],["Nim",0.2025,0.1473,0.2719,158],["Dots and Boxes",0.3667,0.3308,0.4041,660]]}],["Julia-1",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.3562,0.2831,0.4366,146],["Connect Four",0.2227,0.186,0.2644,431],["Nim",0.2706,0.2094,0.3419,170],["Dots and Boxes",0.288,0.2524,0.3263,573]]}],["Random player",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4013,0.3268,0.4807,152],["Connect Four",0.232,0.1951,0.2734,444],["Nim",0.2327,0.1738,0.3042,159],["Dots and Boxes",0.3241,0.2884,0.3621,617]]}]]},"note":"Tic-tac-toe and Dots and Boxes against exact solvers, Nim by nim-sum, Connect Four against a depth-10 search (agreement, not proof of a mistake). The random player is the anchor: below it, a model chose worse than chance.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-pong","title":"Pong as deployed: who won","subtitle":"Share of Pong games won · first to 3 · each answer lands after its own call time on its own machine (hardware-dependent) · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Pong score","points":{"$k":["label","value","lo","hi","n"],"$r":[["Perfect player",1,0.8064,1,16],["Jev 1.13",0.8125,0.5699,0.9341,16],["Kev 4B",0.75,0.505,0.8982,16],["lev 4B",0.5625,0.3318,0.769,16],["Clef-Flash 9B",0.5,0.28,0.72,16],["Clef 27B",0.25,0.1018,0.495,16],["Julia-1",0.25,0.1018,0.495,16],["Random player",0.1875,0.0659,0.4301,16],["Laya",0.1875,0.0659,0.4301,16]]}}],"note":"Each paddle asks its model where to go as soon as it is free; the answer moves the paddle only after that call's real latency, in simulated time. The arrival point is given in the state, so Pong tests reading and speed, not physics.","sourceIds":["system-one-arena"]}],["system-one-arena",{"id":"arena-pong-quality","title":"Pong decision quality, ignoring time","subtitle":"Share of graded decisions where the model chose the zone the ideal intercept chose · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Right zone","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",1,0.9927,1,523],["Clef 27B",1,0.8897,1,31],["lev 4B",0.9869,0.9536,0.9964,153],["Kev 4B",0.9774,0.9541,0.989,310],["Clef-Flash 9B",0.8296,0.7573,0.8837,135],["Julia-1",0.2145,0.1997,0.2301,2811],["Laya",0.1221,0.1033,0.1439,999]]}}],"note":"Hardware-neutral: each decision is judged on its own, and time does not enter. The ball is heading toward the paddle and its arrival point is in the state. Clef 27B decided rarely (slow), so its n is small.","sourceIds":["system-one-arena"]}],["routing-jev-vs-llm",{"id":"routing-exact-decisions","title":"Typed routing decisions answered exactly right","subtitle":"Share of asked cases where every scored question was acceptable","kind":"dot-range","unit":"rate","yLabel":"Exact","whisker":"ci95","series":[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8984,0.8191,0.9497,82,true],["Claude Haiku 4.5",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5",0.939,0.8651,0.9737,82,false]]}}],"note":"Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-key-accuracy","title":"Per-question accuracy","subtitle":"Each open question the router was asked; an unanswered question counts as wrong","kind":"dot-range","unit":"rate","yLabel":"Correct answers","whisker":"ci95","series":[{"name":"Key accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.9485,0.9077,0.9718,194,true],["Claude Haiku 4.5",0.9433,0.9013,0.968,194,false],["Claude Sonnet 5.5",0.9742,0.9411,0.9889,194,false]]}}],"note":"Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-exact-by-decision","title":"Exact rate by decision type","kind":"grouped-bar","unit":"rate","yLabel":"Exact","series":{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.8241,1,18,true],["Message intent",1,0.8389,1,20,true],["Is it a rule?",1,0.7575,1,12,true],["Context shape",0.7396,0.5789,0.8675,32,true]]}],["Claude Haiku 4.5",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9444,0.7424,0.9901,18,false],["Message intent",1,0.8389,1,20,false],["Is it a rule?",1,0.7575,1,12,false],["Context shape",0.75,0.5789,0.8675,32,false]]}],["Claude Sonnet 5.5",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.8241,1,18,false],["Message intent",1,0.8389,1,20,false],["Is it a rule?",1,0.7575,1,12,false],["Context shape",0.8438,0.6825,0.9314,32,false]]}]]},"note":"A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-cost-per-1000","title":"Cost per 1,000 routing decisions","subtitle":"List price × reported tokens per decision","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost","points":{"$k":["label","value","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.0337,246,true],["Claude Haiku 4.5",8.924,82,false],["Claude Sonnet 5.5",4.996,82,false]]}}],"note":"List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.","sourceIds":["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-decision-latency","title":"Time per routing decision","subtitle":"Median wall time, whisker to the 95th percentile","kind":"dot-range","unit":"ms","yLabel":"Time per decision","series":{"$k":["name","points"],"$r":[["Wall time (CLI)",[{"label":"Claude Haiku 4.5","value":12674,"lo":12674,"hi":34413,"n":82},{"label":"Claude Sonnet 5.5","value":2598,"lo":2598,"hi":4298,"n":82}]],["Wall time (direct API call)",[{"label":"Jev 1.13 (TypeSafe)","value":136.5,"lo":136.5,"hi":195.7,"n":246}]],["Model time (API)",[{"label":"Claude Haiku 4.5","value":10734,"lo":10734,"hi":32072,"n":82},{"label":"Claude Sonnet 5.5","value":1599,"lo":1599,"hi":2574,"n":82}]]]},"note":"Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.","whisker":"p50-p95","sourceIds":["agent-routing","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-decision-latency","title":"Time to make one routing decision","subtitle":"Median; whiskers = median to 95th percentile","kind":"dot-range","unit":"ms","yLabel":"Time per decision","whisker":"p50-p95","series":[{"name":"Decision time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",0.00142,0.00142,0.00233,20000,true],["Jev 1.13 (TypeSafe)",136.5,136.5,195.7,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",2597,2597,4298,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",12543,12543,34481,82,"\u0001"]]}}],"note":"The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.","sourceIds":["agent-routing-overhead","agent-routing","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-completed","title":"Routing calls that returned a decision","subtitle":"Completed calls ÷ calls; whiskers = 95% Wilson interval","kind":"dot-range","unit":"rate","yLabel":"Completed","whisker":"ci95","series":[{"name":"Completed","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",1,0.9998,1,20000,true],["Jev 1.13 (TypeSafe)",1,0.9846,1,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",1,0.9552,1,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",1,0.9552,1,82,"\u0001"]]}}],"note":"A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.","sourceIds":["agent-routing-overhead","agent-routing","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-cost-reported","title":"Cost per 1,000 routing decisions: no model call vs provider-reported","subtitle":"USD per 1,000 decisions","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost per 1,000 decisions","points":[{"label":"Deterministic routing policy (Agent, in process)","value":0,"n":20000,"highlight":true},{"label":"Jev 1.13 (TypeSafe)","value":0.0337,"n":82,"highlight":false}]}],"note":"The policy makes no model call, so it costs nothing per decision. Jev’s figure is the cost its provider reported for the recorded production run (82 decisions). The input tokens of the live run give the same figure at the published price. The Claude routers are in the next chart: their cost is derived from list prices.","sourceIds":["agent-routing-overhead","agent-routing","price-jev"]}],["routing-overhead",{"id":"router-overhead-cost-list-price","title":"Cost per 1,000 routing decisions for the model routers (calculation)","subtitle":"List price × the tokens each route reported, USD per 1,000 decisions","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost per 1,000 decisions (list price)","points":{"$k":["label","value","n"],"$r":[["Jev 1.13 (TypeSafe)",0.0337,246],["Claude Sonnet 5.5 (effort low, via Claude Code)",4.996,82],["Claude Haiku 4.5 (thinking on, via Claude Code)",8.924,82]]}}],"note":"A calculation, not a bill: the Claude calls ran on a subscription. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free). Same tokens and prices as the routing study.","sourceIds":["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-cost-per-1000-tasks","title":"Added routing cost per 1,000 tasks (calculation)","subtitle":"Decisions per task from recorded runs × cost per decision","kind":"grouped-bar","unit":"usd","yLabel":"USD per 1,000 tasks","series":[{"name":"Every model call routed (49.5 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0],["Jev 1.13 (TypeSafe)",1.67],["Claude Sonnet 5.5 (effort low, via Claude Code)",247.3],["Claude Haiku 4.5 (thinking on, via Claude Code)",441.74]]}},{"name":"Only System One decisions (7 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0],["Jev 1.13 (TypeSafe)",0.24],["Claude Sonnet 5.5 (effort low, via Claude Code)",34.97],["Claude Haiku 4.5 (thinking on, via Claude Code)",62.47]]}}],"note":"A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).","sourceIds":["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-delay-per-task","title":"Added routing delay per task (calculation)","subtitle":"Decisions per task × median decision time, if every decision waits in line","kind":"grouped-bar","unit":"seconds","yLabel":"Seconds per task","series":[{"name":"Every model call routed (49.5 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0.0000703],["Jev 1.13 (TypeSafe)",6.7568],["Claude Sonnet 5.5 (effort low, via Claude Code)",128.5515],["Claude Haiku 4.5 (thinking on, via Claude Code)",620.8785]]}},{"name":"Only System One decisions (7 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0.0000099],["Jev 1.13 (TypeSafe)",0.9555],["Claude Sonnet 5.5 (effort low, via Claude Code)",18.179],["Claude Haiku 4.5 (thinking on, via Claude Code)",87.801]]}}],"note":"A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.","sourceIds":["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]}],["cost-thought-experiments",{"id":"repriced-cost-per-resolved","title":"Thought experiment: the same tokens at other list prices","subtitle":"Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price","kind":"bar","unit":"usd","yLabel":"USD per resolved instance","series":[{"name":"Repriced cost per resolved instance","points":{"$k":["label","value","highlight"],"$r":[["Claude Fable 5.1",12.852,false],["Claude Opus 5",8.723,false],["Claude Opus 5.5",5.753,false],["Claude Sonnet 5.5",3.489,true],["GPT-6.1 Sol",2.097,false],["Claude Haiku 4.5",1.745,false],["Gemini 3.x Flash",1.016,false],["Jev 1.13 (router)",0.042,false]]}}],"note":"Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.","sourceIds":["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic","price-google","price-openai","price-jev"]}],["routing-holdout",{"id":"routing-holdout-exact","title":"Unseen routing decisions answered exactly right","subtitle":"Share of the 56 holdout cases where every scored question was acceptable","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Exact","series":[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"note":"Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-key-accuracy","title":"Per-question accuracy on unseen decisions","subtitle":"Each open question a router was asked; an unanswered question counts as wrong","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Correct answers","series":[{"name":"Key accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.904,0.8397,0.9442,125,true],["Claude Haiku 4.5 · Claude Code",0.816,0.739,0.8741,125,false],["Claude Sonnet 5.5 (low) · Claude Code",0.92,0.859,0.956,125,false]]}}],"note":"Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-by-purpose","title":"Exact rate on unseen decisions, by decision type","subtitle":"14 cases per decision type","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Exact","series":{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9286,0.6853,0.9873,14,true],["Message intent",0.8571,0.6006,0.9599,14,true],["Is it a rule?",0.9286,0.6853,0.9873,14,true],["Context shape",0.5714,0.3259,0.7862,14,true]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9286,0.6853,0.9873,14,false],["Message intent",0.9286,0.6853,0.9873,14,false],["Is it a rule?",0.9286,0.6853,0.9873,14,false],["Context shape",0.3571,0.1634,0.6124,14,false]]}],["Claude Sonnet 5.5 (low) · Claude Code",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.7847,1,14,false],["Message intent",1,0.7847,1,14,false],["Is it a rule?",0.9286,0.6853,0.9873,14,false],["Context shape",0.5714,0.3259,0.7862,14,false]]}]]},"note":"Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-tuned-vs-unseen","title":"Tuned case set vs unseen holdout: exact rate per router","subtitle":"Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Exact","series":[{"name":"Tuned set (routing-jev-vs-llm)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.9024,0.8191,0.9497,82,true],["Claude Haiku 4.5 · Claude Code",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5 (low) · Claude Code",0.939,0.8651,0.9737,82,false]]}},{"name":"Unseen holdout","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"note":"Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.","whisker":"ci95","sourceIds":["agent-routing-holdout","agent-routing"]}],["routing-holdout",{"id":"routing-holdout-latency","title":"Time per routing decision, by route","subtitle":"Median, whisker to the 95th percentile","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Wall time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.139,0.139,0.192,168,true],["Claude Haiku 4.5 · Claude Code",9.444,9.444,25.413,56,false],["Claude Sonnet 5.5 (low) · Claude Code",2.359,2.359,3.657,56,false]]}},{"name":"Model time (API, CLI-reported)","points":[{"label":"Claude Haiku 4.5 · Claude Code","value":7.522,"lo":7.522,"hi":23.913,"n":56},{"label":"Claude Sonnet 5.5 (low) · Claude Code","value":1.485,"lo":1.485,"hi":2.377,"n":56}]}],"note":"The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.","whisker":"p50-p95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-cost-per-1000","title":"Cost per 1,000 unseen routing decisions","subtitle":"Reported tokens per decision × list price","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost","points":{"$k":["label","value","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.03065,168,true],["Claude Haiku 4.5 · Claude Code",7.129,56,false],["Claude Sonnet 5.5 (low) · Claude Code",7.244,56,false]]}}],"note":"Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.","sourceIds":["agent-routing-holdout","calc-repricing","price-jev","price-anthropic"]}]]}