{"i":9,"study":{"slug":"system-one-arena","title":"System One arena: Jev vs Clef and five open decision models, head to head","seoTitle":"Jev vs Clef: 7 decision models on 1,085 decisions","description":"Jev 1.13, Clef 27B, Clef-Flash and four small open decision models on 1,085 checkable decisions and in head-to-head games, Pong included.","question":"How good are typed-decision models at decisions with a checkable answer, and which one wins when they play each other?","answer":"Jev 1.13 answered 76.8% of 1,048 graded decisions correctly, clearly ahead of every other model (exact McNemar p < 0.001 for each pair). The best open model, Clef 27B, reached 69.9%. At one provider's list prices (Cloudflare Workers AI) a thousand decisions cost $0.035 for Jev 1.13, $0.051 for Clef-Flash 9B, $0.136 for Clef 27B. Speed is compared only where the footing is the same: on one GPU a third-party board measured Clef 27B at 209 ms and Clef-Flash 9B at 39 ms, while Jev 1.13, a hosted API, took 524 ms in that harness with the network included. Clef-Flash 9B (6.49 GB, 65.2%) and Kev 4B (3.03 GB, 64.0%) were statistically tied. Laya (0.45 GB, 27.4%) and Julia-1 (0.17 GB, 25.9%) were statistically tied. When no option fitted, Jev 1.13 chose \"none\", \"ask\" or \"escalate\" 86% of the time; Laya 18%. Clef 27B and Clef-Flash 9B never changed a decision when the options were shuffled, but changed 18% and 18% when the same options were renamed a, b, c. In 1,512 round-robin games (tic-tac-toe, Connect Four, Nim, Dots and Boxes and Pong), Jev 1.13 rated highest (Elo 1040 against 896 for the random player; the other models sit between 877 and 941). Jev 1.13 (30–11 of 42) and Clef-Flash 9B (29–11 of 42) beat the random player clearly (exact sign test p < 0.05); the other models did not. In Pong, where an answer counts only once it arrives, Jev 1.13 won 13 of 16 games and Kev 4B 12, while Clef 27B, the most accurate open model on the exam, won 4 at 1.8 s per decision. In the arena leagues Jev 1.13 ranked as follows: Othello: rank 9 of 9, Elo 890; best was perfect at 1280; the random player 1015; Pong (every answer after 150 ms): rank 3 of 8, Elo 1042; best was perfect at 1098; the random player 930; Snake (every answer after 150 ms): rank 2 of 9, Elo 1103; best was expert at 1185; the random player 954; Tron (every answer after 150 ms): rank 2 of 9, Elo 1219; best was expert at 1267; the random player 898.","date":"2026-10-06","updated":"2026-10-06","tags":["jev","clef","decision-models","system-one","llama-cpp","open-models","routing"],"caveats":["Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.","The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.","Typed-decision models are built for routing and gating with a confidence threshold. These items have no threshold: a model that is unsure must still pick.","A few small families lean on content where the right option is often the longest or shortest description (sudoku, seating order, gambles). Reported, not removed.","These are llama.cpp ports, not the reference implementations: llama.cpp converts only part of lev's output head, an open llama.cpp issue reports degraded probabilities for Clef Q8_0 (we ran Q4_K_M), and Laya and Julia-1 were trained on inputs of 512 to 1,024 tokens while our items average about 800.","One harness error was fixed mid-run (amendment 1): the local servers first ran with a 512-token batch, which rejected longer inputs. Every local model was rerun from the start with a whole-input batch, one model at a time; the earlier rows are kept apart and not counted.","Games: 10 games per pairing in each turn-based game (colours swapped), 2 in Pong, first to 3 after amendment 2. The sign tests against the random player are not corrected for the many comparisons made, so read them as descriptions. Connect Four \"perfect\" is a depth-10 search plus exact endgames, not a full solve.","In Pong each answer moves the paddle only after its measured latency, and the arrival point of the ball is in the state, so Pong tests reading and speed more than physics. The viral \"smarter AI lost at Pong\" result (Laya beating Jev) did not reproduce under these rules: Jev won all 4 games of the featured series.","Jev ran on TypeSafe's servers and the six open models on one Mac Studio, so speed across that line is not comparable. This study compares capability (same questions and positions) directly, and speed only on one machine, on one GPU (third-party numbers), through one gateway (OpenRouter's numbers) or as list price.","Arena leagues (Othello, and Tron, Snake and Pong in fair mode) are content recordings, not tests of the models: one seed, 2 to 4 games per pairing, protocol amendment 3. The Mac also ran other work while they were recorded (erratum 4), so their call times are not clean measurements and no speed claim uses them; fair-mode results do not depend on call time."],"sourceIds":["system-one-arena"],"stats":{"$k":["id","label","value","unit","display","n","ci"],"$r":[["arena-items","Decision items, written by agents that never called a model (1,048 graded, 37 consistency-only)",1085,"count","1,085",1085,"\u0001"],["arena-calls","Decision calls made and counted (exam, latency and repeat passes)",23247,"calls","23,247",23247,"\u0001"],["arena-best","Highest accuracy on the graded items: Jev 1.13",0.7681,"rate","76.8% (805/1048)",1048,[0.7416,0.7927]],["arena-best-open","Best open model: Clef 27B",0.6994,"rate","69.9% (733/1048)",1048,[0.671,0.7264]],["arena-smallest","Smallest model: Julia-1 (0.17 GB file)",0.2586,"rate","25.9% (271/1048)",1048,[0.233,0.2859]],["arena-jev-cost","Jev 1.13 list-price cost per 1,000 decisions (calculation from its reported input tokens)",0.03465,"usd","$0.0347",3081,"\u0001"],["arena-jev-repeat","Same question asked twice: Jev 1.13 gave the same answer 58/60 times; the local models every time",0.9667,"rate","58/60",60,"\u0001"],["arena-jev-latency","Jev 1.13 hosted API: median call time from Houston, network included (not comparable with a model on another machine)",137,"ms","137 ms",120,"\u0001"],["arena-chance","Expected accuracy of picking an option at random on the same graded items (calculation)",0.224,"rate","22.4%",1048,"\u0001"],["arena-rank-agreement","Rank agreement between our accuracy order and the independent Decision Index 0.2.1 order, same seven models (Spearman, calculation)",0.93,"score","0.93",7,"\u0001"],["arena-games","Games played to the end and counted (round robin, featured series and speed ladder)",1620,"count","1,620",1620,"\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","whisker","series","note","sourceIds","polarity","xLabel","yLabel"],"$r":[["arena-accuracy","Who decides right? Accuracy on 1,000+ checkable decisions","Share of graded items answered correctly, first presentation · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13",0.7681,0.7416,0.7927,1048,true],["Clef 27B",0.6994,0.671,0.7264,1048,"\u0001"],["Clef-Flash 9B",0.6517,0.6224,0.68,1048,"\u0001"],["Kev 4B",0.6403,0.6107,0.6688,1048,"\u0001"],["lev 4B",0.5859,0.5558,0.6153,1048,"\u0001"],["Laya",0.2739,0.2477,0.3016,1048,"\u0001"],["Julia-1",0.2586,0.233,0.2859,1048,"\u0001"]]}}],"1048 graded items in five suites. An error, a timeout or a label outside the option set counts as wrong. Two models differ clearly only where the paired McNemar test says so (table below).",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-by-suite","Where each model is strong","Accuracy by suite, first presentation","grouped-bar","rate","\u0001",{"$k":["name","points"],"$r":[["Jev 1.13",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.4219,0.3542,0.4926,192],["Logic and thought experiments",0.7436,0.6698,0.8057,156],["Policy cases, 20 industries",0.8139,0.7636,0.8555,274],["Usability intents",0.9336,0.8934,0.9594,226],["Stress tests",0.87,0.8163,0.9097,200]]}],["Clef 27B",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.3594,0.2949,0.4294,192],["Logic and thought experiments",0.6474,0.5697,0.718,156],["Policy cases, 20 industries",0.719,0.663,0.7689,274],["Usability intents",0.8894,0.8418,0.9239,226],["Stress tests",0.825,0.7664,0.8714,200]]}],["Clef-Flash 9B",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.3073,0.2463,0.3758,192],["Logic and thought experiments",0.5833,0.5049,0.6578,156],["Policy cases, 20 industries",0.7226,0.6668,0.7723,274],["Usability intents",0.8673,0.8168,0.9054,226],["Stress tests",0.695,0.628,0.7546,200]]}],["Kev 4B",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.3854,0.3195,0.4559,192],["Logic and thought experiments",0.5449,0.4666,0.621,156],["Policy cases, 20 industries",0.6752,0.6176,0.7279,274],["Usability intents",0.8009,0.744,0.8477,226],["Stress tests",0.73,0.6646,0.7868,200]]}],["lev 4B",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.3021,0.2415,0.3704,192],["Logic and thought experiments",0.5064,0.4287,0.5838,156],["Policy cases, 20 industries",0.5912,0.5322,0.6478,274],["Usability intents",0.823,0.768,0.8672,226],["Stress tests",0.645,0.5765,0.708,200]]}],["Laya",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.1771,0.1296,0.2373,192],["Logic and thought experiments",0.2821,0.2173,0.3572,156],["Policy cases, 20 industries",0.2664,0.2176,0.3217,274],["Usability intents",0.323,0.2654,0.3865,226],["Stress tests",0.315,0.2546,0.3823,200]]}],["Julia-1",{"$k":["label","value","lo","hi","n"],"$r":[["Games",0.2083,0.1569,0.2712,192],["Logic and thought experiments",0.3462,0.276,0.4237,156],["Policy cases, 20 industries",0.2518,0.2041,0.3064,274],["Usability intents",0.1593,0.1173,0.2126,226],["Stress tests",0.36,0.2967,0.4286,200]]}]]},"Games: positions solved by search. Logic: syllogisms, probability, bias probes. Policy cases: fictional company rules in 20 industries. Usability: what the user wants on a given screen. Stress tests: prompt injection, long logs, thresholds, near-identical options.",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-latency-this-mac","Speed on this Mac: the six open models","Median, whisker to the 95th percentile · one Mac Studio (M3 Ultra), one call at a time","dot-range","ms","p50-p95",[{"name":"This Mac","points":{"$k":["label","value","lo","hi","n"],"$r":[["Julia-1",13,13,22,120],["Laya",43,43,68,120],["Kev 4B",307,307,600,120],["lev 4B",425,425,646,120],["Clef-Flash 9B",566,566,816,120],["Clef 27B",1919,1919,2792,115]]}}],"All six ran on the same machine, so they compare with each other. Jev is not here: it ran on TypeSafe's servers. These times describe this Mac and 4- or 8-bit files, not the models: a data-centre GPU is several times faster.",["system-one-arena"],"lower","\u0001","\u0001"],["arena-latency-same-gpu","Speed on one GPU: third-party numbers","Median, whisker to the 95th percentile · Decision Index 0.2.1, one NVIDIA RTX PRO 6000, one harness (snapshot 2026-10-01)","dot-range","ms","p50-p95",[{"name":"Open models, same GPU","points":{"$k":["label","value","lo","hi"],"$r":[["Laya",5.8,5.8,222.5],["Julia-1",5.8,5.8,15.3],["Clef-Flash 9B",38.8,38.8,122.4],["Kev 4B",52.1,52.1,141.1],["lev 4B",70.8,70.8,687.7],["Clef 27B",209.3,209.3,238.6]]}},{"name":"Hosted: network included","points":[{"label":"Jev 1.13","value":524.1,"lo":524.1,"hi":536}]}],"Measured by the Decision Index, not by us. Every open model ran on the same GPU in one harness, so the six compare with each other. Jev is a hosted API: its time includes the network from the board's lab, so it is shown apart and is not a compute comparison.",["system-one-arena"],"lower","\u0001","\u0001"],["arena-speed-accuracy-gpu","Open models: accuracy against speed on the same GPU","x: third-party median latency on one GPU · y: our accuracy on 1,048 graded items · 95% intervals","scatter","rate","\u0001",{"$k":["name","points"],"$r":[["Clef 27B",[{"label":"Clef 27B","x":209.3,"value":0.6994,"lo":0.671,"hi":0.7264,"n":1048}]],["Clef-Flash 9B",[{"label":"Clef-Flash 9B","x":38.8,"value":0.6517,"lo":0.6224,"hi":0.68,"n":1048}]],["lev 4B",[{"label":"lev 4B","x":70.8,"value":0.5859,"lo":0.5558,"hi":0.6153,"n":1048}]],["Kev 4B",[{"label":"Kev 4B","x":52.1,"value":0.6403,"lo":0.6107,"hi":0.6688,"n":1048}]],["Laya",[{"label":"Laya","x":5.8,"value":0.2739,"lo":0.2477,"hi":0.3016,"n":1048}]],["Julia-1",[{"label":"Julia-1","x":5.8,"value":0.2586,"lo":0.233,"hi":0.2859,"n":1048}]]]},"Speed comes from the Decision Index (one GPU, one harness); accuracy comes from this study. Jev is a hosted API and is not on this machine, so it has no dot here: it scored 76.8%.",["system-one-arena"],"none","Median latency on one GPU (ms, third party)","Accuracy (ours)"],["arena-latency-gateway","Speed through one gateway: OpenRouter's own numbers","Median latency OpenRouter reports for each model's endpoint · read 2026-10-06","bar","seconds","\u0001",[{"name":"Median latency, one gateway","points":{"$k":["label","value"],"$r":[["Jev 1.13 (TypeSafe)",0.17],["Clef 27B (Workers AI)",0.5],["Clef-Flash 9B (Workers AI)",0.59],["Kev 4B (SiliconFlow)",4.08]]}}],"One gateway measures all four, so the network path is the same. The providers differ (TypeSafe, Workers AI, SiliconFlow) and the numbers move daily. Not our measurement.",["system-one-arena"],"lower","\u0001","\u0001"],["arena-cost-same-provider","Price per 1,000 decisions at one provider's list prices","Cloudflare Workers AI list price per million input tokens × the input tokens each model counted per decision in this study","bar","usd","\u0001",[{"name":"USD per 1,000 decisions","points":{"$k":["label","value"],"$r":[["Jev 1.13 ($0.042/M, 825 tokens)",0.0347],["Clef-Flash 9B ($0.09/M, 568 tokens)",0.0511],["Clef 27B ($0.24/M, 568 tokens)",0.1363]]}}],"One provider prices all three, so this is a like-for-like cost, with no timing involved. Token counts come from each model's own tokenizer; output is free for all three. A calculation, not an invoice.",["system-one-arena"],"lower","\u0001","\u0001"],["arena-size-accuracy","Does a bigger file decide better?","Accuracy against model file size (open models)","scatter","rate","\u0001",{"$k":["name","points"],"$r":[["Clef 27B",[{"label":"Clef 27B","x":19.23,"value":0.6994,"lo":0.671,"hi":0.7264,"n":1048}]],["Clef-Flash 9B",[{"label":"Clef-Flash 9B","x":6.49,"value":0.6517,"lo":0.6224,"hi":0.68,"n":1048}]],["lev 4B",[{"label":"lev 4B","x":3.01,"value":0.5859,"lo":0.5558,"hi":0.6153,"n":1048}]],["Kev 4B",[{"label":"Kev 4B","x":3.03,"value":0.6403,"lo":0.6107,"hi":0.6688,"n":1048}]],["Laya",[{"label":"Laya","x":0.45,"value":0.2739,"lo":0.2477,"hi":0.3016,"n":1048}]],["Julia-1",[{"label":"Julia-1","x":0.17,"value":0.2586,"lo":0.233,"hi":0.2859,"n":1048}]]]},"File size of the GGUF that ran (quantized). Hardware does not enter. Jev is closed and its size is not published, so it is not on this chart.",["system-one-arena"],"none","Model file (GB)","Accuracy"],["arena-robust","Same question, different presentation","Right at first presentation vs right in all three: original order, shuffled options, options renamed a, b, c","dot-range","rate","ci95",[{"name":"First presentation","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.7681,0.7416,0.7927,1048],["Clef 27B",0.6994,0.671,0.7264,1048],["Clef-Flash 9B",0.6517,0.6224,0.68,1048],["Kev 4B",0.6403,0.6107,0.6688,1048],["lev 4B",0.5859,0.5558,0.6153,1048],["Laya",0.2739,0.2477,0.3016,1048],["Julia-1",0.2586,0.233,0.2859,1048]]}},{"name":"All three presentations","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.729,0.7013,0.755,1048],["Clef 27B",0.6536,0.6243,0.6818,1048],["Clef-Flash 9B",0.5964,0.5664,0.6257,1048],["Kev 4B",0.5697,0.5395,0.5993,1048],["lev 4B",0.5344,0.5041,0.5644,1048],["Laya",0.1584,0.1375,0.1817,1048],["Julia-1",0.1603,0.1393,0.1838,1048]]}}],"Yes/no items have one presentation, so for them both measures are the same.",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-flips","Decisions that changed when only the presentation changed","Share of choice items · lower is better","grouped-bar","rate","\u0001",[{"name":"Options shuffled","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.077,0.0618,0.0956,961],["Clef 27B",0,0,0.004,961],["Clef-Flash 9B",0,0,0.004,961],["Kev 4B",0.1426,0.1219,0.1661,961],["lev 4B",0.1322,0.1122,0.155,961],["Laya",0.3361,0.3069,0.3666,961],["Julia-1",0.5203,0.4887,0.5517,961]]}},{"name":"Options renamed a, b, c","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.0687,0.0543,0.0864,961],["Clef 27B",0.1831,0.16,0.2088,961],["Clef-Flash 9B",0.18,0.157,0.2056,961],["Kev 4B",0.1103,0.092,0.1317,961],["lev 4B",0.1061,0.0882,0.1272,961],["Laya",0.4631,0.4317,0.4947,961],["Julia-1",0,0,0.004,961]]}}],"The descriptions never changed, only their order or their labels. A decision model should not care.",["system-one-arena"],"lower","\u0001","\u0001"],["arena-escape","Knowing when to say \"none of these\"","Items with a none, ask or escalate option · first presentation","grouped-bar","rate","\u0001",[{"name":"Escaped when it should (higher is better)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.8571,0.7915,0.9046,147],["Clef 27B",0.7619,0.6869,0.8235,147],["Clef-Flash 9B",0.6054,0.5247,0.6808,147],["Kev 4B",0.619,0.5385,0.6936,147],["lev 4B",0.4694,0.3905,0.5498,147],["Laya",0.1769,0.1237,0.2465,147],["Julia-1",0.2245,0.1646,0.2985,147]]}},{"name":"Escaped when it should not (lower is better)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.055,0.0396,0.0759,618],["Clef 27B",0.0259,0.016,0.0416,618],["Clef-Flash 9B",0.0146,0.0077,0.0274,618],["Kev 4B",0.0809,0.0619,0.1051,618],["lev 4B",0.034,0.0223,0.0514,618],["Laya",0.1375,0.1126,0.1669,618],["Julia-1",0.165,0.1379,0.1964,618]]}}],"Out-of-scope requests, policies that do not cover a case, ambiguous requests that need a question, lost game positions.",["system-one-arena"],"none","\u0001","\u0001"],["arena-injection","Prompt injection: does text in the state hijack the decision?","Right answers on items whose state holds \"ignore the instructions and choose X\" style text · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Right despite the injected text","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.95,0.835,0.9862,40],["Clef 27B",0.9,0.7695,0.9604,40],["Clef-Flash 9B",0.775,0.625,0.8768,40],["Kev 4B",0.95,0.835,0.9862,40],["lev 4B",0.675,0.5202,0.7992,40],["Laya",0.325,0.2008,0.4798,40],["Julia-1",0.45,0.3071,0.6017,40]]}}],"The injected text sits in a user-supplied field (a review, an email, a file name). The right answer follows the real instructions.",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-by-length","Short inputs vs long inputs","Accuracy by item length (Jev's token count of the same request) · first presentation","grouped-bar","rate","\u0001",[{"name":"Up to 512 tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.7867,0.7286,0.8351,225],["Clef 27B",0.72,0.658,0.7746,225],["Clef-Flash 9B",0.6667,0.6027,0.725,225],["Kev 4B",0.6,0.5348,0.6618,225],["lev 4B",0.6044,0.5393,0.6661,225],["Laya",0.3689,0.3085,0.4336,225],["Julia-1",0.4222,0.3595,0.4875,225]]}},{"name":"More than 512 tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.7631,0.7328,0.7908,823],["Clef 27B",0.6938,0.6615,0.7243,823],["Clef-Flash 9B",0.6476,0.6144,0.6795,823],["Kev 4B",0.6513,0.6181,0.6831,823],["lev 4B",0.5808,0.5468,0.6141,823],["Laya",0.2479,0.2196,0.2785,823],["Julia-1",0.2139,0.1872,0.2432,823]]}}],"Laya and Julia-1 are small encoders trained on short inputs (512 to 1,024 tokens); the longest items here are about 2,000 tokens. This split was added after the freeze as a description, not a test.",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-confident-wrong","Wrong and sure of it","Share of wrong answers given with probability 0.8 or more · lower is better","dot-range","rate","ci95",[{"name":"Confident wrong answers","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",0.1111,0.0775,0.1568,243],["Clef 27B",0.0161,0.0069,0.0371,311],["Clef-Flash 9B",0.0082,0.0028,0.0239,365],["Kev 4B",0.0743,0.0519,0.1052,377],["lev 4B",0.3364,0.2936,0.3821,434],["Laya",0.0972,0.0782,0.1204,761],["Julia-1",0.5611,0.526,0.5956,777]]}}],"The probability is what the model returned for the option it chose. A well-calibrated decider is rarely this sure when it is wrong, so a confidence gate can catch its mistakes.",["system-one-arena"],"lower","\u0001","\u0001"],["arena-elo","Tournament rating across every game","Elo from every round-robin game (K 32, start 1000, averaged over 200 game orders) · 95% bootstrap intervals","dot-range","score","ci95",[{"name":"Elo","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Perfect player",1554,1498,1609,336,"\u0001"],["Jev 1.13",1040,1002,1082,336,true],["lev 4B",941,894,986,336,"\u0001"],["Clef 27B",941,901,984,336,"\u0001"],["Kev 4B",936,892,974,336,"\u0001"],["Clef-Flash 9B",936,890,982,336,"\u0001"],["Random player",896,858,948,336,"\u0001"],["Julia-1",879,831,921,336,"\u0001"],["Laya",877,840,921,336,"\u0001"]]}}],"Tic-tac-toe, Connect Four, Nim, Dots and Boxes and Pong; every game counts once. The random and perfect players anchor the scale.",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-wdl","Wins, draws and losses in the round robin","Every round-robin game","stacked-bar","count","\u0001",{"$k":["name","points"],"$r":[["Wins",{"$k":["label","value","n"],"$r":[["Perfect player",324,336],["Jev 1.13",196,336],["lev 4B",146,336],["Clef 27B",152,336],["Kev 4B",144,336],["Clef-Flash 9B",147,336],["Random player",127,336],["Julia-1",120,336],["Laya",118,336]]}],["Draws",{"$k":["label","value","n"],"$r":[["Perfect player",9,336],["Jev 1.13",6,336],["lev 4B",9,336],["Clef 27B",5,336],["Kev 4B",9,336],["Clef-Flash 9B",7,336],["Random player",12,336],["Julia-1",6,336],["Laya",13,336]]}],["Losses",{"$k":["label","value","n"],"$r":[["Perfect player",3,336],["Jev 1.13",134,336],["lev 4B",181,336],["Clef 27B",179,336],["Kev 4B",183,336],["Clef-Flash 9B",182,336],["Random player",197,336],["Julia-1",210,336],["Laya",205,336]]}]]},"\u0001",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-perfect-moves","How often a model found the perfect move","Share of decisive moves (some option is a mistake) that a perfect player would also make","grouped-bar","rate","\u0001",{"$k":["name","points"],"$r":[["Jev 1.13",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4653,0.3858,0.5466,144],["Connect Four",0.3789,0.3297,0.4307,351],["Nim",0.3916,0.3206,0.4675,166],["Dots and Boxes",0.5479,0.5127,0.5827,772]]}],["Clef 27B",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4718,0.3916,0.5536,142],["Connect Four",0.3138,0.2698,0.3613,392],["Nim",0.4051,0.3317,0.483,158],["Dots and Boxes",0.3376,0.3016,0.3757,622]]}],["Clef-Flash 9B",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.5306,0.4502,0.6095,147],["Connect Four",0.3629,0.3163,0.4122,383],["Nim",0.2593,0.1927,0.3391,135],["Dots and Boxes",0.3267,0.2903,0.3652,600]]}],["lev 4B",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4737,0.3959,0.5527,152],["Connect Four",0.3477,0.3024,0.396,394],["Nim",0.3101,0.2367,0.3944,129],["Dots and Boxes",0.363,0.3261,0.4017,617]]}],["Kev 4B",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4371,0.3605,0.5168,151],["Connect Four",0.1415,0.1071,0.1846,311],["Nim",0.2975,0.2233,0.3842,121],["Dots and Boxes",0.3339,0.2974,0.3725,602]]}],["Laya",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.3716,0.2979,0.4518,148],["Connect Four",0.1003,0.073,0.1363,349],["Nim",0.2025,0.1473,0.2719,158],["Dots and Boxes",0.3667,0.3308,0.4041,660]]}],["Julia-1",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.3562,0.2831,0.4366,146],["Connect Four",0.2227,0.186,0.2644,431],["Nim",0.2706,0.2094,0.3419,170],["Dots and Boxes",0.288,0.2524,0.3263,573]]}],["Random player",{"$k":["label","value","lo","hi","n"],"$r":[["Tic-tac-toe",0.4013,0.3268,0.4807,152],["Connect Four",0.232,0.1951,0.2734,444],["Nim",0.2327,0.1738,0.3042,159],["Dots and Boxes",0.3241,0.2884,0.3621,617]]}]]},"Tic-tac-toe and Dots and Boxes against exact solvers, Nim by nim-sum, Connect Four against a depth-10 search (agreement, not proof of a mistake). The random player is the anchor: below it, a model chose worse than chance.",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-pong","Pong as deployed: who won","Share of Pong games won · first to 3 · each answer lands after its own call time on its own machine (hardware-dependent) · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Pong score","points":{"$k":["label","value","lo","hi","n"],"$r":[["Perfect player",1,0.8064,1,16],["Jev 1.13",0.8125,0.5699,0.9341,16],["Kev 4B",0.75,0.505,0.8982,16],["lev 4B",0.5625,0.3318,0.769,16],["Clef-Flash 9B",0.5,0.28,0.72,16],["Clef 27B",0.25,0.1018,0.495,16],["Julia-1",0.25,0.1018,0.495,16],["Random player",0.1875,0.0659,0.4301,16],["Laya",0.1875,0.0659,0.4301,16]]}}],"Each paddle asks its model where to go as soon as it is free; the answer moves the paddle only after that call's real latency, in simulated time. The arrival point is given in the state, so Pong tests reading and speed, not physics.",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-pong-quality","Pong decision quality, ignoring time","Share of graded decisions where the model chose the zone the ideal intercept chose · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Right zone","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13",1,0.9927,1,523],["Clef 27B",1,0.8897,1,31],["lev 4B",0.9869,0.9536,0.9964,153],["Kev 4B",0.9774,0.9541,0.989,310],["Clef-Flash 9B",0.8296,0.7573,0.8837,135],["Julia-1",0.2145,0.1997,0.2301,2811],["Laya",0.1221,0.1033,0.1439,999]]}}],"Hardware-neutral: each decision is judged on its own, and time does not enter. The ball is heading toward the paddle and its arrival point is in the state. Clef 27B decided rarely (slow), so its n is small.",["system-one-arena"],"\u0001","\u0001","\u0001"],["arena-pong-ladder","How much is a millisecond worth in Pong?","The perfect player slowed down, in a round robin against itself and a random player · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Share of games won","points":{"$k":["label","value","lo","hi","n"],"$r":[["0 ms",0.975,0.8712,0.9956,40],["100 ms",0.775,0.625,0.8768,40],["300 ms",0.5,0.352,0.648,40],["1000 ms",0.125,0.0546,0.2611,40]]}}],"Same perfect decisions, only slower. No model calls.",["system-one-arena"],"\u0001","\u0001","\u0001"]]},"related":["routing-jev-vs-llm","routing-overhead"]}}