• Jev
  • Clef
  • Decision Models
  • System One

Decision models play each other: Jev vs Clef at Connect Four, Nim and Pong

Seven decision models played 1,620 games. Only Jev and Clef-Flash clearly beat random; speed alone won nothing.

TL;DR

  • Jev 1.13 rated highest: Elo 1,040, against 896 for the random player.
  • Only Jev and Clef-Flash 9B clearly beat random: 30 wins to 11 and 29 to 11 over 42 games each.
  • In Pong as deployed, speed alone won nothing, and neither did accuracy alone.

In part 1 the seven models took a quiz of 1,085 decisions. Here they play each other at tic-tac-toe, Connect Four, Nim, Dots and Boxes and real-time Pong: 1,620 games, every one counted, all on the study page.

How they played

  • Players. The seven models from part 1 (Jev over TypeSafe's API; the six open models in llama.cpp on one Mac Studio, M3 Ultra, 96 GB). The random player is seeded. The perfect player is an exact solver (tic-tac-toe, Nim, Dots and Boxes), a depth-10 search with exact endgames (Connect Four) or an ideal intercept (Pong).
  • One decision per move. The model sees the board and the legal moves and picks one. A move outside the list, or an error, becomes a seeded random legal move and is counted: 0 illegal moves and 2 errors (one Jev, one Clef 27B) in 24,345 model moves across the 1,512 round-robin games.
  • Round robin. Tic-tac-toe, Connect Four, Nim and Dots and Boxes: 10 games per pairing (colours swapped), 360 games each. Pong: 2 games per pairing.
  • Pong is real time. A 60 Hz simulation. An answer takes effect only after that call's measured latency, in simulated time. The ball's arrival point is in the state, so Pong tests reading and speed more than physics. First to 3 (declared before any Pong game; first to 5 needed about three hours).
  • Declared first: protocol and game counts, before the first counted call.

Results

Perfect player
Jev 1.13
lev 4B
Clef 27B
Kev 4B
Clef-Flash 9B
Random player
Julia-1
Laya

9 rows. Highest Perfect player 1,554 (95% interval 1,498–1,609, n 336). Lowest Laya 877 (95% interval 840–921, n 336). Not all intervals overlap.

NotesWhiskers: 95% intervaln = 336 per row

Elo from every round-robin game (K 32, start 1000, averaged over 200 game orders) · 95% bootstrap intervals

Tic-tac-toe, Connect Four, Nim, Dots and Boxes and Pong; every game counts once. The random and perfect players anchor the scale.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Jev leads with Elo 1,040. The six other models sit between 877 and 941, the random player at 896 and the perfect player at 1,554. Over 42 games each (10 per game, 2 in Pong), against the random player Jev won 30, lost 11 and drew 1; Clef-Flash won 29, lost 11 and drew 2. Both are clear (exact sign test p below 0.01). None of the other five is clear (wins to losses: Clef 27B 24 to 18, Laya 22 to 17, lev 4B 21 to 21, Kev 4B 17 to 23, Julia-1 17 to 24). Jev is also ahead of Clef-Flash (30 to 11), Clef 27B (30 to 12), Laya (28 to 12) and Julia-1 (29 to 12), but not clearly of Kev 4B (25 to 17) or lev 4B (24 to 18). The tests are not corrected for multiple comparisons: read them as description.

  • Tic-tac-toe and Nim were a coin flip. Jev won 37 of 80 tic-tac-toe games (random player 33) and 31 of 80 Nim games (random 29). The study page has every model's record in each game.
  • Connect Four and Dots and Boxes separated the models. Jev won 56 of 80 (70%) and 59 of 80 (74%).
  • Two models read Connect Four worse than random. Kev 4B (14%) and Laya (10%) found the best Connect Four move less often than the random player (23%). Jev found the best Dots and Boxes move 55% of the time against 32% for random.
  • The quiz predicted the extremes: the two tiniest models finished last in both. The middle is too noisy to rank models 5 points apart.
Connect Four after four random opening moves: Jev (first) beats Clef 27B in 27 moves. Best move found: Jev 6 of 12, Clef 4 of 10.

Pong: a millisecond is worth something

Each Pong answer takes effect after its own call time on its own machine, and Jev's includes the network. These results are as deployed: what each setup did, not what each model can do. Two measures stay fair: the right-zone share (judged per decision, no clock) and the speed ladder (the same perfect decisions, only delayed). Part 3 reports fair mode, where every answer takes the same delay.

  • Fast and wrong. Julia-1 (12 ms median, 7,451 decisions) picked the right zone on 21% of its 2,811 graded decisions (603), Laya (36 ms, 2,460 decisions) on 12% of 999 (122). They won 4 and 3 of 16 games, about what the random player wins (3 of 16).
  • Right and late. Clef 27B picked the right zone every time it was graded (31 of 31), but it answers in 1.8 s, decides only 84 times, returns 26% of shots and won 4 of 16. On a data-centre GPU it would answer about nine times faster (209 vs 1,919 ms); Julia-1 and Laya would still be wrong.
  • In between (median, right zone, games, shots returned). Jev: 132 ms, 523 of 523, 13-3, 86%. Kev 4B: 285 ms, 98% (303/310), 12-4, 80%. lev 4B: 418 ms, 99% (151/153), 9-7, 59%. Clef-Flash 9B: 528 ms, 83% (112/135), 8-8, 49%.
0 ms
100 ms
300 ms
1000 ms

4 rows. Highest 0 ms 98% (95% interval 87%–100%, n 40). Lowest 1000 ms 13% (95% interval 5.5%–26%, n 40). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 40 per row

The perfect player slowed down, in a round robin against itself and a random player · 95% Wilson intervals

Same perfect decisions, only slower. No model calls.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

To isolate speed, we slowed the perfect player: the same perfect decisions, delayed. At 0 ms it won 97.5% of its games; at 100 ms, 77.5%; at 300 ms, 50%; at 1 second, 12.5%, the same as the random player. A perfect decision a second late is worth nothing. Clef 27B's median answer, 1.8 s, is longer than our slowest rung.

The rematch. A LinkedIn post showed Laya beating Jev 5 to 1 at Pong under "the smarter AI lost" (Jev 93% right at 311 ms, Laya 78% right at 38 ms). We played the same pairing 4 times. Jev won all four, 3 to 0 each time. In the featured match Jev decides in about 124 ms and Laya in 37 ms. We do not know the original setup, and our rules differ (first to 3, zone answers, arrival hint, a different Mac), so the claim is sensitive to setup. Jev also beat Clef 27B 2 to 0; Clef 27B and Clef-Flash split 1 to 1.

The rematch, as deployed: Jev beats Laya 3 to 0.

What this means for a router

A router makes one decision, not a game, so ignore the tic-tac-toe record. Two points hold up:

  1. A fast wrong answer is not a fast right answer. The two fastest models were right on 21% and 12% of their graded decisions. Gate on accuracy before latency.
  2. A slow right answer can still lose. In a loop with a clock, a 1.8 s model must be much better than a 130 ms one. Clef 27B was as accurate as anyone in Pong and lost 12 of 16.

Limits, changes and errors

  • Ten games per pairing is small. Elo intervals are about 80 points either side, and a model "not clearly above random" may be a little better.
  • Models read the board as text; another format might play better.
  • Jev ran hosted and the open models quantized on one Mac, so Pong as deployed is not a like-for-like speed test. The ladder and part 3's fair mode remove the machine.
  • Amendment 2 (before any Pong game): Pong to 3 instead of 5; featured series cut to 4 games for Jev vs Laya and 2 for each other pair. Amendment 1 is in part 1.
  • Errata. Erratum 2: the verifier reported 10 problems among 14,928 re-solved moves, all in Connect Four: the reference solver reused deeper search results for shallower queries. We fixed it and regraded all 3,815 Connect Four moves: 14 optimal sets and 4 flags changed; no pick, probability, game result or rating changed. The verifier now reports 0 problems across 1,620 matches, 14,928 re-solved moves and 151,325 Pong frames. Erratum 3 (wording): a few low-priority data scripts of ours ran during Pong, about a minute of one core in total; we measured no effect.

The data behind this post

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.