• Jev
  • Clef
  • Decision Models
  • System One

Jev in fair mode: 2nd of 9 at Tron and Snake, last of 9 at Othello

The same 150 ms for every answer: Jev ranks 2 of 9 in Tron and Snake, 3 of 8 in Pong, 9 of 9 in Othello.

TL;DR

  • Jev 1.13 ranks 2 of 9 in Tron and Snake, 3 of 8 in Pong and 9 of 9 in Othello.
  • Fair mode: the same simulated 150 ms for every answer, so hardware does not decide.
  • Recordings, not tests: one seed, 388 games, 2 to 4 per pairing (Pong 1).

Part 3 adds four league recordings: Tron, Snake and Pong in fair mode, and Othello, which has no clock. Part 1 was a quiz of 1,085 decisions and part 2 a five-game round robin; replay every game at /arena.

Fair mode and the leagues

  • Why. In part 2, each Pong answer took effect after its own call time on its own machine (Jev hosted, the open models on one Mac Studio), so the table mixed decision quality with hardware. Now the engine gives every model the same delay: 150 ms after the question, whatever its real call time. The models still run for real. Pong also uses a 250 ms minimum think (protocol amendment 3).
  • What it does not fix. The models differ in more than speed: the open models are 4-bit or 8-bit files on one Mac, and Jev's hardware is not public. Tron and Snake move in 250 ms steps, so every answer lands one step after the position the model read. The expert and perfect players are ceilings, not peers: in the replays the expert's answers take effect at once. Othello is turn-based and needs no fair mode.
  • The leagues. Each is a round robin with one seed. Tron: 9 players, 4 games per pairing, 144 games. Snake: 9 players, 2 per pairing, 72 games. Pong: 8 players, 1 per pairing, 28 games. Othello: 9 players, 4 per pairing, 144 games. The players are the seven models (six in Pong), the random player and one reference player: the expert player in Tron and Snake, the perfect player in Othello and Pong. Elo: K 32, start 1,000, the average of 200 seeded game orders, 95% interval from a bootstrap over games.

Where Jev ranks

LeagueJev's rankElo (95% interval)W-D-L
Tron2 of 91,219 (1,170 to 1,280)27-3-2
Snake2 of 91,103 (1,029 to 1,159)12-1-3
Pong3 of 81,042 (978 to 1,109)5-0-2
Othello9 of 9890 (800 to 972)9-1-22

Every player's Elo, interval and record is on Tron, Snake, Pong and Othello.

Tron, fair mode: Jev beats Clef-Flash 9B.
  • Tron. Only the expert player is higher (1,267, 30-2-0). Jev lost 2 of its 32 games, both to the expert, and drew 2 with the expert and 1 with Kev 4B. It won all 4 games against each of Clef 27B, Clef-Flash 9B, lev 4B, Laya, Julia-1 and the random player. Its interval overlaps none of the other models' or the random player's.
  • Snake. Behind the expert player (1,185, 16-0-0). Each pairing had only 2 games, so Jev's interval overlaps those of the expert, Clef 27B, Clef-Flash 9B, Julia-1 and the random player. It is clearly above lev 4B, Kev 4B and Laya.
  • Pong. Kev 4B is rated higher (1,070, 6-0-1) and beat Jev 3 to 0 in their one game. The perfect player is first (1,098, 7-0-0). Each player played 7 games, and Jev's interval overlaps all others except Laya's.
  • Othello. Jev and Kev 4B share the record 9-1-22, with ratings 1 point apart (Kev 4B 891). The random player is at 1,015 (902 to 1,106), Clef 27B second at 1,082 (20-2-10), the perfect player first at 1,280. Jev lost all 4 games to Clef 27B and all 4 to the perfect player, and won 2 of 4 against the random player. Six of the seven models rate below the random player, but no model's interval is clearly separate from it: read "last" as the order of point estimates.
Othello: Clef 27B beats Jev 20 to 16.

Pong: as deployed and in fair mode

In part 2, Jev won 13 of its 16 round-robin Pong games and all 6 featured games (4 against Laya, 2 against Clef 27B), as deployed. In fair mode it won 5 of the 8 games it played (7 in the league, 1 featured). Six pairings keep their direction. Two changed: Kev 4B (as deployed 1-1, fair 0-1) and Clef 27B (2-0, then 0-1 in the featured game, which ended 3 to 2). Laya lost to Jev 3 to 0 again. Each change is one game, so we cannot tell a real change from luck.

The right-zone share has no clock in it. In the fair league Jev picked the right zone on 188 of 188 graded decisions, lev 4B on 188 of 191 (98.4%) and Kev 4B on 213 of 217 (98.2%). Clef-Flash 9B managed 178 of 228 (78.1%), Julia-1 29 of 100, Laya 13 of 97 and the random player 23 of 92. Three models decide almost perfectly, so with one game per pairing the order of Kev 4B (6-0-1), Jev (5-0-2) and lev 4B (4-0-3) tells us little.

What the games show

How often a player found the best move where some option was a mistake, graded by the expert player in Tron and Snake, a depth-8 search in Othello and the ideal intercept in Pong.

GameJev 1.13Best other modelRandom player
Tron64.6% (422/653)Kev 4B 56.2% (294/523)36.2% (124/343)
Snake99.3% (1,450/1,460)Kev 4B 61.2% (128/209)48.7% (133/273)
Pong100% (188/188)lev 4B 98.4% (188/191)25.0% (23/92)
Othello26.2% (81/309)Clef 27B 30.4% (97/319)29.2% (90/308)
  • Tron. Jev has the highest share and the highest Elo of the models. Laya matched the expert on 4 of 276 decisive moves, fewer than a random mover.
  • Snake. A player that lives longer makes more decisions, so the counts differ a lot, and the share does not follow the results: Clef-Flash 9B has 16.8% and won 9 of 16 games. We do not rank by it.
  • Othello. No model found the best move clearly more often than the random player (shares 20.9% to 30.4%). Our Wilson intervals assume independent moves; moves in one game are not independent, so the true intervals are wider.
  • A guess, untested. Tron and Snake may suit a typed decision model: three turns, mostly a local choice, with the free cells straight ahead in the state. Othello needs look-ahead. Picking the most flips is not only Jev's habit: in the 14 recorded Othello games Jev did it on all 125 of its decisions, Clef 27B on all 70 (the replay lists up to five options, and the 14 games are a curated set). So the habit does not explain Jev's rank, and Jev's 32 games cannot separate it from five of the other six models.

Limits

  • Small samples; content, not tests. One seed. Tron and Othello have 4 games per pairing (32 per player), Snake 2 (16), Pong 1 (7). Snake and Pong ranks are the order of point estimates. Protocol amendment 3 (written after the Othello and Tron leagues, before Snake and Pong) says we fixed the players, game counts, seeds and fair-mode setting before the first league started.
  • No speed claim. The Mac also ran other work during the Othello, Tron and Snake leagues (erratum 4: builds, type checks, tests, dev servers, other sessions' jobs, a virtual machine), so the logged call times are not clean and we use none. Fair-mode results do not depend on call time. We do not compare Jev's speed with the local models anywhere in this post.
  • Fair is not identical. The clock is equal; the model files, hardware and training are not. The expert player is a heuristic, not a solver. The Othello perfect player is a depth-8 search that solves the last 12 squares exactly, not a proof of perfect play.
  • Coverage. Clef 27B is not in the Pong league; it played three featured fair games. Not tested: other delays than 150 ms (and 250 ms for the Pong think time), other seeds, more games per pairing, other hosted models, and the real-clock Tron and Snake pairings (2 and 3 recordings, replays only, in no standing). The Othello flips, Tron and Snake free-space and Pong arrival hints were on for every model.
  • Clean run. The four leagues hold 13,111 model moves (our sum); the engine counts 0 fallbacks, 0 illegal moves and 0 errors.

Every league game is a replay on /arena. More: Tron, the expert player beats Jev in 17.3 s, Pong, Kev 4B beats Jev 3 to 0, Pong, Jev beats Laya 3 to 0, Othello, Julia-1 beats Jev. All numbers come from leagues.json and the replay files next to it; the study page has the league table.

The data behind this post

Compare the systems in this post

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.