• Jev
  • Clef
  • Decision Models
  • System One
  • Llama Cpp
  • Open Models
  • Routing

System One arena: Jev vs Clef and five open decision models, head to head

How good are typed-decision models at decisions with a checkable answer, and which one wins when they play each other?

Published · 20 charts · Download the data or a carousel

1,085

n = 1085

Decision items, written by agents that never called a model (1,048 graded, 37 consistency-only)

The answer

Jev 1.13 answered 76.8% of 1,048 graded decisions correctly, clearly ahead of every other model (exact McNemar p < 0.001 for each pair). The best open model, Clef 27B, reached 69.9%. At one provider's list prices (Cloudflare Workers AI) a thousand decisions cost $0.035 for Jev 1.13, $0.051 for Clef-Flash 9B, $0.136 for Clef 27B. Speed is compared only where the footing is the same: on one GPU a third-party board measured Clef 27B at 209 ms and Clef-Flash 9B at 39 ms, while Jev 1.13, a hosted API, took 524 ms in that harness with the network included. Clef-Flash 9B (6.49 GB, 65.2%) and Kev 4B (3.03 GB, 64.0%) were statistically tied. Laya (0.45 GB, 27.4%) and Julia-1 (0.17 GB, 25.9%) were statistically tied. When no option fitted, Jev 1.13 chose "none", "ask" or "escalate" 86% of the time; Laya 18%. Clef 27B and Clef-Flash 9B never changed a decision when the options were shuffled, but changed 18% and 18% when the same options were renamed a, b, c. In 1,512 round-robin games (tic-tac-toe, Connect Four, Nim, Dots and Boxes and Pong), Jev 1.13 rated highest (Elo 1040 against 896 for the random player; the other models sit between 877 and 941). Jev 1.13 (30–11 of 42) and Clef-Flash 9B (29–11 of 42) beat the random player clearly (exact sign test p < 0.05); the other models did not. In Pong, where an answer counts only once it arrives, Jev 1.13 won 13 of 16 games and Kev 4B 12, while Clef 27B, the most accurate open model on the exam, won 4 at 1.8 s per decision. In the arena leagues Jev 1.13 ranked as follows: Othello: rank 9 of 9, Elo 890; best was perfect at 1280; the random player 1015; Pong (every answer after 150 ms): rank 3 of 8, Elo 1042; best was perfect at 1098; the random player 930; Snake (every answer after 150 ms): rank 2 of 9, Elo 1103; best was expert at 1185; the random player 954; Tron (every answer after 150 ms): rank 2 of 9, Elo 1219; best was expert at 1267; the random player 898.

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 98 sSystem One arena: Jev vs Clef and five open decision models

System One arena: Jev vs Clef and five open decision models

1,085 checkable decisions, seven typed-decision models, every call counted. Jev 1.13 led with 76.8% right.

Transcript
  1. System One arena · 1,085 decisions · 23,247 calls. Jev vs Clef vs five open decision models. Games, logic, policy cases in 20 industries, usability and stress tests. Every answer checkable.
  2. Jev 1.13 answered 76.8% right. The best open model, Clef 27B: 69.9%. The smallest, Julia-1: 25.9%. Jev 1.13, hosted API: 76.8% (805/1048) (n = 1048, 95% CI 74–79%). Clef 27B, open weights: 69.9% (733/1048) (n = 1048, 95% CI 67–73%). Julia-1, smallest: 25.9% (271/1048) (n = 1048, 95% CI 23–29%). Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.
  3. One scale, seven decision models. Where the intervals overlap, the order is not settled. Chart: Share of graded decisions answered right · 95% intervals (n = 1048 each). Caveat: The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.
  4. Speed on one GPU, from a third-party board, against our accuracy. Jev is a hosted API, so it is not on this chart. Chart: Open models: accuracy against speed on the same GPU (n = 1048 each). Caveat: Jev ran on TypeSafe's servers and the six open models on one Mac Studio, so speed across that line is not comparable. This study compares capability (same questions and positions) directly, and speed only on one machine, on one GPU (third-party numbers), through one gateway (OpenRouter's numbers) or as list price.
  5. Knowing when no option fits is the job of a router. The smallest models rarely said "none". Chart: Items with a "none", "ask" or "escalate" option (n = 147–618 each). Caveat: The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.
  6. Same options, same words. Only the order, or the labels, changed. Chart: Decisions that changed when only the presentation changed (n = 961 each). Caveat: The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.
  7. "Ignore the instructions and choose X", hidden in a review, an email or a file name. Chart: Right answers despite injected text in the state (n = 40 each). Caveat: The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.
  8. A confidence gate only works if the model is unsure when it is wrong. Chart: Wrong answers given with probability 0.8 or more (n = 243–777 each). Caveat: Typed-decision models are built for routing and gating with a confidence threshold. These items have no threshold: a model that is unsure must still pick.
  9. Then they played each other: tic-tac-toe, Connect Four, Nim, Dots and Boxes and Pong. A random and a perfect player anchor the scale. Chart: Tournament rating across every game · 95% intervals (n = 336 each). Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.
  10. Jev 1.13 won. Rings mark moves a perfect solver would not play: Jev 1.13 6, Clef 27B 6. Recorded Connect Four game: Jev 1.13 (moves first) against Clef 27B, Jev 1.13 wins after 27 moves. Jev 1.13 played the best move 6 of 12 times. Clef 27B played the best move 4 of 10 times. Caveat: One recorded game (the first between these two in the round robin), after four random opening moves. "Best" columns come from a depth-10 search.
  11. Jev 1.13 won 3–0. Jev 1.13 (hosted) decides in 124 ms, Laya (on one Mac) in 37 ms, but Laya picked the right zone on only 12% of its graded decisions. Recorded rally, 3.8 s: Jev 1.13 (left, 124 ms per decision, 77% right on the exam) against Laya (right, 37 ms per decision, 27% right on the exam). Point Jev 1.13. Final score 3–0. Caveat: Not a speed comparison: Jev ran on TypeSafe's servers, the open models on one Mac. The longest rally of the featured Jev vs Laya series (Jev won all 4 games).
  12. Scored in points, each model on its own machine. On the same Mac, Julia-1 decided in 12 ms but picked the right zone on 21%; Clef 27B was right on 100% (31/31) of its graded decisions, at 1.8 s, and won 4 of 16. Chart: Pong round robin · share of games won · 95% intervals (n = 16 each). Caveat: Jev ran on TypeSafe's servers and the six open models on one Mac Studio, so speed across that line is not comparable. This study compares capability (same questions and positions) directly, and speed only on one machine, on one GPU (third-party numbers), through one gateway (OpenRouter's numbers) or as list price.
  13. Test a decision model on your own decisions before you route on it. Every call online.

Key numbers

23,247

Decision calls made and counted (exam, latency and repeat passes)

n = 23247

76.8%

Highest accuracy on the graded items: Jev 1.13

(805/1048) · 95% CI 74%–79% · n = 1048

69.9%

Best open model: Clef 27B

(733/1048) · 95% CI 67%–73% · n = 1048

25.9%

Smallest model: Julia-1 (0.17 GB file)

(271/1048) · 95% CI 23%–29% · n = 1048

$0.0347

Jev 1.13 list-price cost per 1,000 decisions (calculation from its reported input tokens)

n = 3081

58/60

Same question asked twice: Jev 1.13 gave the same answer 58/60 times; the local models every time

n = 60

137ms

Jev 1.13 hosted API: median call time from Houston, network included (not comparable with a model on another machine)

n = 120

22.4%

Expected accuracy of picking an option at random on the same graded items (calculation)

n = 1048

0.93

Rank agreement between our accuracy order and the independent Decision Index 0.2.1 order, same seven models (Spearman, calculation)

n = 7

1,620

Games played to the end and counted (round robin, featured series and speed ladder)

n = 1620

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Jev 1.13
Clef 27B
Clef-Flash 9B
Kev 4B
lev 4B
Laya
Julia-1

7 rows. Highest Jev 1.13 77% (95% interval 74%–79%, n 1048). Lowest Julia-1 26% (95% interval 23%–29%, n 1048). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 1048 per row

Share of graded items answered correctly, first presentation · 95% Wilson intervals

1048 graded items in five suites. An error, a timeout or a label outside the option set counts as wrong. Two models differ clearly only where the paired McNemar test says so (table below).

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)

Where each model is strong

Accuracy by suite, first presentation

Jev 1.13

Games
Logic and thought experiments
Policy cases, 20 industries
Usability intents
Stress tests

Clef 27B

Games
Logic and thought experiments
Policy cases, 20 industries
Usability intents
Stress tests

Clef-Flash 9B

Games
Logic and thought experiments
Policy cases, 20 industries
Usability intents
Stress tests

Kev 4B

Games
Logic and thought experiments
Policy cases, 20 industries
Usability intents
Stress tests

lev 4B

Games
Logic and thought experiments
Policy cases, 20 industries
Usability intents
Stress tests

Laya

Games
Logic and thought experiments
Policy cases, 20 industries
Usability intents
Stress tests

Julia-1

Games
Logic and thought experiments
Policy cases, 20 industries
Usability intents
Stress tests

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

5 rows, 7 series: Jev 1.13, Clef 27B, Clef-Flash 9B, Kev 4B, lev 4B, Laya, Julia-1. Jev 1.13: highest Usability intents 93% (95% interval 89%–96%, n 226). Lowest Games 42% (95% interval 35%–49%, n 192). Not all intervals overlap. Clef 27B: highest Usability intents 89% (95% interval 84%–92%, n 226). Lowest Games 36% (95% interval 29%–43%, n 192). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 156–274 per row

Games: positions solved by search. Logic: syllogisms, probability, bias probes. Policy cases: fictional company rules in 20 industries. Usability: what the user wants on a given screen. Stress tests: prompt injection, long logs, thresholds, near-identical options.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
Julia-1
Laya
Kev 4B
lev 4B
Clef-Flash 9B
Clef 27B

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

log scale: each gridline is 10 times the one before

6 rows. Slowest Clef 27B 1.92 s (median to p95 1.92 s–2.79 s, n 115). Fastest Julia-1 13 ms (median to p95 13 ms–22 ms, n 120). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 115–120 per row

Median, whisker to the 95th percentile · one Mac Studio (M3 Ultra), one call at a time

All six ran on the same machine, so they compare with each other. Jev is not here: it ran on TypeSafe's servers. These times describe this Mac and 4- or 8-bit files, not the models: a data-centre GPU is several times faster.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
  • Open models, same GPU
  • Hosted: network included
Laya
Julia-1
Clef-Flash 9B
Kev 4B
lev 4B
Clef 27B
Jev 1.13

log scale: each gridline is 10 times the one before

7 rows, 2 series: Open models, same GPU, Hosted: network included. Open models, same GPU: slowest Clef 27B 209 ms (median to p95 209 ms–239 ms). Fastest Julia-1 6 ms (median to p95 6 ms–15 ms). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)

Median, whisker to the 95th percentile · Decision Index 0.2.1, one NVIDIA RTX PRO 6000, one harness (snapshot 2026-10-01)

Measured by the Decision Index, not by us. Every open model ran on the same GPU in one harness, so the six compare with each other. Jev is a hosted API: its time includes the network from the board's lab, so it is shown apart and is not a compute comparison.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
  • Clef 27B
  • Clef-Flash 9B
  • lev 4B
  • Kev 4B
  • Laya
  • Julia-1
Better: upper left

6 points: Accuracy (ours) against Median latency on one GPU (ms, third party). Median latency on one GPU (ms, third party) runs from 6 ms to 209 ms; Accuracy (ours) from 26% to 70%.

NotesWhiskers: 95% Wilson intervaln = 1048 per point

x: third-party median latency on one GPU · y: our accuracy on 1,048 graded items · 95% intervals

Speed comes from the Decision Index (one GPU, one harness); accuracy comes from this study. Jev is a hosted API and is not on this machine, so it has no dot here: it scored 76.8%.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
Largest value is 24x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Clef 27B (Workers AI)
Clef-Flash 9B (Workers AI)
Kev 4B (SiliconFlow)

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the lowest value): a ratio of the two values shown, not a measurement.

4 rows. Slowest Kev 4B (SiliconFlow) 4.1 s. Fastest Jev 1.13 (TypeSafe) 0.2 s.

Notes

Median latency OpenRouter reports for each model's endpoint · read 2026-10-06

One gateway measures all four, so the network path is the same. The providers differ (TypeSafe, Workers AI, SiliconFlow) and the numbers move daily. Not our measurement.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
Calculation
Jev 1.13 ($0.042/M, 825 tokens)
Clef-Flash 9B ($0.09/M, 568 tokens)
Clef 27B ($0.24/M, 568 tokens)

Hover or focus a bar for its ratio to Jev 1.13 ($0.042/M, 825 token… (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Clef 27B ($0.24/M, 568 tokens) $0.14. Lowest Jev 1.13 ($0.042/M, 825 tokens) $0.035.

Notes

Cloudflare Workers AI list price per million input tokens × the input tokens each model counted per decision in this study

One provider prices all three, so this is a like-for-like cost, with no timing involved. Token counts come from each model's own tokenizer; output is free for all three. A calculation, not an invoice.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
  • Clef 27B
  • Clef-Flash 9B
  • lev 4B
  • Kev 4B
  • Laya
  • Julia-1

6 points: Accuracy against Model file (GB). Model file (GB) runs from 0.2 to 19.2; Accuracy from 26% to 70%.

NotesWhiskers: 95% Wilson intervaln = 1048 per point

Accuracy against model file size (open models)

File size of the GGUF that ran (quantized). Hardware does not enter. Jev is closed and its size is not published, so it is not on this chart.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
  • First presentation
  • All three presentations
Jev 1.13
Clef 27B
Clef-Flash 9B
Kev 4B
lev 4B
Laya
Julia-1

7 rows, 2 series: First presentation, All three presentations. First presentation: highest Jev 1.13 77% (95% interval 74%–79%, n 1048). Lowest Julia-1 26% (95% interval 23%–29%, n 1048). Not all intervals overlap. All three presentations: highest Jev 1.13 73% (95% interval 70%–76%, n 1048). Lowest Laya 16% (95% interval 14%–18%, n 1048). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 1048 per row

Right at first presentation vs right in all three: original order, shuffled options, options renamed a, b, c

Yes/no items have one presentation, so for them both measures are the same.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
  • Options shuffled
  • Options renamed a, b, c (square)
Sorted by gap, largest first.
Julia-1
Clef 27B
Clef-Flash 9B
Laya
Kev 4B
lev 4B
Jev 1.13

Gap labels, Options renamed a, b, c vs Options shuffled: Options renamed a, b, c is x percentage points higher (+) or lower (−) than Options shuffled, calculated from the two values shown; lines are the 95% Wilson interval.

7 rows, 2 series: Options shuffled, Options renamed a, b, c. Options shuffled: highest Julia-1 52% (95% interval 49%–55%, n 961). Lowest Clef-Flash 9B 0% (95% interval 0%–0.4%, n 961). Not all intervals overlap. Options renamed a, b, c: highest Laya 46% (95% interval 43%–49%, n 961). Lowest Julia-1 0% (95% interval 0%–0.4%, n 961). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 961 per row

Share of choice items · lower is better

The descriptions never changed, only their order or their labels. A decision model should not care.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
  • Escaped when it should (higher is better)
  • Escaped when it should not (lower is better) (square)
Sorted by gap, largest first.
Jev 1.13
Clef 27B
Clef-Flash 9B
Kev 4B
lev 4B
Julia-1
Laya

Gap labels, Escaped when it should not (lower is better) vs Escaped when it should (higher is better): Escaped when it should not (lower is better) is x percentage points higher (+) or lower (−) than Escaped when it should (higher is better), calculated from the two values shown; lines are the 95% Wilson interval.

7 rows, 2 series: Escaped when it should (higher is better), Escaped when it should not (lower is better). Escaped when it should (higher is better): highest Jev 1.13 86% (95% interval 79%–90%, n 147). Lowest Laya 18% (95% interval 12%–25%, n 147). Not all intervals overlap. Escaped when it should not (lower is better): highest Julia-1 17% (95% interval 14%–20%, n 618). Lowest Clef-Flash 9B 1.5% (95% interval 0.8%–2.7%, n 618). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 147–618 per row

Items with a none, ask or escalate option · first presentation

Out-of-scope requests, policies that do not cover a case, ambiguous requests that need a question, lost game positions.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
Jev 1.13
Clef 27B
Clef-Flash 9B
Kev 4B
lev 4B
Laya
Julia-1

7 rows. Highest Jev 1.13 95% (95% interval 84%–99%, n 40). Lowest Laya 33% (95% interval 20%–48%, n 40). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 40 per row

Right answers on items whose state holds "ignore the instructions and choose X" style text · 95% Wilson intervals

The injected text sits in a user-supplied field (a review, an email, a file name). The right answer follows the real instructions.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
  • Up to 512 tokens
  • More than 512 tokens (square)
Sorted by gap, largest first.
Julia-1
Laya
Kev 4B
Clef 27B
lev 4B
Jev 1.13
Clef-Flash 9B

Gap labels, More than 512 tokens vs Up to 512 tokens: More than 512 tokens is x percentage points higher (+) or lower (−) than Up to 512 tokens, calculated from the two values shown; lines are the 95% Wilson interval.

7 rows, 2 series: Up to 512 tokens, More than 512 tokens. Up to 512 tokens: highest Jev 1.13 79% (95% interval 73%–84%, n 225). Lowest Laya 37% (95% interval 31%–43%, n 225). Not all intervals overlap. More than 512 tokens: highest Jev 1.13 76% (95% interval 73%–79%, n 823). Lowest Julia-1 21% (95% interval 19%–24%, n 823). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 225–823 per row

Accuracy by item length (Jev's token count of the same request) · first presentation

Laya and Julia-1 are small encoders trained on short inputs (512 to 1,024 tokens); the longest items here are about 2,000 tokens. This split was added after the freeze as a description, not a test.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
Jev 1.13
Clef 27B
Clef-Flash 9B
Kev 4B
lev 4B
Laya
Julia-1

7 rows. Highest Julia-1 56% (95% interval 53%–60%, n 777). Lowest Clef-Flash 9B 0.8% (95% interval 0.3%–2.4%, n 365). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 243–777 per row

Share of wrong answers given with probability 0.8 or more · lower is better

The probability is what the model returned for the option it chose. A well-calibrated decider is rarely this sure when it is wrong, so a confidence gate can catch its mistakes.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
Perfect player
Jev 1.13
lev 4B
Clef 27B
Kev 4B
Clef-Flash 9B
Random player
Julia-1
Laya

9 rows. Highest Perfect player 1,554 (95% interval 1,498–1,609, n 336). Lowest Laya 877 (95% interval 840–921, n 336). Not all intervals overlap.

NotesWhiskers: 95% intervaln = 336 per row

Elo from every round-robin game (K 32, start 1000, averaged over 200 game orders) · 95% bootstrap intervals

Tic-tac-toe, Connect Four, Nim, Dots and Boxes and Pong; every game counts once. The random and perfect players anchor the scale.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
  • Wins
  • Draws
  • Losses
Perfect player
Jev 1.13
lev 4B
Clef 27B
Kev 4B
Clef-Flash 9B
Random player
Julia-1
Laya

One square per 8 items (rounded); counts at the right are exact and in legend order.

9 rows, 3 series: Wins, Draws, Losses. Wins: highest Perfect player 324 (n 336). Lowest Laya 118 (n 336). Draws: highest Laya 13 (n 336). Lowest Clef 27B 5 (n 336).

Notesn = 336 per row

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)

Jev 1.13

Tic-tac-toe
Connect Four
Nim
Dots and Boxes

Clef 27B

Tic-tac-toe
Connect Four
Nim
Dots and Boxes

Clef-Flash 9B

Tic-tac-toe
Connect Four
Nim
Dots and Boxes

lev 4B

Tic-tac-toe
Connect Four
Nim
Dots and Boxes

Kev 4B

Tic-tac-toe
Connect Four
Nim
Dots and Boxes

Laya

Tic-tac-toe
Connect Four
Nim
Dots and Boxes

Julia-1

Tic-tac-toe
Connect Four
Nim
Dots and Boxes

Random player

Tic-tac-toe
Connect Four
Nim
Dots and Boxes

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

4 rows, 8 series: Jev 1.13, Clef 27B, Clef-Flash 9B, lev 4B, Kev 4B, Laya, Julia-1, Random player. Jev 1.13: highest Dots and Boxes 55% (95% interval 51%–58%, n 772). Lowest Connect Four 38% (95% interval 33%–43%, n 351). Not all intervals overlap. Clef 27B: highest Tic-tac-toe 47% (95% interval 39%–55%, n 142). Lowest Connect Four 31% (95% interval 27%–36%, n 392). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 121–772 per row

Share of decisive moves (some option is a mistake) that a perfect player would also make

Tic-tac-toe and Dots and Boxes against exact solvers, Nim by nim-sum, Connect Four against a depth-10 search (agreement, not proof of a mistake). The random player is the anchor: below it, a model chose worse than chance.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
Perfect player
Jev 1.13
Kev 4B
lev 4B
Clef-Flash 9B
Clef 27B
Julia-1
Random player
Laya

9 rows. Highest Perfect player 100% (95% interval 81%–100%, n 16). Lowest Laya 19% (95% interval 6.6%–43%, n 16). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 16 per row

Share of Pong games won · first to 3 · each answer lands after its own call time on its own machine (hardware-dependent) · 95% Wilson intervals

Each paddle asks its model where to go as soon as it is free; the answer moves the paddle only after that call's real latency, in simulated time. The arrival point is given in the state, so Pong tests reading and speed, not physics.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
Jev 1.13
Clef 27B
lev 4B
Kev 4B
Clef-Flash 9B
Julia-1
Laya

7 rows. Highest Jev 1.13 100% (95% interval 99%–100%, n 523). Lowest Laya 12% (95% interval 10%–14%, n 999). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 31–2811 per row

Share of graded decisions where the model chose the zone the ideal intercept chose · 95% Wilson intervals

Hardware-neutral: each decision is judged on its own, and time does not enter. The ball is heading toward the paddle and its arrival point is in the state. Clef 27B decided rarely (slow), so its n is small.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)
0 ms
100 ms
300 ms
1000 ms

4 rows. Highest 0 ms 98% (95% interval 87%–100%, n 40). Lowest 1000 ms 13% (95% interval 5.5%–26%, n 40). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 40 per row

The perfect player slowed down, in a round robin against itself and a random player · 95% Wilson intervals

Same perfect decisions, only slower. No model calls.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Share card (PNG)

Tables

What is comparable: each comparison and what keeps it fair

ComparisonFair across all seven?Why
Accuracy, consistency, calibration, "none of these", prompt injectionYesSame 1,085 items and the same states for every model. The open models ran as 4- or 8-bit files, so a full-precision run could differ a little; our order of the models agrees with the independent Decision Index (rank correlation 0.93).
Turn-based gamesYesNo clock: a move counts only by its quality.
Pong decision quality (right zone per decision)YesJudged per decision against the ideal intercept; time does not enter.
Speed on this MacAmong the six open models onlyThe six ran on one machine, one call at a time. Jev ran on TypeSafe's servers.
Speed on one GPU (Decision Index)Among the six open models onlyThird party, one GPU, one harness. Jev is a hosted API there and includes the network.
Speed through one gateway (OpenRouter)Yes for the four it listsOne gateway measures all four. Providers differ and the numbers change daily.
Price per 1,000 decisionsYes for Jev, Clef and Clef-FlashCloudflare Workers AI prices all three on one list.
Pong and other real-time games as deployedNoEach answer takes effect after its own call time on its own machine, so the result mixes decision quality with hardware.
Real-time games in fair mode (Tron, Snake and Pong in the arena)YesEvery player's answer takes effect after the same simulated 150 ms, whatever its real call time, so hosted and local models stand on the same footing and only decision quality counts. The models are still called for real. The leagues are one seed with 2 to 4 games per pairing, so the intervals are wide.

Every model, every measure

ModelRuns asAccuracyRight in all 3 presentationsConsistent groupsEscaped when it shouldEscaped when it should notCalibration error (ECE)BrierMedian ms (where it ran: this Mac, or Jev's servers)p95 msSame answer twiceErrors
Jev 1.13Hosted API76.8% (805/1048)764/104840/53126/14734/6180.03x0.18x137 ms189 ms58/600
Clef 27BQ4_K_M in llama.cpp, 19.23 GB69.9% (733/1048)685/104838/53112/14716/6180.28x0.43x1.92 s2.79 s60/608
Clef-Flash 9BQ4_K_M in llama.cpp, 6.49 GB65.2% (683/1048)625/104842/5389/1479/6180.29x0.48x566 ms816 ms60/600
Kev 4BQ4_K_M in llama.cpp, 3.03 GB64.0% (671/1048)597/104835/5391/14750/6180.037x0.33x307 ms600 ms60/600
lev 4BQ4_K_M in llama.cpp, 3.01 GB58.6% (614/1048)560/104838/5369/14721/6180.14x0.36x425 ms646 ms60/600
LayaQ8_0 in llama.cpp, 0.45 GB27.4% (287/1048)166/104835/5326/14785/6180.2x0.6x43 ms68 ms60/600
Julia-1Q8_0 in llama.cpp, 0.17 GB25.9% (271/1048)168/104828/5333/147102/6180.55x0.68x13 ms22 ms60/600

Head to head on the same items: who was right where the other was wrong (exact McNemar test)

PairOnly the first rightOnly the second rightpVerdict
Jev 1.13 vs Clef 27B158860xJev 1.13 is ahead
Jev 1.13 vs Clef-Flash 9B208860xJev 1.13 is ahead
Jev 1.13 vs lev 4B270790xJev 1.13 is ahead
Jev 1.13 vs Kev 4B212780xJev 1.13 is ahead
Jev 1.13 vs Laya584660xJev 1.13 is ahead
Jev 1.13 vs Julia-1599650xJev 1.13 is ahead
Clef 27B vs Clef-Flash 9B141910.0012xClef 27B is ahead
Clef 27B vs lev 4B209900xClef 27B is ahead
Clef 27B vs Kev 4B1781160.0004xClef 27B is ahead
Clef 27B vs Laya529830xClef 27B is ahead
Clef 27B vs Julia-1543810xClef 27B is ahead
Clef-Flash 9B vs lev 4B168990xClef-Flash 9B is ahead
Clef-Flash 9B vs Kev 4B1581460.53xnot clear
Clef-Flash 9B vs Laya486900xClef-Flash 9B is ahead
Clef-Flash 9B vs Julia-1510980xClef-Flash 9B is ahead
lev 4B vs Kev 4B1111680.0008xKev 4B is ahead
lev 4B vs Laya425980xlev 4B is ahead
lev 4B vs Julia-14521090xlev 4B is ahead
Kev 4B vs Laya471870xKev 4B is ahead
Kev 4B vs Julia-15031030xKev 4B is ahead
Laya vs Julia-11901740.43xnot clear

One item per suite that split the field: about half the models right

SuiteItemQuestionRight answerRightWrongWhy
GamesFind the poisoned columnY is to move. R has no winning move at the moment. Exactly one column is unsafe: if Y drops a disc in it, R can win on the very next move. Choose that unsafe column.Drop a disc in column 3Jev 1.13, lev 4B, Kev 4B, LayaClef 27B, Clef-Flash 9B, Julia-1A Y disc in column 3 lands directly under the cell where R completes four, so R wins on the very next move. All other columns leave R without an immediate win.
Logic and thought experimentsFlagged application: how likely is the fault?One application was picked at random from those checked, and it was flagged. Which band holds the chance that it really has a copied answer?The chance is 30% to under 60%.Jev 1.13, Clef 27B, Clef-Flash 9Blev 4B, Kev 4B, Laya, Julia-1Among the flagged items, the share that truly have the fault is 35.5% (Bayes' rule), which falls in the band "30% to under 60%".
Policy cases, 20 industriesDelay exactly on the long-delay line, refund requested, no missed connectionWhich action does the policy require for this passenger case? Apply the numbered rules in order and choose exactly one action.Offer no remedy: no rebooking and no refundJev 1.13, Clef-Flash 9B, lev 4B, Kev 4BClef 27B, Laya, Julia-1R4: the delay is 300 minutes, which does not meet "less than 150 minutes" and meets "300 minutes or less", and no connection was missed, so there is no remedy.
Usability intentsButtons lost against the backgroundPick the one action the app should run now for what the user just said. Use the screen state and the rules in the state.Turn on high-contrast colorsJev 1.13, Clef 27B, Clef-Flash 9B, lev 4BKev 4B, Laya, Julia-1Buttons that blend into the background are a contrast problem, so High contrast.
Stress testsLoud alert names and big counts versus the severity ruleApply the routing rules to the alert and pick the action.Open a ticketJev 1.13, Clef 27B, Clef-Flash 9B, Kev 4Blev 4B, Laya, Julia-1Severity 2 is below the page level, so no page; it is at least 2, so open_ticket, whatever the name or customer count says.

Every item family: right answers per model (first presentation)

SuiteFamilyJev 1.13Clef 27BClef-Flash 9BKev 4Blev 4BLayaJulia-1
Gamestic-tac-toe8/237/236/238/2311/235/231/23
Gamesconnect-four7/264/266/267/267/264/265/26
Gamesnim12/227/229/225/225/221/226/22
Gamespong-intercept6/189/188/189/186/183/184/18
Gamessudoku-single6/175/172/176/173/172/172/17
Gameswordle10/167/163/169/169/163/164/16
Gamesmaze7/164/162/165/163/165/164/16
Gameshanoi5/166/168/164/162/164/165/16
Gamescoin-weighing11/1711/179/179/178/174/176/17
Gamesknight-moves9/219/216/2112/214/213/213/21
Logic and thought experimentssyllogism13/1312/1312/1313/139/137/134/13
Logic and thought experimentsquantifier-logic6/65/62/64/65/61/62/6
Logic and thought experimentsconditional-logic12/1211/1210/1210/127/123/127/12
Logic and thought experimentswason-selection5/65/66/65/63/61/61/6
Logic and thought experimentsknights-knaves4/122/122/122/122/123/123/12
Logic and thought experimentsseating-order7/125/123/125/123/121/124/12
Logic and thought experimentsmonty-hall7/108/107/104/109/103/103/10
Logic and thought experimentsbase-rate9/109/103/102/103/102/102/10
Logic and thought experimentsdice-cards9/106/106/105/106/104/104/10
Logic and thought experimentsgamble-ev6/99/96/99/95/91/92/9
Logic and thought experimentssunk-cost5/65/64/65/65/62/62/6
Logic and thought experimentsgamblers-fallacy9/119/117/117/118/114/114/11
Logic and thought experimentsdenominator-neglect1/31/31/31/32/31/32/3
Logic and thought experimentsconjunction4/52/54/52/52/52/52/5
Logic and thought experimentssimpson2/40/41/41/42/41/41/4
Logic and thought experimentssample-size4/54/55/51/51/51/53/5
Logic and thought experimentsanchoring4/42/43/42/42/42/41/4
Logic and thought experimentsunits4/93/94/93/92/92/95/9
Logic and thought experimentscalendar-time5/93/95/94/93/93/92/9
Policy cases, 20 industriesinsurance10/1711/1715/179/178/178/174/17
Policy cases, 20 industriesretail9/137/1310/137/137/132/131/13
Policy cases, 20 industriesbanking12/1613/1610/168/1611/163/161/16
Policy cases, 20 industriesfraud12/1410/1411/1411/1411/143/145/14
Policy cases, 20 industriesairline14/1411/1413/1410/1411/142/144/14
Policy cases, 20 industrieshotel14/1611/1610/168/166/164/164/16
Policy cases, 20 industrieslogistics13/1310/1310/139/139/131/134/13
Policy cases, 20 industriessaas13/1313/1311/1310/1311/133/132/13
Policy cases, 20 industriestelecom15/1513/1511/1512/159/154/153/15
Policy cases, 20 industriesutilities8/1410/1410/1411/149/145/143/14
Policy cases, 20 industrieshr-leave13/1312/139/1312/1310/133/133/13
Policy cases, 20 industrieshr-expense10/1211/1211/1210/129/124/125/12
Policy cases, 20 industriesrealestate9/154/157/159/152/156/155/15
Policy cases, 20 industrieseducation13/158/159/158/156/155/156/15
Policy cases, 20 industriesfood10/1513/1510/1512/1511/154/155/15
Policy cases, 20 industriesmanufacturing9/126/129/125/124/123/125/12
Policy cases, 20 industrieshealth-billing10/1210/128/127/126/123/122/12
Policy cases, 20 industrieshealth-scheduling3/73/73/75/75/72/70/7
Policy cases, 20 industriespermits14/1412/1412/1411/149/142/142/14
Policy cases, 20 industriesitsec12/149/149/1411/148/146/145/14
Usability intentsclear intent28/3129/3131/3127/3130/3116/318/31
Usability intentsscreen or history context38/3933/3934/3929/3933/397/393/39
Usability intentsout of scope18/1918/1917/1917/1914/1912/190/19
Usability intentsambiguous request17/2013/207/203/206/201/208/20
Usability intentsdestructive needs confirm27/2724/2725/2723/2725/275/274/27
Usability intentsslang and typos17/1818/1817/1816/1817/185/183/18
Usability intentsmultilingual33/3429/3428/3431/3429/3413/346/34
Usability intentsaccessibility13/1716/1717/1715/1717/175/171/17
Usability intentsundo, redo and back20/2121/2120/2120/2115/219/213/21
Stress testsinjection38/4036/4031/4038/4027/4013/4018/40
Stress testsneedle22/2221/2215/2216/2215/227/228/22
Stress testsdistractor15/1615/1613/1612/169/163/164/16
Stress testsnegation17/2119/2118/2113/2116/218/216/21
Stress teststhreshold13/2513/257/257/2512/2512/2512/25
Stress testsout-of-scope19/1918/1916/1916/1915/194/199/19
Stress testsnear-options12/149/1410/1412/1410/144/144/14
Stress testsmany-options16/1614/169/1614/1611/161/160/16
Stress testsparaphrase22/2720/2720/2718/2714/2711/2711/27

Policy cases by industry: right answers per model (first presentation)

IndustryJev 1.13Clef 27BClef-Flash 9BKev 4Blev 4BLayaJulia-1
Airline rebooking14/1411/1413/1410/1411/142/144/14
Banking and card disputes12/1613/1610/168/1611/163/161/16
Education admissions and registration13/158/159/158/156/155/156/15
HR leave and expense policy23/2523/2520/2522/2519/257/258/25
Healthcare administration13/1913/1911/1912/1911/195/192/19
Hotel booking changes14/1611/1610/168/166/164/164/16
IT security access requests12/149/149/1411/148/146/145/14
Insurance claims10/1711/1715/179/178/178/174/17
Logistics and shipping exceptions13/1310/1310/139/139/131/134/13
Manufacturing quality control9/126/129/125/124/123/125/12
Payments fraud flags12/1410/1411/1411/1411/143/145/14
Public-sector permits14/1412/1412/1411/149/142/142/14
Real-estate rental applications9/154/157/159/152/156/155/15
Restaurant and food delivery10/1513/1510/1512/1511/154/155/15
Retail returns9/137/1310/137/137/132/131/13
SaaS customer support routing13/1313/1311/1310/1311/133/132/13
Telecom plan changes15/1513/1511/1512/159/154/153/15
Utilities and energy billing8/1410/1410/1411/149/145/143/14

Round robin by game: wins-draws-losses

ModelTic-tac-toeConnect FourNimDots and BoxesIllegal moves
Jev 1.1337-6-3756-0-2431-0-4959-0-210
Clef 27B38-5-3742-0-3845-0-3523-0-570
Clef-Flash 9B34-7-3938-0-4236-0-4431-0-490
lev 4B29-9-4245-0-3529-0-5134-0-460
Kev 4B33-9-3827-0-5338-0-4234-0-460
Laya20-13-4717-0-6334-0-4644-0-360
Julia-129-6-4526-0-5438-0-4223-0-570

Pong round robin: every model, speed and accuracy side by side

ModelGames W-D-LMedian decision (ms)DecisionsRight zone when the ball cameShots returned
Jev 1.1313-0-3132 ms1,436100% (523/523)86%
Kev 4B12-0-4285 ms81598% (303/310)80%
lev 4B9-0-7418 ms45199% (151/153)59%
Clef-Flash 9B8-0-8528 ms33683% (112/135)49%
Clef 27B4-0-121.79 s84100% (31/31)26%
Julia-14-0-1212 ms7,45121% (603/2811)21%
Laya3-0-1336 ms2,46012% (122/999)12%

Featured Pong match (Jev 1.13 vs Laya): who played which side

SideModelIdMedian decision time (ms)Decisions in the matchPoints
leftJev 1.13jev124 ms543
rightLayalaya37 ms1560

Featured Connect Four game: players and result

SideModelResult
AJev 1.13won
BClef 27Blost

Featured Pong rally (Jev 1.13 vs Laya): ball and paddle positions every 80 ms, 0 to 1

SecondsBall xBall yLeft paddleRight paddle
0 s0.5x0.5x0.5x0.5x
83 ms0.5x0.5x0.5x0.5x
0.2 s0.5x0.5x0.5x0.5x
0.3 s0.5x0.5x0.5x0.5x
0.3 s0.5x0.5x0.5x0.5x
0.4 s0.5x0.5x0.5x0.5x
0.5 s0.5x0.5x0.5x0.5x
0.6 s0.55x0.47x0.5x0.48x
0.7 s0.6x0.44x0.5x0.43x
0.8 s0.65x0.41x0.49x0.38x
0.8 s0.69x0.38x0.44x0.33x
0.9 s0.74x0.35x0.39x0.3x
1 s0.79x0.32x0.34x0.3x
1.1 s0.84x0.29x0.29x0.3x
1.2 s0.89x0.26x0.24x0.3x
1.3 s0.91x0.23x0.19x0.28x
1.3 s0.86x0.21x0.14x0.25x
1.4 s0.8x0.18x0.13x0.3x
1.5 s0.75x0.15x0.18x0.3x
1.6 s0.69x0.12x0.23x0.3x
1.7 s0.64x0.091x0.28x0.3x
1.8 s0.58x0.062x0.3x0.29x
1.8 s0.53x0.033x0.3x0.3x
1.9 s0.47x0.035x0.3x0.3x
2 s0.42x0.064x0.3x0.3x
2.1 s0.37x0.093x0.3x0.3x
2.2 s0.31x0.12x0.3x0.3x
2.3 s0.26x0.15x0.3x0.3x
2.3 s0.2x0.18x0.3x0.3x
2.4 s0.15x0.21x0.3x0.3x
2.5 s0.094x0.24x0.3x0.3x
2.6 s0.11x0.23x0.3x0.3x
2.7 s0.17x0.19x0.3x0.27x
2.8 s0.23x0.16x0.3x0.26x
2.8 s0.29x0.12x0.3x0.23x
2.9 s0.35x0.085x0.28x0.18x
3 s0.41x0.05x0.23x0.13x
3.1 s0.47x0.026x0.18x0.1x
3.2 s0.53x0.061x0.13x0.1x
3.3 s0.58x0.097x0.1x0.1x
3.3 s0.64x0.13x0.1x0.1x
3.4 s0.7x0.17x0.1x0.1x
3.5 s0.76x0.2x0.1x0.1x
3.6 s0.82x0.24x0.1x0.1x
3.7 s0.88x0.27x0.1x0.1x
3.8 s0.94x0.31x0.1x0.1x
3.8 s1x0.34x0.15x0.11x

Featured Connect Four game (Jev 1.13 vs Clef 27B): every move. Rule: The longest Connect Four game between jev and clef in the turn-based round robin (most moves, openings included); ties go to the earliest.

MoveRandom opening moveSideColumnProbabilityPerfect moveBest columns
1YesA4———
2YesB4———
3YesA3———
4YesB3———
5NoA40.44xNo2,5
6NoB40.19xNo5
7NoA30.39xNo2,5
8NoB30.19xNo5
9NoA30.44xNo2,5
10NoB40.18xNo2
11NoA30.58xNo2,5
12NoB40.25xNo2
13NoA50.62xYes2,5
14NoB20.22x——
15NoA50.6xNo6
16NoB60.21xYes6
17NoA50.61xYes5
18NoB50.24xYes5
19NoA50.67xNo2
20NoB50.22xYes5
21NoA60.48xYes6
22NoB60.32xYes6
23NoA60.74xYes6
24NoB60.36xNo1
25NoA60.62xYes6
26NoB20.48xNo7
27NoA20.54xYes2

Arena leagues: wins-draws-losses and Elo per model

PlayerOthelloPong, fair mode 150 msSnake, fair mode 150 msTron, fair mode 150 ms
perfect32-0-0 (1280)7-0-0 (1098)––
Clef 27B20-2-10 (1082)–11-0-5 (1068)12-3-17 (956)
random17-0-15 (1015)1-0-6 (930)6-0-10 (954)8-4-20 (898)
Laya15-1-16 (990)0-0-7 (902)1-1-14 (850)7-1-24 (852)
Julia-113-2-17 (965)2-0-5 (958)9-0-7 (1022)2-1-29 (761)
lev 4B14-0-18 (964)4-0-3 (1014)4-0-12 (908)14-3-15 (991)
Clef-Flash 9B11-1-20 (924)3-0-4 (986)9-1-6 (1034)17-1-14 (1025)
Kev 4B9-1-22 (891)6-0-1 (1070)2-1-13 (875)17-2-13 (1031)
Jev 1.139-1-22 (890)5-0-2 (1042)12-1-3 (1103)27-3-2 (1219)
expert––16-0-0 (1185)30-2-0 (1267)

Method

  1. Seven models with the same API (a JSON state in, a choice or a yes/no with probabilities out): TypeSafe Jev 1.13 over its hosted API, and Cloudflare Clef 27B, Clef-Flash 9B, lev 4B, Kev 4B, Laya and Julia-1 as GGUF files (Q4_K_M or Q8_0, checked against the published SHA-256) in llama.cpp 0.6.0 on one Mac Studio (M3 Ultra, 96 GB).
  2. Five suites, 1,085 items, written by five agents that never called a model: games solved by search (tic-tac-toe, Connect Four, Nim, Pong intercepts, Sudoku, Wordle, mazes, Hanoi, coin weighing, knight moves), logic and thought experiments (syllogisms, knights and knaves, Monty Hall, base rates, expected value, bias probes, framing pairs), fictional company policies in 20 industries, usability intents on 14 kinds of app, and stress tests (prompt injection, long logs, thresholds, near-identical options, paraphrases).
  3. Every answer comes from search, arithmetic, logic or a rule written in the item. Each suite has its own verifier that re-derives every answer. A second agent re-solved every item of the first versions of the suites (978) and found no wrong answer; the ambiguities and shortcuts it found were fixed before the freeze, adding 107 items and rewriting others. Consistency-only items (no right answer) are scored only on agreement inside their group.
  4. Exam pass: every item, every model; choice items also with shuffled options and with options renamed a, b, c. Latency pass: 120 seeded items, one model and one call at a time. Repeat pass: 60 items asked twice.
  5. Matches: a round robin of tic-tac-toe, Connect Four, Nim and Dots and Boxes (10 games per pairing, colours swapped), real-time Pong (first to 3 points after amendment 2) where each answer takes effect only after its measured latency, and a random and a perfect player as anchors. Elo: K 32, averaged over 200 seeded game orders.
  6. The protocol, the item hashes and the game counts were declared before the first counted call. Two amendments, both declared before the data they affect, are in the published protocol: the local-server batch size (amendment 1) and the Pong match length (amendment 2). One erratum: the Connect Four reference solver graded some moves in a search-order-dependent way; it was fixed and every move regraded, and no pick, probability, result or rating changed (erratum 2). Every call and every game is kept; errors count as wrong.

Caveats

  • Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.
  • The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.
  • Typed-decision models are built for routing and gating with a confidence threshold. These items have no threshold: a model that is unsure must still pick.
  • A few small families lean on content where the right option is often the longest or shortest description (sudoku, seating order, gambles). Reported, not removed.
  • These are llama.cpp ports, not the reference implementations: llama.cpp converts only part of lev's output head, an open llama.cpp issue reports degraded probabilities for Clef Q8_0 (we ran Q4_K_M), and Laya and Julia-1 were trained on inputs of 512 to 1,024 tokens while our items average about 800.
  • One harness error was fixed mid-run (amendment 1): the local servers first ran with a 512-token batch, which rejected longer inputs. Every local model was rerun from the start with a whole-input batch, one model at a time; the earlier rows are kept apart and not counted.
  • Games: 10 games per pairing in each turn-based game (colours swapped), 2 in Pong, first to 3 after amendment 2. The sign tests against the random player are not corrected for the many comparisons made, so read them as descriptions. Connect Four "perfect" is a depth-10 search plus exact endgames, not a full solve.
  • In Pong each answer moves the paddle only after its measured latency, and the arrival point of the ball is in the state, so Pong tests reading and speed more than physics. The viral "smarter AI lost at Pong" result (Laya beating Jev) did not reproduce under these rules: Jev won all 4 games of the featured series.
  • Jev ran on TypeSafe's servers and the six open models on one Mac Studio, so speed across that line is not comparable. This study compares capability (same questions and positions) directly, and speed only on one machine, on one GPU (third-party numbers), through one gateway (OpenRouter's numbers) or as list price.
  • Arena leagues (Othello, and Tron, Snake and Pong in fair mode) are content recordings, not tests of the models: one seed, 2 to 4 games per pairing, protocol amendment 3. The Mac also ran other work while they were recorded (erratum 4), so their call times are not clean measurements and no speed claim uses them; fair-mode results do not depend on call time.

Sources

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “System One arena: Jev vs Clef and five open decision models, head to head”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/system-one-arena.

Models and comparisons in this study

More studies

All benchmarks
Live storyIncludes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Jev

Jev vs Claude routers on unseen decisions: a blind holdout

Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.

82% (46/56)Jev 1.13 (TypeSafe): exact on unseen decisions · n = 56

6 chartsUpdated October 6, 2026

Live story
  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.