• Jev
  • Clef
  • Decision Models
  • System One

Jev vs Clef and five open decision models: 1,085 checkable decisions, tested

Seven decision models, 1,085 checkable decisions. The hosted model led; size bought less than you think.

TL;DR

  • Jev 1.13 led: 76.8% right. The best open model, Clef 27B, reached 69.9%.
  • Size bought less than you think: Kev 4B (3 GB) tied Clef-Flash 9B, 64.0% vs 65.2%.
  • No speed ranking across machines: Jev is hosted; the open models ran on one Mac.

We tested 7 typed-decision models on 1,085 decisions with a checkable answer (1,048 graded, 37 that test consistency only). That made 23,247 calls, every one counted; every item and call is in the study data.

How we tested

  • Models. TypeSafe's hosted Jev 1.13; Cloudflare's Clef 27B and Clef-Flash 9B; the small open models Kev 4B, lev 4B, Laya and Julia-1. A decision model reads a JSON state, answers a typed question and returns a probability for each option. It writes no text and bills input tokens only (Jev: $0.042 per million).
  • Items. Five suites, written by five agents that never called a model: games solved by search, logic and thought experiments, company policies in 20 industries, usability intents on 14 kinds of app, and stress tests. A second agent re-solved the first 978 items and found no wrong answer; the ambiguities and shortcuts it found led to 107 more items. Each choice item ran three ways: original order, shuffled, and renamed a, b, c.
  • Hardware. The six open models ran as GGUF files (Q4_K_M or Q8_0, SHA-256 checked) in llama.cpp 0.6.0 on one Mac Studio (M3 Ultra, 96 GB). Jev ran on TypeSafe's API from Houston. Latency came from a separate pass of 120 items; repeatability from 60 items asked twice.
  • Fair footing. Accuracy, consistency, calibration, "none of these", injection and games have no clock, so they compare directly. Our order matches the independent Decision Index (rank correlation 0.93, same seven models), so 4- and 8-bit files do not appear to distort it. We never compare speed across the hosted/local line.
  • Declared first. The protocol and item hashes came before the first counted call. One amendment: llama.cpp's default 512-token batch rejected every longer request, so we reran every local model.

Results

Jev 1.13 answered 805 of 1,048 graded items right (76.8%, 95% interval 74.2% to 79.3%); Clef 27B answered 733. Jev is ahead of every open model (exact McNemar test on the same items, p < 0.001 for each pair). Kev 4B and Clef-Flash 9B are tied (p = 0.53), and lev 4B trails. Laya (27.4%) and Julia-1 (25.9%) sit near chance: picking an option at random scores 22.4%.

Jev 1.13
Clef 27B
Clef-Flash 9B
Kev 4B
lev 4B
Laya
Julia-1

7 rows. Highest Jev 1.13 77% (95% interval 74%–79%, n 1048). Lowest Julia-1 26% (95% interval 23%–29%, n 1048). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 1048 per row

Share of graded items answered correctly, first presentation · 95% Wilson intervals

1048 graded items in five suites. An error, a timeout or a label outside the option set counts as wrong. Two models differ clearly only where the paired McNemar test says so (table below).

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

"None of these" is where models differ most. 147 items had "none", "ask" or "escalate" as the right answer. Jev chose it 126 times (86%), Clef 27B 112 (76%), Kev 4B 91 (62%), Clef-Flash 89 (61%), lev 4B 69 (47%), Julia-1 33 and Laya 26 (18%). On the 618 items where escaping was wrong, Clef-Flash escaped 1.5% of the time and Jev 5.5%; the tiny models 14% to 17%.

  • Escaped when it should (higher is better)
  • Escaped when it should not (lower is better) (square)
Sorted by gap, largest first.
Jev 1.13
Clef 27B
Clef-Flash 9B
Kev 4B
lev 4B
Julia-1
Laya

Gap labels, Escaped when it should not (lower is better) vs Escaped when it should (higher is better): Escaped when it should not (lower is better) is x percentage points higher (+) or lower (−) than Escaped when it should (higher is better), calculated from the two values shown; lines are the 95% Wilson interval.

7 rows, 2 series: Escaped when it should (higher is better), Escaped when it should not (lower is better). Escaped when it should (higher is better): highest Jev 1.13 86% (95% interval 79%–90%, n 147). Lowest Laya 18% (95% interval 12%–25%, n 147). Not all intervals overlap. Escaped when it should not (lower is better): highest Julia-1 17% (95% interval 14%–20%, n 618). Lowest Clef-Flash 9B 1.5% (95% interval 0.8%–2.7%, n 618). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 147–618 per row

Items with a none, ask or escalate option · first presentation

Out-of-scope requests, policies that do not cover a case, ambiguous requests that need a question, lost game positions.

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Games were hard for everyone: Jev 42%, Kev 4B 39%, Clef 27B 36%. Usability intents were easiest: Jev 93%, Clef 27B 89%, Clef-Flash 87%.

Tic-tac-toe, Jev vs Clef 27B: each pick shows the probability the model gave it.
  • Presentation. Shuffling changed none of Clef 27B's or Clef-Flash's decisions (0 of 961 choice items each). Renaming the options a, b, c changed about 18%: your labels are part of the prompt. Jev changed 7.7% on a shuffle and 6.9% on a rename; Laya and Julia-1 a third to a half on one of the two. Right in all three presentations: Jev 72.9%, Clef 27B 65.4%, Kev 4B 57.0%.
  • Prompt injection. 40 items hide "ignore the instructions and choose X" in a field a user controls. Jev and Kev 4B were right on 38 (95%), Clef 27B on 36, Clef-Flash on 31, lev 4B on 27, Julia-1 on 18 and Laya on 13.
  • Confidence. Jev and Kev 4B are well calibrated (expected calibration error 0.031 and 0.037; Kev's model card claims 0.029). The Clef models rarely give a wrong answer a high probability (under 2%), but their probabilities are low across the board, so a fixed 0.8 gate rejects many right answers. Julia-1 gave 56% of its wrong answers a probability of 0.8 or more, lev 4B 34%.
  • Long inputs. Laya and Julia-1 are small encoders trained on 512 to 1,024 tokens; three quarters of our items are longer. On the shorter quarter they score 37% and 42%. Larger models barely change with length.
  • Repeatability. The six open models repeated every answer (60 of 60 each); Jev repeated 58 of 60.

Each source used its own data. Ours are synthetic items, so a gap is a reason to test your own decisions, not proof that a claim is false.

ClaimWho says itOn our 1,048 items
Clef leads JevCloudflare's launch (Decision Index 0.2.1)Jev 76.8%, Clef 27B 69.9% (p < 0.001); the Decision Index 0.3 board agrees
Clef-Flash is weak on "none of these"Cloudflare's model cardAgreed: 61% vs Jev's 86%
Laya beats JevLaya's model card (a fine-tuned variant)The base model llama.cpp serves: 27.4%
Julia-1 matches JevSupersonic Labs25.9%
Kev 4B is well calibrated (ECE 0.029)Kev's model cardAgreed: ECE 0.037
Clef-Flash is many times faster than JevCloudflare (hardware not named)Depends on where each runs; see below

Speed and price, on equal footing

Jev ran on TypeSafe's servers and the open models on one Mac, so side-by-side times would measure the machines. These numbers share a footing.

  • One GPU (third-party Decision Index, one NVIDIA RTX PRO 6000, one harness): Laya and Julia-1 6 ms, Clef-Flash 39 ms, Kev 4B 52 ms, lev 4B 71 ms, Clef 27B 209 ms. Jev's 524 ms in that harness includes the network, so it stands apart.
  • One gateway (OpenRouter's own medians; they move daily, so a snapshot, not a ranking): Jev 0.17 s, Clef 0.50 s, Clef-Flash 0.59 s, Kev 4B 4.08 s (a slower provider).
  • One provider's list prices (Cloudflare Workers AI, at the tokens each model counted here): a thousand decisions cost $0.035 for Jev, $0.051 for Clef-Flash and $0.136 for Clef 27B.
  • This Mac (the open models compare only with each other): Julia-1 13 ms, Laya 43 ms, Kev 4B 307 ms, lev 4B 425 ms, Clef-Flash 566 ms, Clef 27B 1.9 s. The Mac is roughly 2 to 15 times slower than the board's GPU (Clef-Flash 566 vs 39 ms, Clef 27B 1,919 vs 209 ms), so these are not server speeds.

We cannot say whether Jev is faster than Clef on equal hardware: Jev's hardware is not public and nobody has published all seven on one machine. Jev's hosted speed (0.17 s through OpenRouter, 0.14 s in our own run) is in the same range as the open models on a GPU; on a Mac the 27B model is not a real-time router. The 19 GB Clef 27B led the 6.5 GB Clef-Flash and the 3 GB Kev 4B by 5 to 6 points.

Which to use, and how to run it

  • Our pick: Jev 1.13 if a hosted API is fine (about $0.035 per 1,000 decisions at our request size); Kev 4B if decisions must stay on your hardware. Clef 27B is more accurate if a GPU makes its 1.9 s go away.
  • Either way: keep a "none of these" option, gate on the probability, keep your labels stable and test a few hundred of your own decisions first.
llama-server -hf ggml-org/Kev-4B-GGUF:Q4_K_M --port 8094 -c 8192 -b 8192 -ub 8192
curl -s localhost:8094/v1/systemone -H 'content-type: application/json' -d '{"state":{"ticket":"Charged twice"},"questions":{"team":{"type":"choice","instructions":"Which team?","criteria":{"billing":"Payments","none":"None of these"}}}}'
  • Set -b and -ub above your longest request. With the default 512-token batch, every longer request fails with HTTP 500; our first run lost about 40% of the local calls this way. Do not go far higher: Clef-Flash reserves about 1 MiB of memory per batch token.
  • Run one model at a time. Seven at once made each local model 3 to 4 times slower.
  • Files: Clef 27B Q4_K_M 19.2 GB, Clef-Flash 6.5 GB, Kev 4B and lev 4B about 3 GB each, Laya 0.45 GB, Julia-1 0.17 GB. GET /v1/models names the loaded file; check it, because one of our ports was taken by another app.

Limits

  • The local models ran quantized on one Mac; vendors measured full-precision weights on data-centre GPUs. llama.cpp's ports are not the reference code: it converts only part of lev's output head, and an open issue reports degraded probabilities for Clef at Q8_0 (we ran Q4_K_M).
  • The items are synthetic and agent-written: clear, checkable decisions, not messy product ones.
  • Speed numbers are third-party or labelled "this Mac", and move with hardware and load.

Next: part 2, the models play each other, then part 3, fair-mode leagues.

The data behind this post

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.