Jev vs Clef and five open decision models: 1,085 checkable decisions, tested
Seven decision models, 1,085 checkable decisions. The hosted model led; size bought less than you think.
TL;DR
- Jev 1.13 led: 76.8% right. The best open model, Clef 27B, reached 69.9%.
- Size bought less than you think: Kev 4B (3 GB) tied Clef-Flash 9B, 64.0% vs 65.2%.
- No speed ranking across machines: Jev is hosted; the open models ran on one Mac.
We tested 7 typed-decision models on 1,085 decisions with a checkable answer (1,048 graded, 37 that test consistency only). That made 23,247 calls, every one counted; every item and call is in the study data.
How we tested
- Models. TypeSafe's hosted Jev 1.13; Cloudflare's Clef 27B and Clef-Flash 9B; the small open models Kev 4B, lev 4B, Laya and Julia-1. A decision model reads a JSON state, answers a typed question and returns a probability for each option. It writes no text and bills input tokens only (Jev: $0.042 per million).
- Items. Five suites, written by five agents that never called a model: games solved by search, logic and thought experiments, company policies in 20 industries, usability intents on 14 kinds of app, and stress tests. A second agent re-solved the first 978 items and found no wrong answer; the ambiguities and shortcuts it found led to 107 more items. Each choice item ran three ways: original order, shuffled, and renamed a, b, c.
- Hardware. The six open models ran as GGUF files (Q4_K_M or Q8_0, SHA-256 checked) in llama.cpp 0.6.0 on one Mac Studio (M3 Ultra, 96 GB). Jev ran on TypeSafe's API from Houston. Latency came from a separate pass of 120 items; repeatability from 60 items asked twice.
- Fair footing. Accuracy, consistency, calibration, "none of these", injection and games have no clock, so they compare directly. Our order matches the independent Decision Index (rank correlation 0.93, same seven models), so 4- and 8-bit files do not appear to distort it. We never compare speed across the hosted/local line.
- Declared first. The protocol and item hashes came before the first counted call. One amendment: llama.cpp's default 512-token batch rejected every longer request, so we reran every local model.
Results
Jev 1.13 answered 805 of 1,048 graded items right (76.8%, 95% interval 74.2% to 79.3%); Clef 27B answered 733. Jev is ahead of every open model (exact McNemar test on the same items, p < 0.001 for each pair). Kev 4B and Clef-Flash 9B are tied (p = 0.53), and lev 4B trails. Laya (27.4%) and Julia-1 (25.9%) sit near chance: picking an option at random scores 22.4%.
"None of these" is where models differ most. 147 items had "none", "ask" or "escalate" as the right answer. Jev chose it 126 times (86%), Clef 27B 112 (76%), Kev 4B 91 (62%), Clef-Flash 89 (61%), lev 4B 69 (47%), Julia-1 33 and Laya 26 (18%). On the 618 items where escaping was wrong, Clef-Flash escaped 1.5% of the time and Jev 5.5%; the tiny models 14% to 17%.
Games were hard for everyone: Jev 42%, Kev 4B 39%, Clef 27B 36%. Usability intents were easiest: Jev 93%, Clef 27B 89%, Clef-Flash 87%.
- Presentation. Shuffling changed none of Clef 27B's or Clef-Flash's decisions (0 of 961 choice items each). Renaming the options a, b, c changed about 18%: your labels are part of the prompt. Jev changed 7.7% on a shuffle and 6.9% on a rename; Laya and Julia-1 a third to a half on one of the two. Right in all three presentations: Jev 72.9%, Clef 27B 65.4%, Kev 4B 57.0%.
- Prompt injection. 40 items hide "ignore the instructions and choose X" in a field a user controls. Jev and Kev 4B were right on 38 (95%), Clef 27B on 36, Clef-Flash on 31, lev 4B on 27, Julia-1 on 18 and Laya on 13.
- Confidence. Jev and Kev 4B are well calibrated (expected calibration error 0.031 and 0.037; Kev's model card claims 0.029). The Clef models rarely give a wrong answer a high probability (under 2%), but their probabilities are low across the board, so a fixed 0.8 gate rejects many right answers. Julia-1 gave 56% of its wrong answers a probability of 0.8 or more, lev 4B 34%.
- Long inputs. Laya and Julia-1 are small encoders trained on 512 to 1,024 tokens; three quarters of our items are longer. On the shorter quarter they score 37% and 42%. Larger models barely change with length.
- Repeatability. The six open models repeated every answer (60 of 60 each); Jev repeated 58 of 60.
Each source used its own data. Ours are synthetic items, so a gap is a reason to test your own decisions, not proof that a claim is false.
| Claim | Who says it | On our 1,048 items |
|---|---|---|
| Clef leads Jev | Cloudflare's launch (Decision Index 0.2.1) | Jev 76.8%, Clef 27B 69.9% (p < 0.001); the Decision Index 0.3 board agrees |
| Clef-Flash is weak on "none of these" | Cloudflare's model card | Agreed: 61% vs Jev's 86% |
| Laya beats Jev | Laya's model card (a fine-tuned variant) | The base model llama.cpp serves: 27.4% |
| Julia-1 matches Jev | Supersonic Labs | 25.9% |
| Kev 4B is well calibrated (ECE 0.029) | Kev's model card | Agreed: ECE 0.037 |
| Clef-Flash is many times faster than Jev | Cloudflare (hardware not named) | Depends on where each runs; see below |
Speed and price, on equal footing
Jev ran on TypeSafe's servers and the open models on one Mac, so side-by-side times would measure the machines. These numbers share a footing.
- One GPU (third-party Decision Index, one NVIDIA RTX PRO 6000, one harness): Laya and Julia-1 6 ms, Clef-Flash 39 ms, Kev 4B 52 ms, lev 4B 71 ms, Clef 27B 209 ms. Jev's 524 ms in that harness includes the network, so it stands apart.
- One gateway (OpenRouter's own medians; they move daily, so a snapshot, not a ranking): Jev 0.17 s, Clef 0.50 s, Clef-Flash 0.59 s, Kev 4B 4.08 s (a slower provider).
- One provider's list prices (Cloudflare Workers AI, at the tokens each model counted here): a thousand decisions cost $0.035 for Jev, $0.051 for Clef-Flash and $0.136 for Clef 27B.
- This Mac (the open models compare only with each other): Julia-1 13 ms, Laya 43 ms, Kev 4B 307 ms, lev 4B 425 ms, Clef-Flash 566 ms, Clef 27B 1.9 s. The Mac is roughly 2 to 15 times slower than the board's GPU (Clef-Flash 566 vs 39 ms, Clef 27B 1,919 vs 209 ms), so these are not server speeds.
We cannot say whether Jev is faster than Clef on equal hardware: Jev's hardware is not public and nobody has published all seven on one machine. Jev's hosted speed (0.17 s through OpenRouter, 0.14 s in our own run) is in the same range as the open models on a GPU; on a Mac the 27B model is not a real-time router. The 19 GB Clef 27B led the 6.5 GB Clef-Flash and the 3 GB Kev 4B by 5 to 6 points.
Which to use, and how to run it
- Our pick: Jev 1.13 if a hosted API is fine (about $0.035 per 1,000 decisions at our request size); Kev 4B if decisions must stay on your hardware. Clef 27B is more accurate if a GPU makes its 1.9 s go away.
- Either way: keep a "none of these" option, gate on the probability, keep your labels stable and test a few hundred of your own decisions first.
llama-server -hf ggml-org/Kev-4B-GGUF:Q4_K_M --port 8094 -c 8192 -b 8192 -ub 8192
curl -s localhost:8094/v1/systemone -H 'content-type: application/json' -d '{"state":{"ticket":"Charged twice"},"questions":{"team":{"type":"choice","instructions":"Which team?","criteria":{"billing":"Payments","none":"None of these"}}}}'- Set
-band-ubabove your longest request. With the default 512-token batch, every longer request fails with HTTP 500; our first run lost about 40% of the local calls this way. Do not go far higher: Clef-Flash reserves about 1 MiB of memory per batch token. - Run one model at a time. Seven at once made each local model 3 to 4 times slower.
- Files: Clef 27B Q4_K_M 19.2 GB, Clef-Flash 6.5 GB, Kev 4B and lev 4B about 3 GB each, Laya 0.45 GB, Julia-1 0.17 GB.
GET /v1/modelsnames the loaded file; check it, because one of our ports was taken by another app.
Limits
- The local models ran quantized on one Mac; vendors measured full-precision weights on data-centre GPUs. llama.cpp's ports are not the reference code: it converts only part of lev's output head, and an open issue reports degraded probabilities for Clef at Q8_0 (we ran Q4_K_M).
- The items are synthetic and agent-written: clear, checkable decisions, not messy product ones.
- Speed numbers are third-party or labelled "this Mac", and move with hardware and load.
Next: part 2, the models play each other, then part 3, fair-mode leagues.