{"method":["Cases: the labelled decision suites the platform uses (failure class, message intent, is-it-a-rule, context shape), only the cases production asks a router about.","Scoring: the repository’s own decision-eval runner. Exact means every scored question in a case was acceptable; key accuracy counts each question.","Jev numbers come from a live run on 2026-10-06: 3 repeats of the same 82 decisions (246 counted calls), sent one at a time over HTTPS to the TypeSafe API from one Apple M3 Ultra Mac on a home network. A failed call would count as wrong; there were 0. Latency is client wall time, so the network is inside it, and the API reports no server time. Jev’s earlier recorded production run (2026-10-05, same case versions and runner) scored 74 of 82 and has no per-call latency. Cost is a calculation: reported input tokens × the published price.","Claude routers ran through the Claude Code CLI with the production system text and schema, one call per decision, one pass over the 82 decisions.","Stability: 73 of 82 decisions were exact in every repeat, 8 in none and 1 in some (a Context shape case, wrong in repeat 2). Jev can return different probabilities for the same request; the repeats show how much.","Economics: recorded tokens of the benchmark runs repriced at list prices for each model mix (a calculation)."]}