{"method":["Cases: 56 new cases, 14 per decision type, in the production case shape. The types are failure class, message intent, is-it-a-rule and context shape. Production asks a router about each case. We score only questions the rule leaves open: 125 labelled questions.","Blindness: the author says they opened no router answer before the label freeze. File times support this order but cannot prove blindness. A script compared every case with existing cases. The author rewrote 17 close cases before the freeze.","File times: the protocol file dates from 21:35:45 UTC on 2026-10-06. Case drafts and one uncounted second-labeller probe came first. The first counted labelling call came at 21:35:49 UTC. The freeze came at 21:37:02 UTC. The frozen file and its checksum preceded every router call.","Second labeller: GPT-6.1 Sol through the Codex CLI, 4 counted calls. It labelled every case without the author labels.\n\nOur set accepts its label on 115 of 125 (92%, 95% interval 86% to 96%). It matches our first label on 100 of 125 (80%, 95% interval 72% to 86%). Of these questions, 24 accept two or three labels. The lowest agreement was on “artifacts”; see the caveats. Secondary scoring keeps only questions where our set accepts its label.","Routers: Jev 1.13 used direct HTTPS with the production request body. It made 168 calls in 3 repetitions, with no retries.\n\nClaude Haiku 4.5 used CLI default effort; Claude Sonnet 5.5 used effort low. Both used the Claude Code CLI with the production system text and schema. Each made 56 counted calls, one per decision. Each route had one uncounted probe.\n\nTotal calls stayed within the caps: Jev 169/200, Claude 114/120, labeller 5/5. No arm stopped early or lost cases.","Scoring: the repository’s decision-eval runner. Exact means every scored question in a case was acceptable. Key accuracy counts each question. We use Wilson 95% intervals and exact McNemar tests on paired cases. The paired tests have no multiple-test correction. Errors count as wrong.","Cost: reported tokens × list price per 1,000 decisions, a calculation."]}