{"method":["Routing uses the platform’s labelled decision suites: Failure class (failure-class-2026-10-05.1, 18 cases); Message intent (frontdoor-intent-2026-10-04.2, 20 cases); Is it a rule? (memory-is-rule-2026-10-04.2, 12 cases); Context shape (context-shape-2026-10-05.1, 32 cases).","Production rules select the cases and leave the scored keys open. This gives 82 decisions and 194 scored questions.","Each decision uses one Claude Code CLI call. Calls use JSON output, tools off, no MCP and no saved session. They use the production system text, prompt and JSON schema. Calls run one at a time; temperature cannot be set.","The decision-eval runner scores each reply. Exact means every scored question is acceptable. Per-question accuracy counts each scored question. Errors and unanswered questions count as wrong.","Thinking off sets MAX_THINKING_TOKENS=0 on the claude process. Two uncounted probes check the setting, one per route. The probes and all new counted calls reported zero thinking tokens.","Thinking off uses new calls. Thinking on uses the CLI default in the recorded routing run. Sonnet 5.5 uses low effort in that recorded run. Both reference arms are reused; they were not rerun.","The exact two-sided McNemar test uses case-level pairs that differ (binomial, p = 0.5). Rate intervals are 95% Wilson intervals.","Timing uses wall time and the CLI’s reported API time. The same code computes nearest-rank p50 and p95 from each arm’s call log. Cost uses reported input, cache and output tokens times list price (a calculation).","Hard tasks use the 8 unchanged tasks and sandboxed deterministic validators from the hard head-to-head.","Controls ran before the first new model call, but before protocol creation. All 8 reference answers passed (8 tested). All 26 planted wrong answers failed. All 8 wrapped references were flagged as format misses.","Strict pass means the whole reply passes. A format miss means a lenient extractor finds a passing answer; it never counts as a strict pass. Errors and incomplete attempts count as failures.","Thinking off used 24 calls: 8 tasks, three repetitions each. Calls ran one at a time, with no effort flag, tools off, no MCP and one turn. The timeout was 300 s; the output-token cap was 16,000. Thinking on reuses 24 recorded default Haiku receipts.","Cost per strict pass divides all call costs in a cell by its strict passes. Each call cost uses list price times reported tokens (a calculation).","Recorded order: controls at 2026-10-06T21:27:56.463Z; protocol creation at 2026-10-06T21:29:01.607477Z; then the two probes. The first new counted call started at 2026-10-06T21:29:57.270Z. Reused reference calls predate both the controls and this protocol.","Declared caps: 106 new counted calls and 112 calls including probes. Observed: 106 counted calls plus 2 probes; the total stayed within both call caps.","Attempts: 106 counted calls, 2 uncounted probes. Nothing was trimmed or retried, and no run hit a usage limit. Every attempt is in the published extract."]}