{"method":["Protocol declared before the first call. A follow-up to the five-task head-to-head (/benchmarks/model-head-to-head), where pass rate hit a ceiling.","8 hard tasks, each with a deterministic validator that runs in a sandbox without network: Fix an interval-merge function (off-by-one and edge cases); Fix a time-zone day-length function (DST); Write a CSV parser (quoted newlines, strict errors); Predict JavaScript event-loop output order; Solve a multi-constraint room schedule; Write a strict SemVer 2.0.0 regex; Refactor to remove duplication, keep 20 tests green; Write a SQLite reporting query (fan-out, ties, boundaries).","Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 references wrapped in a fence or prose are flagged as format misses.","Configurations that reached a model: Sonnet, Opus, Opus (high), GPT-6.1 Sol (medium), Fable, GPT-6.1 Sol (high), Haiku. Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task).","Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no code fence and no other text. This is stricter than the five-task study, which removed one wrapping fence.","Format miss: the strict check failed, but a lenient extractor (fenced block, outer JSON, first code line, one output line) finds an answer that passes the same validator. Reported apart from wrong answers, never as a pass.","Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call. One call at a time per account.","Attempts: Claude Code: 120 attempts, 120 reached a model, 0 blocked; Codex CLI: 62 attempts, 32 reached a model, 30 blocked. Nothing was trimmed or retried, and no run hit a usage or rate limit.","One Codex batch was resumed after its process ended (declared in the protocol before the resume): the resume skipped every task, repetition and effort the batch file already held, so no recorded call was repeated or replaced.","Cost per strict pass: list price × reported tokens for every call in the configuration (cache reads and writes priced as in the five-task study), divided by its strict passes. A calculation."]}