{"method":["Protocols declared before the first call, one per route. A follow-up to the hard head-to-head (/benchmarks/hard-model-head-to-head), on the same 8 tasks, validators, controls and CLI flags.","New cells: Claude Sonnet 5.5 (low) · Claude Code; Claude Sonnet 5.5 (medium) · Claude Code; Claude Sonnet 5.5 (high) · Claude Code; Claude Opus 5.5 (low) · Claude Code; Claude Opus 5.5 (medium) · Claude Code; GPT-6.1 Sol (low) · Codex CLI. 96 calls (8 tasks × 2 repetitions per cell), one call at a time per account; order rep-major, then task, then configuration.","Reference cells (5): Claude Sonnet 5.5 · Claude Code; Claude Opus 5.5 (high) · Claude Code; Claude Opus 5.5 · Claude Code; GPT-6.1 Sol (medium) · Codex CLI; GPT-6.1 Sol (high) · Codex CLI, from the hard head-to-head. Reused, not rerun. Claude reference cells keep repetitions 1-2 (n = 16) so every cell has the same design; all 24 of their calls passed in the hard study.","Effort: \"--effort <level>\" for Claude Code and the thread effort for Codex CLI. \"default\" means the flag was not passed and the CLI chose the level.","Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.","Strict pass, format miss and wrong answer as in the hard head-to-head. Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call.","Stop rules: stop at the first usage-limit or rate-limit text. No batch stopped early, nothing was trimmed or retried, and no call failed.","Cost per strict pass: list price × reported tokens for every call in the cell, divided by its strict passes. A calculation."]}