{"method":["Protocol declared before the first counted call. The same 8 tasks, prompts, validators and strict grading as the hard head-to-head (/benchmarks/hard-model-head-to-head).","Single call: the task prompt, tools off, one turn, 300 s timeout. Agent loop: the same prompt plus one paragraph (\"You may create and run files in the current folder to test your answer. Your final message must be only the answer, in the format stated above.\"). The final message is graded exactly like a single call.","New cells: Claude Haiku 4.5 (agent loop) · Claude Code (24); Claude Sonnet 5.5 (agent loop) · Claude Code (16); GPT-6 Luna (single call) · Codex CLI (16); GPT-6 Luna (agent loop) · Codex CLI (16). Reference cells: Claude Haiku 4.5 (single call) · Claude Code (24); Claude Sonnet 5.5 (single call) · Claude Code (24), from the hard head-to-head. One call or session at a time per account; order rep-major, then task, then configuration.","Agent-loop sandbox: a fresh empty work folder per attempt, outside the temp folder. Claude Code: tools Bash, Read, Edit, Write, Glob and Grep only, the Claude Code sandbox on (writes only in the work folder, no network, no unsandboxed commands), no MCP servers, no user settings, at most 30 turns, the same 16,000-token output cap per response as the single calls. Codex CLI: exec with the workspace-write sandbox (network off), approvals never. 10-minute limit per session.","Sandbox probes before the matrix (not counted): Claude Code: network blocked (the command was refused before it ran), write to the parent folder refused, file in the work folder created; Codex CLI: network blocked, write to the parent folder refused, file in the work folder created.","Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.","Audit of every agent-loop transcript: file-edit tool calls, read tool paths and paths in shell commands. An attempt that read a file outside its work folder is contaminated: kept in the raw extract, left out of the rates. Edits outside the work folder that ran must be 0; a call the CLI refused before it ran is counted as an attempt, not an access. The reference answers were locked (no read access) while agents ran.","Stop rules: stop a route at the first usage-limit or rate-limit message; errors, time-outs and turn-limit stops count as fails. No batch stopped early and nothing was trimmed or retried.","Cost per strict pass: list price × reported tokens for every attempt in the cell, divided by its strict passes. A calculation."]}