{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/harder-tasks-head-to-head",
  "study": {
    "slug": "harder-tasks-head-to-head",
    "title": "GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks",
    "seoTitle": "GPT-6.1 Sol vs Claude Opus 5.5 on harder tasks",
    "description": "56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.",
    "question": "On a task set built so that Claude Sonnet 5.5 did not pass it every time, does pass rate separate GPT-6.1 Sol (Codex CLI) from Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 (Claude Code)?",
    "answer": "22 of 56 counted calls passed strictly (39%, 95% Wilson interval 28% to 52%) on 4 tasks. The Sonnet pilot did not pass these twice. GPT-6.1 Sol (medium) passed 11/16 strictly (69%, 95% interval 44% to 86%). Opus 5.5 passed 5/12 strictly (42%, 95% interval 19% to 68%). Sonnet 5.5 passed 6/16 strictly (38%, 95% interval 18% to 61%). Haiku 4.5 passed 0/12 strictly (0%, 95% interval 0% to 24%). GPT-6.1 Sol (medium) is ahead of Haiku 4.5 (the 95% intervals do not overlap). The other 5 of 6 pairs overlap, so this set cannot rank them. The lenient reading counts format misses. Opus 5.5 is also ahead of Haiku 4.5 on the lenient reading. Opus 5.5 6/12 (n = 12, 95% Wilson interval 25.4% to 74.6%); Haiku 4.5 0/12 (n = 12, 95% Wilson interval 0.0% to 24.2%). The intervals miss by 1.1 points (calculation), so this is fragile. Haiku 4.5 passed none (a floor for that configuration on this set). No configuration passed every call across the full set. Some per-task cells still hit a ceiling; see the task chart. Of 34 non-passes, 1 was a format miss with the right answer. Another 21 were wrong answers. 12 gave no answer; 10 hit the 300 s timeout. 11 of 56 calls tried a tool although tools were off. Per configuration: GPT-6.1 Sol (medium) 0 of 16, Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12. None of them passed. The Codex runner also tells the model not to call tools; the Claude Code runner does not. Median total time per completed call (wrong answers and format misses included; timeouts and tool-call parse errors excluded; ranges are not intervals): GPT-6.1 Sol (medium) 120.2 s (n = 13, range 46.2 s to 273.5 s). Opus 5.5 80.3 s (n = 9, range 3.8 s to 279.5 s). Sonnet 5.5 70.4 s (n = 12, range 4.3 s to 210.1 s). Haiku 4.5 109.0 s (n = 10, range 25.7 s to 223.9 s). Every configuration’s fastest-to-slowest range overlaps every other, so the medians describe this run and are not a tested ranking. Cost is a list-price calculation; the calls ran on subscriptions. The lowest recorded lower bound was GPT-6.1 Sol (medium) · Codex CLI: $0.083 (a lower bound: 3 timed-out calls report no tokens; $0.100 if each had cost a median call, an assumption). Selection effect: the study picked tasks with mixed or failed Sonnet pilot results. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls. The counted calls are new calls. Selection can produce this pattern; the run does not establish its cause.",
    "date": "2026-10-07",
    "updated": "2026-10-07",
    "tags": [
      "head-to-head",
      "harder-tasks",
      "gpt-6-1-sol",
      "codex-cli",
      "claude-sonnet",
      "claude-opus",
      "claude-haiku",
      "reasoning",
      "selection-effect",
      "format-misses",
      "latency"
    ],
    "method": [
      "The original protocol file predates the first probe and counted call. This review checked file birth times. Later changes are amendments. This follows the hard head-to-head (/benchmarks/hard-model-head-to-head), which hit a pass-rate ceiling for the strong models.",
      "The study started with 16 candidate tasks, written in two rounds (12 first, then 4 replacements; code, SQL, reasoning, spec, simulation, numeric). Each is one call with no tools, an exact output format and a deterministic validator that runs in a sandbox without network. Claude Sonnet 5.5 at default effort ran each candidate twice (the pilot, 32 calls). It passed 12 of the 16 candidates 2 of 2. The study dropped them.",
      "The dropped candidates include 10 of 10 code, SQL, spec, numeric and simulation tasks.",
      "The protocol declared the selection rule before the pilot. Keep at most 8 tasks: first those Sonnet passed 1 of 2, then 0 of 2. Drop tasks with 2 of 2 passes. Only 1 of the first 12 qualified, below the threshold of 6. The study wrote 4 replacement candidates once and piloted them twice; 3 qualified.",
      "The counted set is 4 tasks: 10x10 nonogram; Sudoku, 22 givens; 6x6 Skyscrapers; Seeded shuffle output. The study ran no second round of replacements.",
      "Controls ran in stages. The first 12 candidates had controls before the first probe. After a module-export validator defect, controls ran again and stored pilot replies were re-scored without new calls. Final controls covered all 16 candidates before the replacement pilot and all counted calls.\n\n16/16 references pass. 66/66 wrong answers fail. 16/16 wrapped references are format misses. Independent solvers checked unique reasoning answers. Python integers confirmed the code-reading answer.",
      "Counted cells: GPT-6.1 Sol (medium) · Codex CLI (16 calls, 4 per task). Claude Opus 5.5 · Claude Code (12 calls, 3 per task). Claude Sonnet 5.5 · Claude Code (16 calls, 4 per task). Claude Haiku 4.5 · Claude Code (12 calls, 3 per task). Rep-major, round-robin order, one call at a time per account, 300 s timeout.",
      "Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no other text. Format miss: the strict check fails, but an extracted answer passes the same validator. The extractor reads fenced blocks, answer lines or grid rows. For exact tasks it reads only the last answer-shaped candidate, not an earlier guess. Reported apart from wrong answers, never as a pass.",
      "An error or timeout counts as a non-pass. Tool use is a separate flag, not a score. It marks tool-call markup or a CLI tool-call parse error.",
      "Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn. The Codex runner adds a developer instruction to answer directly and not to call tools. The Claude Code runner adds no such instruction. Claude Code ran with an output-token cap setting of 16,000; the Codex CLI had none.",
      "Default effort means the effort flag was not passed. GPT-6.1 Sol ran at medium effort.",
      "Counted calls: Claude Code: 40 counted calls, 40 reached a model; Codex CLI: 16 counted calls, 16 reached a model. The study trimmed no calls. The study repeated no counted key. No run reported a usage or rate limit. The Claude CLI reported its own failed retry on tool-call parse errors; those receipts remain failures.\n\nThe Claude batch stopped 1 time on a tool-call parse error. A later amendment kept that failure but allowed the lane to continue. The lane repeated no counted key.",
      "Uncounted probes: 2, kept apart from 32 pilot calls and 56 counted attempts. Call caps, including probes and pilots: Claude Code: 73/105 calls; Codex CLI: 17/33 calls.",
      "Cost per strict pass is a calculation: total reported-token cost divided by strict passes. It includes failed calls with tokens and prices cache reads and writes. Timeouts report no tokens, so the cost is a lower bound."
    ],
    "caveats": [
      "Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.",
      "Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.",
      "Small samples: GPT-6.1 Sol (medium) n = 16, Opus 5.5 n = 12, Sonnet 5.5 n = 16, Haiku 4.5 n = 12. Per-task cells have 3 to 4 calls. The intervals are wide, and the calls on one task are not independent, so the intervals are likely narrower than the truth.",
      "Each row pairs a CLI with a model on one subscription. The CLIs use different subscriptions, system prompts and start-up steps. A Claude-vs-GPT row compares route + model pairs, not the models alone.",
      "Both CLIs disabled tools. Every prompt asks for the answer only. The runners differ in one way: the Codex runner adds a developer instruction not to call tools, and the Claude Code runner adds none. 11 of 40 Claude Code calls tried a tool anyway. By model: Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12.\n\nNone of them passed. None of the 16 Codex CLI calls did. We did not test the instruction, so we cannot say how much of the gap it explains. With tools on, the Claude Code models might have run code and passed on these tasks. This study does not test that either.",
      "The runner stopped 10 of 56 calls at the 300 s timeout and counted them as no answer. Counts: GPT-6.1 Sol (medium) 3, Opus 5.5 1, Sonnet 5.5 4 and Haiku 4.5 2. A longer limit might turn some into answers, passes or fails.\n\nTiming medians cover completed calls only. They include wrong answers and format misses. They omit only timeouts and tool-call parse errors in this run.",
      "Claude Code ran with an output-token cap setting of 16,000. Reported totals exceeded 16,000 in 11 of 40 Claude Code calls. The setting did not bound reported totals. We cannot tell whether it cut any reply. The tool-call parse errors came at 16,746 and 16,859 reported output tokens. The Codex CLI calls had no such cap.",
      "The kept tasks are mostly long exact-search or computation tasks. In the pilot Claude Sonnet 5.5 passed 10 of 10 code, SQL, spec, numeric and simulation candidates 2 of 2. This set says little about everyday coding.",
      "Claude Haiku 4.5 · Claude Code passed none: a floor on this set, not a general rating.",
      "Strict format rules decide part of the result: a reply with extra text fails. The chart shows the lenient reading next to the strict score.",
      "Arena servers shared the Mac during part of the run. Host load was not controlled. CLI timings include start-up and each CLI’s system prompt. These times do not isolate model speed.",
      "The protocol first said to drop tasks with validator defects. An amendment instead repaired the JavaScript validators and re-scored stored pilot replies. None of those code-writing tasks entered the counted set.",
      "No configuration passed every call across the full set. Some per-task cells hit a ceiling: Sonnet and Opus on the nonogram, and Sol on Skyscrapers. Their perfect cells do not prove equal ability.",
      "Wilson intervals treat calls as independent. The same four tasks repeat, so these are descriptive call-level intervals, not population intervals for coding tasks. The study ran no task-level paired significance test.",
      "Default effort means the effort flag was not passed; the CLI chose. Median reasoning tokens per completed call, where the CLI reports them (n and ranges are in the token chart): GPT-6.1 Sol (medium) 4,971, Opus 5.5 8,352, Sonnet 5.5 6,557, Haiku 4.5 12,483. They are part of the output tokens.",
      "List-price costs are calculations; the calls used flat subscriptions. Opus 5.5 figures are provisional: its cache-read price is under re-check. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. Those cells show a lower bound on cost per pass."
    ],
    "sourceIds": [
      "agent-harder-tasks",
      "calc-repricing",
      "price-anthropic",
      "price-openai"
    ],
    "stats": [
      {
        "id": "harder-h2h-pass-all",
        "label": "Counted calls that passed strictly (harder set)",
        "value": 0.3929,
        "unit": "rate",
        "display": "39% (22/56)",
        "n": 56,
        "ci": [
          0.2758,
          0.5237
        ]
      },
      {
        "id": "harder-h2h-correct-all",
        "label": "Counted calls with a correct answer, format misses included (lenient reading)",
        "value": 0.4107,
        "unit": "rate",
        "display": "41% (23/56)",
        "n": 56,
        "ci": [
          0.2917,
          0.5412
        ]
      },
      {
        "id": "harder-h2h-pass-sol",
        "label": "GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks",
        "value": 0.6875,
        "unit": "rate",
        "display": "69% (11/16)",
        "n": 16,
        "ci": [
          0.444,
          0.8584
        ]
      },
      {
        "id": "harder-h2h-pass-opus",
        "label": "Opus 5.5 (Claude Code): strict pass rate on the harder tasks",
        "value": 0.4167,
        "unit": "rate",
        "display": "42% (5/12)",
        "n": 12,
        "ci": [
          0.1933,
          0.6805
        ]
      },
      {
        "id": "harder-h2h-pass-sonnet",
        "label": "Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks",
        "value": 0.375,
        "unit": "rate",
        "display": "38% (6/16)",
        "n": 16,
        "ci": [
          0.1848,
          0.6136
        ]
      },
      {
        "id": "harder-h2h-pass-haiku",
        "label": "Haiku 4.5 (Claude Code): strict pass rate on the harder tasks",
        "value": 0,
        "unit": "rate",
        "display": "0% (0/12)",
        "n": 12,
        "ci": [
          0,
          0.2425
        ]
      },
      {
        "id": "harder-h2h-format-misses",
        "label": "Non-passes that were format misses, not wrong answers",
        "value": 1,
        "unit": "count",
        "display": "1 of 34 non-passes (21 wrong answers, 12 no answer)",
        "n": 34
      },
      {
        "id": "harder-h2h-tool-attempts",
        "label": "Counted calls that tried a tool although tools were off (a behaviour, not a quality score)",
        "value": 0.1964,
        "unit": "rate",
        "display": "20% (11/56)",
        "n": 56,
        "ci": [
          0.1134,
          0.3184
        ],
        "note": "Per configuration: GPT-6.1 Sol (medium) 0 of 16, Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12. None of these calls passed. The flag comes from the reply text: tool-call markup, or a CLI message that a tool call could not be parsed. The Codex runner also tells the model not to call tools; the Claude Code runner does not."
      },
      {
        "id": "harder-h2h-timeouts",
        "label": "Counted calls that ran past the 300 s limit and gave no answer",
        "value": 0.1786,
        "unit": "rate",
        "display": "18% (10/56)",
        "n": 56,
        "ci": [
          0.1,
          0.2984
        ],
        "note": "Per configuration: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12, Sonnet 5.5 4 of 16 and Haiku 4.5 2 of 12. A timeout counts as a non-pass."
      },
      {
        "id": "harder-h2h-pilot-sonnet",
        "label": "Pilot (not scored): strict passes of Claude Sonnet 5.5 on all candidate tasks",
        "value": 0.8125,
        "unit": "rate",
        "display": "81% (26/32)",
        "n": 32,
        "ci": [
          0.6469,
          0.9111
        ],
        "note": "Two calls per candidate. Its rates have a selection effect and repeated-task dependence. The pilot decided which tasks were kept; it is not part of any counted cell."
      },
      {
        "id": "harder-h2h-pilot-sonnet-kept",
        "label": "Pilot (not scored): strict passes of Claude Sonnet 5.5 on the tasks that were kept",
        "value": 0.25,
        "unit": "rate",
        "display": "25% (2/8)",
        "n": 8,
        "ci": [
          0.0715,
          0.5907
        ],
        "note": "The same four tasks the counted calls use. The counted Sonnet rate is in harder-h2h-pass-sonnet. Selection against Sonnet predicts a lower pilot rate than a fresh rate."
      },
      {
        "id": "harder-h2h-tasks-kept",
        "label": "Candidate tasks kept for the counted set",
        "value": 4,
        "unit": "count",
        "display": "4 of 16 (Sonnet passed 12 candidates 2 of 2, including 10 of 10 code, SQL, spec, numeric and simulation tasks)",
        "n": 16
      },
      {
        "id": "harder-h2h-cheapest-per-pass",
        "label": "Lowest recorded cost lower bound per strict pass (calculation)",
        "value": 0.08293,
        "unit": "usd",
        "display": "GPT-6.1 Sol (medium) · Codex CLI: $0.0829 (lower bound)",
        "n": 16,
        "note": "All passing configurations have unpriced timeout calls. Unknown costs can change their order. This is a recorded lower bound, not a cost ranking."
      }
    ],
    "charts": [
      {
        "id": "harder-h2h-pass-rate",
        "title": "Pass rate on 4 harder tasks",
        "subtitle": "Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts",
        "kind": "dot-range",
        "unit": "rate",
        "polarity": "higher",
        "yLabel": "Passed",
        "whisker": "ci95",
        "series": [
          {
            "name": "Strict pass",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.6875,
                "lo": 0.444,
                "hi": 0.8584,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.4167,
                "lo": 0.1933,
                "hi": 0.6805,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.375,
                "lo": 0.1848,
                "hi": 0.6136,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0,
                "lo": 0,
                "hi": 0.2425,
                "n": 12
              }
            ]
          },
          {
            "name": "Lenient (format misses counted)",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.6875,
                "lo": 0.444,
                "hi": 0.8584,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.5,
                "lo": 0.2538,
                "hi": 0.7462,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.375,
                "lo": 0.1848,
                "hi": 0.6136,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0,
                "lo": 0,
                "hi": 0.2425,
                "n": 12
              }
            ]
          }
        ],
        "note": "Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row.\n\nCounted calls are new calls.",
        "sourceIds": [
          "agent-harder-tasks"
        ]
      },
      {
        "id": "harder-h2h-outcomes",
        "title": "What happened on every call",
        "subtitle": "Counts per configuration: strict passes, format misses, wrong answers and calls with no answer",
        "kind": "stacked-bar",
        "unit": "count",
        "yLabel": "Calls",
        "series": [
          {
            "name": "Strict pass",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 11,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 5,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 6,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0,
                "n": 12
              }
            ]
          },
          {
            "name": "Format miss (correct answer, wrong format)",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 1,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0,
                "n": 12
              }
            ]
          },
          {
            "name": "Wrong answer",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 2,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 3,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 6,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 10,
                "n": 12
              }
            ]
          },
          {
            "name": "No answer (timeout or error)",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 3,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 3,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 4,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 2,
                "n": 12
              }
            ]
          }
        ],
        "note": "A format miss fails strictly but has an extracted answer that passes the same validator. Extra working or grids can cause this outcome. It is not a pass. A call with no answer is a timeout or an error; it counts as a non-pass.",
        "sourceIds": [
          "agent-harder-tasks"
        ]
      },
      {
        "id": "harder-h2h-tool-attempts",
        "title": "Calls that tried a tool although tools were off",
        "subtitle": "Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse",
        "kind": "dot-range",
        "unit": "rate",
        "yLabel": "Calls with a tool attempt",
        "polarity": "none",
        "whisker": "ci95",
        "series": [
          {
            "name": "Tool attempt",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0,
                "lo": 0,
                "hi": 0.1936,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.4167,
                "lo": 0.1933,
                "hi": 0.6805,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.3125,
                "lo": 0.1416,
                "hi": 0.556,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0.0833,
                "lo": 0.0149,
                "hi": 0.3539,
                "n": 12
              }
            ]
          }
        ],
        "note": "Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator.\n\nIt is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.",
        "sourceIds": [
          "agent-harder-tasks"
        ]
      },
      {
        "id": "harder-h2h-pass-by-task",
        "title": "Strict pass rate by task",
        "subtitle": "One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide",
        "kind": "grouped-bar",
        "unit": "rate",
        "polarity": "higher",
        "yLabel": "Strict pass rate",
        "whisker": "ci95",
        "series": [
          {
            "name": "GPT-6.1 Sol (medium) · Codex CLI",
            "points": [
              {
                "label": "10x10 nonogram",
                "value": 0.75,
                "lo": 0.3006,
                "hi": 0.9544,
                "n": 4
              },
              {
                "label": "Sudoku, 22 givens",
                "value": 0.25,
                "lo": 0.0456,
                "hi": 0.6994,
                "n": 4
              },
              {
                "label": "6x6 Skyscrapers",
                "value": 1,
                "lo": 0.5101,
                "hi": 1,
                "n": 4
              },
              {
                "label": "Seeded shuffle output",
                "value": 0.75,
                "lo": 0.3006,
                "hi": 0.9544,
                "n": 4
              }
            ]
          },
          {
            "name": "Claude Opus 5.5 · Claude Code",
            "points": [
              {
                "label": "10x10 nonogram",
                "value": 1,
                "lo": 0.4385,
                "hi": 1,
                "n": 3
              },
              {
                "label": "Sudoku, 22 givens",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              },
              {
                "label": "6x6 Skyscrapers",
                "value": 0.3333,
                "lo": 0.0615,
                "hi": 0.7923,
                "n": 3
              },
              {
                "label": "Seeded shuffle output",
                "value": 0.3333,
                "lo": 0.0615,
                "hi": 0.7923,
                "n": 3
              }
            ]
          },
          {
            "name": "Claude Sonnet 5.5 · Claude Code",
            "points": [
              {
                "label": "10x10 nonogram",
                "value": 1,
                "lo": 0.5101,
                "hi": 1,
                "n": 4
              },
              {
                "label": "Sudoku, 22 givens",
                "value": 0,
                "lo": 0,
                "hi": 0.4899,
                "n": 4
              },
              {
                "label": "6x6 Skyscrapers",
                "value": 0,
                "lo": 0,
                "hi": 0.4899,
                "n": 4
              },
              {
                "label": "Seeded shuffle output",
                "value": 0.5,
                "lo": 0.15,
                "hi": 0.85,
                "n": 4
              }
            ]
          },
          {
            "name": "Claude Haiku 4.5 · Claude Code",
            "points": [
              {
                "label": "10x10 nonogram",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              },
              {
                "label": "Sudoku, 22 givens",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              },
              {
                "label": "6x6 Skyscrapers",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              },
              {
                "label": "Seeded shuffle output",
                "value": 0,
                "lo": 0,
                "hi": 0.5615,
                "n": 3
              }
            ]
          }
        ],
        "note": "Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.",
        "sourceIds": [
          "agent-harder-tasks"
        ]
      },
      {
        "id": "harder-h2h-total-latency",
        "title": "Total time per call on harder tasks",
        "subtitle": "Median per configuration; whiskers = fastest and slowest call",
        "kind": "dot-range",
        "unit": "seconds",
        "yLabel": "Seconds",
        "whisker": "minmax",
        "series": [
          {
            "name": "Total time per call",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 120.24,
                "lo": 46.24,
                "hi": 273.46,
                "n": 13,
                "highlight": false
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 80.34,
                "lo": 3.82,
                "hi": 279.5,
                "n": 9,
                "highlight": false
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 70.43,
                "lo": 4.32,
                "hi": 210.08,
                "n": 12,
                "highlight": false
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 108.98,
                "lo": 25.73,
                "hi": 223.95,
                "n": 10,
                "highlight": false
              }
            ]
          }
        ],
        "note": "Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session.\n\nArena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.",
        "sourceIds": [
          "agent-harder-tasks"
        ]
      },
      {
        "id": "harder-h2h-output-tokens",
        "title": "Output tokens per call on harder tasks",
        "subtitle": "Median per configuration; reasoning tokens as the CLI reports them",
        "kind": "grouped-bar",
        "unit": "tokens",
        "yLabel": "Tokens",
        "polarity": "none",
        "whisker": "minmax",
        "series": [
          {
            "name": "Output tokens",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 4994,
                "lo": 2099,
                "hi": 13413,
                "n": 13
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 8420,
                "lo": 279,
                "hi": 40044,
                "n": 9
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 9287,
                "lo": 407,
                "hi": 27921,
                "n": 12
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 12508,
                "lo": 2965,
                "hi": 26532,
                "n": 10
              }
            ]
          },
          {
            "name": "Reasoning tokens",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 4971,
                "lo": 2070,
                "hi": 13372,
                "n": 13
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 8352,
                "lo": 21,
                "hi": 9897,
                "n": 9
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 6557,
                "lo": 63,
                "hi": 27902,
                "n": 12
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 12483,
                "lo": 2924,
                "hi": 26510,
                "n": 10
              }
            ]
          }
        ],
        "note": "Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured.\n\nClaude Code used an output-token cap setting of 16,000. Some reported totals exceeded it.\n\nCodex CLI had no cap. More tokens is not better or worse by itself.",
        "sourceIds": [
          "agent-harder-tasks"
        ]
      },
      {
        "id": "harder-h2h-cost-per-pass",
        "title": "List-price cost per strict pass on harder tasks (calculation)",
        "subtitle": "All calls in a configuration, failures and format misses included, divided by its strict passes",
        "kind": "bar",
        "unit": "usd",
        "yLabel": "USD per strict pass",
        "series": [
          {
            "name": "Cost per strict pass",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.08293,
                "n": 16,
                "highlight": true
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.23843,
                "n": 16,
                "highlight": false
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.59333,
                "n": 12,
                "highlight": false
              }
            ]
          }
        ],
        "note": "Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. These cells show a lower bound.\n\nAssume each unpriced call cost its cell’s median priced call. This sensitivity calculation gives GPT-6.1 Sol (medium) $0.100, Opus 5.5 $0.633 and Sonnet 5.5 $0.303. Opus 5.5 figures are provisional: its cache-read price is under re-check.\n\nHighlights mark the observed frontier of these lower-bound costs. Unknown timeout costs can change it; this is not a cost ranking. Claude Haiku 4.5 · Claude Code had no strict pass, so it has no cost per pass.",
        "sourceIds": [
          "agent-harder-tasks",
          "calc-repricing",
          "price-anthropic",
          "price-openai"
        ]
      },
      {
        "id": "harder-h2h-frontier",
        "title": "Observed quality vs cost frontier (calculation)",
        "subtitle": "Strict pass rate against list-price cost per strict pass",
        "kind": "scatter",
        "unit": "rate",
        "polarity": "higher",
        "xLabel": "USD per strict pass (list-price calculation)",
        "yLabel": "Strict pass rate",
        "whisker": "ci95",
        "series": [
          {
            "name": "Codex CLI",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.6875,
                "lo": 0.444,
                "hi": 0.8584,
                "n": 16,
                "x": 0.08293,
                "highlight": true
              }
            ]
          },
          {
            "name": "Claude Code",
            "points": [
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.4167,
                "lo": 0.1933,
                "hi": 0.6805,
                "n": 12,
                "x": 0.59333,
                "highlight": false
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.375,
                "lo": 0.1848,
                "hi": 0.6136,
                "n": 16,
                "x": 0.23843,
                "highlight": false
              }
            ]
          }
        ],
        "note": "Upper-left has a higher observed pass rate and lower recorded cost per pass. Highlights mark the observed frontier of lower-bound costs. Unknown timeout costs can change it. This is not a tested ranking. Frontier: GPT-6.1 Sol (medium) · Codex CLI.\n\nCosts are calculations from tokens; calls that timed out are not priced (GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16), so those cells are lower bounds. The 95% Wilson intervals are listed below; the plot shows point estimates.\n\nStrict rates: GPT-6.1 Sol (medium) 11/16 (n = 16, 95% interval 44.4% to 85.8%); Opus 5.5 5/12 (n = 12, 95% interval 19.3% to 68.0%); Sonnet 5.5 6/16 (n = 16, 95% interval 18.5% to 61.4%).",
        "sourceIds": [
          "agent-harder-tasks",
          "calc-repricing",
          "price-anthropic",
          "price-openai"
        ]
      }
    ],
    "tables": [
      {
        "id": "harder-h2h-cells",
        "title": "Cells: configuration, calls and outcomes",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "calls",
            "label": "Calls",
            "unit": "count"
          },
          {
            "key": "strict",
            "label": "Strict passes",
            "unit": "text"
          },
          {
            "key": "interval",
            "label": "95% Wilson interval",
            "unit": "text"
          },
          {
            "key": "formatMisses",
            "label": "Format misses",
            "unit": "count"
          },
          {
            "key": "wrong",
            "label": "Wrong answers",
            "unit": "count"
          },
          {
            "key": "noAnswer",
            "label": "No answer",
            "unit": "count"
          },
          {
            "key": "priced",
            "label": "Calls with a token report",
            "unit": "count"
          },
          {
            "key": "completed",
            "label": "Completed calls (timing and token n)",
            "unit": "count"
          },
          {
            "key": "medianTotal",
            "label": "Median completed-call total (s)",
            "unit": "seconds"
          },
          {
            "key": "rangeTotal",
            "label": "Completed-call time range (s, not an interval)",
            "unit": "text"
          },
          {
            "key": "p95Total",
            "label": "p95 total (s)",
            "unit": "seconds"
          },
          {
            "key": "medianOut",
            "label": "Median output tokens",
            "unit": "tokens"
          }
        ],
        "rows": [
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "calls": 16,
            "strict": "11/16",
            "interval": "44% to 86%",
            "formatMisses": 0,
            "wrong": 2,
            "noAnswer": 3,
            "priced": 13,
            "completed": 13,
            "rangeTotal": "46.24 to 273.46",
            "medianTotal": 120.24,
            "p95Total": 240.94,
            "medianOut": 4994
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "calls": 12,
            "strict": "5/12",
            "interval": "19% to 68%",
            "formatMisses": 1,
            "wrong": 3,
            "noAnswer": 3,
            "priced": 11,
            "completed": 9,
            "rangeTotal": "3.82 to 279.5",
            "medianTotal": 80.34,
            "p95Total": 208.56,
            "medianOut": 8420
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "calls": 16,
            "strict": "6/16",
            "interval": "18% to 61%",
            "formatMisses": 0,
            "wrong": 6,
            "noAnswer": 4,
            "priced": 12,
            "completed": 12,
            "rangeTotal": "4.32 to 210.08",
            "medianTotal": 70.43,
            "p95Total": 208.72,
            "medianOut": 9287
          },
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "calls": 12,
            "strict": "0/12",
            "interval": "0% to 24%",
            "formatMisses": 0,
            "wrong": 10,
            "noAnswer": 2,
            "priced": 10,
            "completed": 10,
            "rangeTotal": "25.73 to 223.95",
            "medianTotal": 108.98,
            "p95Total": 202.31,
            "medianOut": 12508
          }
        ]
      },
      {
        "id": "harder-h2h-tasks",
        "title": "Candidate tasks, pilot result and selection (prompts not shown)",
        "columns": [
          {
            "key": "task",
            "label": "Task",
            "unit": "text"
          },
          {
            "key": "kind",
            "label": "Kind",
            "unit": "text"
          },
          {
            "key": "checks",
            "label": "Validator checks",
            "unit": "count"
          },
          {
            "key": "round",
            "label": "Pilot round",
            "unit": "text"
          },
          {
            "key": "pilot",
            "label": "Sonnet 5.5 pilot (strict passes)",
            "unit": "text"
          },
          {
            "key": "selected",
            "label": "In the counted set",
            "unit": "text"
          },
          {
            "key": "controls",
            "label": "Controls (wrong answers rejected, wrapped reference flagged)",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "task": "Write a TTL and LRU cache class (random operation sequences)",
            "kind": "code",
            "checks": 13,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "5/5; yes"
          },
          {
            "task": "Write a glob matcher (braces, **, character sets, hidden files)",
            "kind": "code",
            "checks": 68,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "4/4; yes"
          },
          {
            "task": "Write a unified diff (minimal script, fixed tie-break, hunk headers)",
            "kind": "code",
            "checks": 20,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "4/4; yes"
          },
          {
            "task": "Evaluate Python-style integer expressions (precedence, chained comparisons)",
            "kind": "code",
            "checks": 64,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "5/5; yes"
          },
          {
            "task": "SQLite sessions report (gaps and islands, median, logout rule)",
            "kind": "SQL",
            "checks": 7,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "5/5; yes"
          },
          {
            "task": "SQLite as-of price and currency report (missing days, rounding)",
            "kind": "SQL",
            "checks": 14,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "5/5; yes"
          },
          {
            "task": "Solve a 6x6 Skyscrapers puzzle (14 clues, one solution)",
            "kind": "reasoning",
            "checks": 1,
            "round": "first 12 candidates",
            "pilot": "0/2 (+2 format misses)",
            "selected": "yes",
            "controls": "3/3; yes"
          },
          {
            "task": "Pick the most profitable jobs for two machines (one optimum)",
            "kind": "reasoning",
            "checks": 1,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "3/3; yes"
          },
          {
            "task": "Decode UTF-8 with one U+FFFD per maximal subpart",
            "kind": "spec",
            "checks": 35,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "4/4; yes"
          },
          {
            "task": "Next run of a cron expression in UTC (day-of-month or day-of-week rule)",
            "kind": "spec",
            "checks": 37,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "4/4; yes"
          },
          {
            "task": "Simulate a retry queue (priorities, timeouts, backoff, ties)",
            "kind": "simulation",
            "checks": 1,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "4/4; yes"
          },
          {
            "task": "Correctly rounded sum of doubles (ties, overflow, negative zero)",
            "kind": "numeric",
            "checks": 32,
            "round": "first 12 candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "4/4; yes"
          },
          {
            "task": "Solve a 10x10 nonogram (one solution)",
            "kind": "reasoning",
            "checks": 1,
            "round": "replacement candidates",
            "pilot": "1/2",
            "selected": "yes",
            "controls": "4/4; yes"
          },
          {
            "task": "Pick the most profitable jobs for three machines (30 jobs, one optimum)",
            "kind": "reasoning",
            "checks": 1,
            "round": "replacement candidates",
            "pilot": "2/2",
            "selected": "no",
            "controls": "5/5; yes"
          },
          {
            "task": "Predict the output of a seeded shuffle (32-bit integer arithmetic)",
            "kind": "code reading",
            "checks": 1,
            "round": "replacement candidates",
            "pilot": "0/2",
            "selected": "yes",
            "controls": "4/4; yes"
          },
          {
            "task": "Solve a 9x9 Sudoku with 22 givens (one solution)",
            "kind": "reasoning",
            "checks": 1,
            "round": "replacement candidates",
            "pilot": "1/2",
            "selected": "yes",
            "controls": "3/3; yes"
          }
        ]
      }
    ],
    "related": [
      "hard-model-head-to-head",
      "effort-ladder"
    ]
  },
  "sources": [
    {
      "id": "price-anthropic",
      "title": "Anthropic list prices (Claude models)",
      "kind": "price-list",
      "date": "2026-09-21",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "note": "Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price."
    },
    {
      "id": "price-openai",
      "title": "OpenAI list prices",
      "kind": "price-list",
      "date": "2026-10-03",
      "url": "https://developers.openai.com/api/docs/pricing",
      "note": "Token prices as listed by the vendor on 2026-10-03."
    },
    {
      "id": "calc-repricing",
      "title": "Repricing calculation",
      "kind": "calculation",
      "date": "2026-10-05",
      "note": "Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes."
    },
    {
      "id": "agent-harder-tasks",
      "title": "Harder tasks head-to-head",
      "kind": "run",
      "date": "2026-10-07",
      "data": [
        "/benchmarks/raw/harder-tasks/receipts.json"
      ],
      "note": "Frozen tasks selected with a Sonnet pilot. Fresh counted calls retain failures, format misses and timeouts. Replies and expected answers are not published."
    }
  ]
}
