{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/hard-model-head-to-head",
  "study": {
    "slug": "hard-model-head-to-head",
    "title": "Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks",
    "seoTitle": "Hard tasks: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol",
    "description": "152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.",
    "question": "On 8 hard tasks with deterministic validators, does pass rate separate the Claude Code models and GPT-6.1 Sol through the Codex CLI, and what do speed, tokens and cost per pass add?",
    "answer": "139 of 152 calls that reached a model passed strictly (91%). 6 of 7 configurations passed every call: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Fable 5.1 · Claude Code (24/24 each, 95% interval 86% to 100%) and GPT-6.1 Sol (medium) · Codex CLI, GPT-6.1 Sol (high) · Codex CLI (16/16 each, 95% interval 81% to 100%), so the hard set still has a ceiling for these models and pass rate does not separate them. Claude Haiku 4.5 · Claude Code passed 11/24 strictly (46%, 95% interval 28% to 65%). 5 more replies had the right answer in the wrong format (for example inside a code fence), so 16/24 on a lenient reading (47% to 82%); 8 replies were wrong. It passed none of the tasks “Predict JavaScript event-loop output order”, “Solve a multi-constraint room schedule” and “Write a SQLite reporting query (fan-out, ties, boundaries)”. Median total time per call was Sonnet 7.7 s, Opus 9.2 s, Opus (high) 11.0 s, GPT-6.1 Sol (medium) 13.1 s, Fable 16.1 s, GPT-6.1 Sol (high) 18.1 s, Haiku 39.0 s. The fastest and slowest single calls of every configuration overlap with every other, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on a subscription), the lowest cost per strict pass was Claude Sonnet 5.5 · Claude Code at $0.0143; the quality-vs-cost frontier is Claude Sonnet 5.5 · Claude Code. 30 earlier attempts were blocked before any model call (Codex CLI: the CLI reported no signed-in account); they are reported, not scored, and that route ran in a later batch.",
    "date": "2026-10-06",
    "updated": "2026-10-06",
    "tags": [
      "head-to-head",
      "hard-tasks",
      "claude-haiku",
      "claude-sonnet",
      "claude-opus",
      "claude-fable",
      "gpt-6-1-sol",
      "codex-cli",
      "format-misses",
      "latency"
    ],
    "method": [
      "Protocol declared before the first call. A follow-up to the five-task head-to-head (/benchmarks/model-head-to-head), where pass rate hit a ceiling.",
      "8 hard tasks, each with a deterministic validator that runs in a sandbox without network: Fix an interval-merge function (off-by-one and edge cases); Fix a time-zone day-length function (DST); Write a CSV parser (quoted newlines, strict errors); Predict JavaScript event-loop output order; Solve a multi-constraint room schedule; Write a strict SemVer 2.0.0 regex; Refactor to remove duplication, keep 20 tests green; Write a SQLite reporting query (fan-out, ties, boundaries).",
      "Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 references wrapped in a fence or prose are flagged as format misses.",
      "Configurations that reached a model: Sonnet, Opus, Opus (high), GPT-6.1 Sol (medium), Fable, GPT-6.1 Sol (high), Haiku. Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task).",
      "Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no code fence and no other text. This is stricter than the five-task study, which removed one wrapping fence.",
      "Format miss: the strict check failed, but a lenient extractor (fenced block, outer JSON, first code line, one output line) finds an answer that passes the same validator. Reported apart from wrong answers, never as a pass.",
      "Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call. One call at a time per account.",
      "Attempts: Claude Code: 120 attempts, 120 reached a model, 0 blocked; Codex CLI: 62 attempts, 32 reached a model, 30 blocked. Nothing was trimmed or retried, and no run hit a usage or rate limit.",
      "One Codex batch was resumed after its process ended (declared in the protocol before the resume): the resume skipped every task, repetition and effort the batch file already held, so no recorded call was repeated or replaced.",
      "Cost per strict pass: list price × reported tokens for every call in the configuration (cache reads and writes priced as in the five-task study), divided by its strict passes. A calculation."
    ],
    "caveats": [
      "6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.",
      "Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.",
      "Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.",
      "The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.",
      "Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.",
      "CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.",
      "Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.",
      "List-price costs are calculations; the calls used a flat subscription."
    ],
    "sourceIds": [
      "agent-provider-h2h-hard",
      "calc-repricing",
      "price-anthropic",
      "price-openai"
    ],
    "stats": [
      {
        "id": "hard-h2h-pass-all",
        "label": "Calls that passed strictly (hard set)",
        "value": 0.9145,
        "unit": "rate",
        "display": "91% (139/152)",
        "n": 152,
        "ci": [
          0.8592,
          0.9493
        ]
      },
      {
        "id": "hard-h2h-correct-all",
        "label": "Calls with a correct answer, format misses included (lenient reading)",
        "value": 0.9474,
        "unit": "rate",
        "display": "95% (144/152)",
        "n": 152,
        "ci": [
          0.8996,
          0.9731
        ]
      },
      {
        "id": "hard-h2h-format-misses",
        "label": "Non-passes that were format misses, not wrong answers",
        "value": 5,
        "unit": "count",
        "display": "5 of 13 non-passes (8 wrong answers)",
        "n": 13
      },
      {
        "id": "hard-h2h-perfect-configs",
        "label": "Configurations that passed every call",
        "value": 6,
        "unit": "count",
        "display": "6 of 7 (4 at 24/24, 2 at 16/16)",
        "n": 7
      },
      {
        "id": "hard-h2h-fastest-perfect",
        "label": "Lowest observed median time among configurations that passed every call (separate batches)",
        "value": 7.75,
        "unit": "seconds",
        "display": "Claude Sonnet 5.5 · Claude Code: 7.7 s",
        "n": 24,
        "note": "The counted Claude and Codex batches ran hours apart on one host and network. Host load was not controlled; this does not isolate model speed."
      },
      {
        "id": "hard-h2h-cheapest-per-pass",
        "label": "Lowest list-price cost per strict pass (calculation)",
        "value": 0.01435,
        "unit": "usd",
        "display": "Claude Sonnet 5.5 · Claude Code: $0.0143",
        "n": 24
      },
      {
        "id": "hard-h2h-blocked",
        "label": "Attempts blocked before any model call (not scored)",
        "value": 30,
        "unit": "count",
        "display": "30 (Codex CLI; 0 model calls)",
        "n": 182
      }
    ],
    "charts": [
      {
        "id": "hard-h2h-pass-rate",
        "title": "Pass rate on eight hard tasks",
        "subtitle": "Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts",
        "kind": "dot-range",
        "unit": "rate",
        "yLabel": "Passed",
        "series": [
          {
            "name": "Strict pass",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "lo": 0.862,
                "hi": 1,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 1,
                "lo": 0.862,
                "hi": 1,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 1,
                "lo": 0.862,
                "hi": 1,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 1,
                "lo": 0.862,
                "hi": 1,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0.4583,
                "lo": 0.2789,
                "hi": 0.6493,
                "n": 24
              }
            ]
          },
          {
            "name": "Lenient (format misses counted)",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "lo": 0.862,
                "hi": 1,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 1,
                "lo": 0.862,
                "hi": 1,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 1,
                "lo": 0.862,
                "hi": 1,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 1,
                "lo": 0.862,
                "hi": 1,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0.6667,
                "lo": 0.4671,
                "hi": 0.8203,
                "n": 24
              }
            ]
          }
        ],
        "note": "Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.",
        "sourceIds": [
          "agent-provider-h2h-hard"
        ]
      },
      {
        "id": "hard-h2h-outcomes",
        "title": "What happened on every call",
        "subtitle": "Counts per configuration: strict passes, format misses and wrong answers",
        "kind": "stacked-bar",
        "unit": "count",
        "yLabel": "Calls",
        "series": [
          {
            "name": "Strict pass",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 24,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 24,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 24,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 16,
                "n": 16
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 24,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 16,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 11,
                "n": 24
              }
            ]
          },
          {
            "name": "Format miss (correct answer, wrong format)",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0,
                "n": 16
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 0,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 5,
                "n": 24
              }
            ]
          },
          {
            "name": "Wrong answer",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0,
                "n": 16
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 0,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 8,
                "n": 24
              }
            ]
          }
        ],
        "note": "A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.",
        "sourceIds": [
          "agent-provider-h2h-hard"
        ]
      },
      {
        "id": "hard-h2h-total-latency",
        "title": "Total time per call on hard tasks (separate batches)",
        "subtitle": "Median per configuration; whiskers = fastest and slowest call",
        "kind": "dot-range",
        "unit": "seconds",
        "yLabel": "Seconds",
        "series": [
          {
            "name": "Total time per call on hard tasks (separate batches)",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 7.75,
                "lo": 2.26,
                "hi": 34.79,
                "n": 24,
                "highlight": true
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 9.18,
                "lo": 4.24,
                "hi": 27.21,
                "n": 24,
                "highlight": true
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 11.03,
                "lo": 3.63,
                "hi": 63,
                "n": 24,
                "highlight": true
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 13.11,
                "lo": 8.54,
                "hi": 61.6,
                "n": 16,
                "highlight": true
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 16.13,
                "lo": 4.46,
                "hi": 90,
                "n": 24,
                "highlight": true
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 18.12,
                "lo": 11.67,
                "hi": 92.21,
                "n": 16,
                "highlight": true
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 39.01,
                "lo": 15.27,
                "hi": 75.13,
                "n": 24,
                "highlight": false
              }
            ]
          }
        ],
        "note": "One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.",
        "sourceIds": [
          "agent-provider-h2h-hard"
        ]
      },
      {
        "id": "hard-h2h-first-useful-latency",
        "title": "Time to first useful output on hard tasks",
        "subtitle": "Median per configuration; whiskers = fastest and slowest call",
        "kind": "dot-range",
        "unit": "seconds",
        "yLabel": "Seconds",
        "series": [
          {
            "name": "Time to first useful output on hard tasks",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 5.95,
                "lo": 0.86,
                "hi": 30.57,
                "n": 24,
                "highlight": true
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 6.78,
                "lo": 2.39,
                "hi": 21.77,
                "n": 24,
                "highlight": true
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 7.13,
                "lo": 2.15,
                "hi": 56.23,
                "n": 24,
                "highlight": true
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 10.23,
                "lo": 6.09,
                "hi": 40.41,
                "n": 16,
                "highlight": true
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 11.63,
                "lo": 2,
                "hi": 85.33,
                "n": 24,
                "highlight": true
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 12.69,
                "lo": 8.93,
                "hi": 75.91,
                "n": 16,
                "highlight": true
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 35.54,
                "lo": 12.88,
                "hi": 70.31,
                "n": 24,
                "highlight": false
              }
            ]
          }
        ],
        "note": "One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.",
        "sourceIds": [
          "agent-provider-h2h-hard"
        ]
      },
      {
        "id": "hard-h2h-output-tokens",
        "title": "Output tokens per call on hard tasks",
        "subtitle": "Median per configuration; reasoning tokens as the CLI reports them",
        "kind": "grouped-bar",
        "unit": "tokens",
        "yLabel": "Tokens",
        "series": [
          {
            "name": "Output tokens",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1050,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 945,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 1052,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 335,
                "n": 16
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 1366,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 436,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 5064,
                "n": 24
              }
            ]
          },
          {
            "name": "Reasoning tokens",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 585,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 529,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 614,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 150,
                "n": 16
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 889,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 225,
                "n": 16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 4556,
                "n": 24
              }
            ]
          }
        ],
        "note": "Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.",
        "sourceIds": [
          "agent-provider-h2h-hard"
        ]
      },
      {
        "id": "hard-h2h-cost-per-pass",
        "title": "List-price cost per strict pass on hard tasks (calculation)",
        "subtitle": "All calls in a configuration, failures and format misses included, divided by its strict passes",
        "kind": "bar",
        "unit": "usd",
        "yLabel": "USD per strict pass",
        "series": [
          {
            "name": "Cost per strict pass",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.01435,
                "n": 24,
                "highlight": true
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 0.01514,
                "n": 16,
                "highlight": false
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.02564,
                "n": 16,
                "highlight": false
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.02824,
                "n": 24,
                "highlight": false
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 0.03337,
                "n": 24,
                "highlight": false
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0.0672,
                "n": 24,
                "highlight": false
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 0.09331,
                "n": 24,
                "highlight": false
              }
            ]
          }
        ],
        "note": "Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.",
        "sourceIds": [
          "agent-provider-h2h-hard",
          "calc-repricing",
          "price-anthropic",
          "price-openai"
        ]
      },
      {
        "id": "hard-h2h-frontier",
        "title": "Quality vs cost frontier on hard tasks",
        "subtitle": "Strict pass rate against list-price cost per strict pass",
        "kind": "scatter",
        "unit": "rate",
        "xLabel": "USD per strict pass (list-price calculation)",
        "yLabel": "Strict pass rate",
        "series": [
          {
            "name": "Claude Code",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "x": 0.01435,
                "value": 1,
                "n": 24,
                "highlight": true
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "x": 0.02824,
                "value": 1,
                "n": 24,
                "highlight": false
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "x": 0.03337,
                "value": 1,
                "n": 24,
                "highlight": false
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "x": 0.09331,
                "value": 1,
                "n": 24,
                "highlight": false
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "x": 0.0672,
                "value": 0.4583,
                "n": 24,
                "highlight": false
              }
            ]
          },
          {
            "name": "Codex CLI",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "x": 0.02564,
                "value": 1,
                "n": 16,
                "highlight": false
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "x": 0.01514,
                "value": 1,
                "n": 16,
                "highlight": false
              }
            ]
          }
        ],
        "note": "Upper-left is better. Highlighted points are on the frontier: no other configuration passes at least as often for at most the same cost per pass. Frontier: Claude Sonnet 5.5 · Claude Code. Costs are calculations from tokens. Pass rates with their 95% intervals are in the pass-rate chart.",
        "sourceIds": [
          "agent-provider-h2h-hard",
          "calc-repricing",
          "price-anthropic",
          "price-openai"
        ]
      }
    ],
    "tables": [
      {
        "id": "hard-h2h-pass-matrix",
        "title": "Pass matrix on hard tasks: configuration × task (strict passes)",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "c0",
            "label": "Fix an interval-merge function (off-by-one and edge cases)",
            "unit": "text"
          },
          {
            "key": "c1",
            "label": "Fix a time-zone day-length function (DST)",
            "unit": "text"
          },
          {
            "key": "c2",
            "label": "Write a CSV parser (quoted newlines, strict errors)",
            "unit": "text"
          },
          {
            "key": "c3",
            "label": "Predict JavaScript event-loop output order",
            "unit": "text"
          },
          {
            "key": "c4",
            "label": "Solve a multi-constraint room schedule",
            "unit": "text"
          },
          {
            "key": "c5",
            "label": "Write a strict SemVer 2.0.0 regex",
            "unit": "text"
          },
          {
            "key": "c6",
            "label": "Refactor to remove duplication, keep 20 tests green",
            "unit": "text"
          },
          {
            "key": "c7",
            "label": "Write a SQLite reporting query (fan-out, ties, boundaries)",
            "unit": "text"
          },
          {
            "key": "strict",
            "label": "Strict total",
            "unit": "text"
          },
          {
            "key": "formatMisses",
            "label": "Format misses",
            "unit": "count"
          },
          {
            "key": "wrong",
            "label": "Wrong answers",
            "unit": "count"
          },
          {
            "key": "medianTotal",
            "label": "Median total (s)",
            "unit": "seconds"
          }
        ],
        "rows": [
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "strict": "24/24",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 7.75,
            "c0": "3/3",
            "c1": "3/3",
            "c2": "3/3",
            "c3": "3/3",
            "c4": "3/3",
            "c5": "3/3",
            "c6": "3/3",
            "c7": "3/3"
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "strict": "24/24",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 9.18,
            "c0": "3/3",
            "c1": "3/3",
            "c2": "3/3",
            "c3": "3/3",
            "c4": "3/3",
            "c5": "3/3",
            "c6": "3/3",
            "c7": "3/3"
          },
          {
            "config": "Claude Opus 5.5 (high) · Claude Code",
            "strict": "24/24",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 11.03,
            "c0": "3/3",
            "c1": "3/3",
            "c2": "3/3",
            "c3": "3/3",
            "c4": "3/3",
            "c5": "3/3",
            "c6": "3/3",
            "c7": "3/3"
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "strict": "16/16",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 13.11,
            "c0": "2/2",
            "c1": "2/2",
            "c2": "2/2",
            "c3": "2/2",
            "c4": "2/2",
            "c5": "2/2",
            "c6": "2/2",
            "c7": "2/2"
          },
          {
            "config": "Claude Fable 5.1 · Claude Code",
            "strict": "24/24",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 16.13,
            "c0": "3/3",
            "c1": "3/3",
            "c2": "3/3",
            "c3": "3/3",
            "c4": "3/3",
            "c5": "3/3",
            "c6": "3/3",
            "c7": "3/3"
          },
          {
            "config": "GPT-6.1 Sol (high) · Codex CLI",
            "strict": "16/16",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 18.12,
            "c0": "2/2",
            "c1": "2/2",
            "c2": "2/2",
            "c3": "2/2",
            "c4": "2/2",
            "c5": "2/2",
            "c6": "2/2",
            "c7": "2/2"
          },
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "strict": "11/24",
            "formatMisses": 5,
            "wrong": 8,
            "medianTotal": 39.01,
            "c0": "3/3",
            "c1": "1/3",
            "c2": "2/3 (+1 format miss)",
            "c3": "0/3",
            "c4": "0/3 (+2 format misses)",
            "c5": "3/3",
            "c6": "2/3",
            "c7": "0/3 (+2 format misses)"
          }
        ]
      },
      {
        "id": "hard-h2h-controls",
        "title": "Validator controls run before the first model call",
        "columns": [
          {
            "key": "task",
            "label": "Task",
            "unit": "text"
          },
          {
            "key": "reference",
            "label": "Reference answer",
            "unit": "text"
          },
          {
            "key": "checks",
            "label": "Checks",
            "unit": "count"
          },
          {
            "key": "wrong",
            "label": "Plausible wrong answers rejected",
            "unit": "text"
          },
          {
            "key": "wrapped",
            "label": "Wrapped reference flagged as format miss",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "task": "Fix an interval-merge function (off-by-one and edge cases)",
            "reference": "passes",
            "checks": 18,
            "wrong": "5/5",
            "wrapped": "yes"
          },
          {
            "task": "Fix a time-zone day-length function (DST)",
            "reference": "passes",
            "checks": 36,
            "wrong": "3/3",
            "wrapped": "yes"
          },
          {
            "task": "Write a CSV parser (quoted newlines, strict errors)",
            "reference": "passes",
            "checks": 22,
            "wrong": "3/3",
            "wrapped": "yes"
          },
          {
            "task": "Predict JavaScript event-loop output order",
            "reference": "passes",
            "checks": 1,
            "wrong": "3/3",
            "wrapped": "yes"
          },
          {
            "task": "Solve a multi-constraint room schedule",
            "reference": "passes",
            "checks": 14,
            "wrong": "3/3",
            "wrapped": "yes"
          },
          {
            "task": "Write a strict SemVer 2.0.0 regex",
            "reference": "passes",
            "checks": 30,
            "wrong": "3/3",
            "wrapped": "yes"
          },
          {
            "task": "Refactor to remove duplication, keep 20 tests green",
            "reference": "passes",
            "checks": 25,
            "wrong": "2/2",
            "wrapped": "yes"
          },
          {
            "task": "Write a SQLite reporting query (fan-out, ties, boundaries)",
            "reference": "passes",
            "checks": 8,
            "wrong": "4/4",
            "wrapped": "yes"
          }
        ]
      }
    ],
    "related": [
      "model-head-to-head"
    ]
  },
  "sources": [
    {
      "id": "agent-provider-h2h-hard",
      "title": "Provider head-to-head, hard set: eight hard tasks with strict validators",
      "kind": "run",
      "date": "2026-10-06",
      "note": "Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.",
      "data": [
        "/benchmarks/raw/provider-h2h-hard/receipts.json"
      ]
    },
    {
      "id": "price-anthropic",
      "title": "Anthropic list prices (Claude models)",
      "kind": "price-list",
      "date": "2026-09-21",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "note": "Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price."
    },
    {
      "id": "price-openai",
      "title": "OpenAI list prices",
      "kind": "price-list",
      "date": "2026-10-03",
      "url": "https://developers.openai.com/api/docs/pricing",
      "note": "Token prices as listed by the vendor on 2026-10-03."
    },
    {
      "id": "calc-repricing",
      "title": "Repricing calculation",
      "kind": "calculation",
      "date": "2026-10-05",
      "note": "Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes."
    }
  ]
}
