{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/effort-ladder",
  "study": {
    "slug": "effort-ladder",
    "title": "Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks",
    "seoTitle": "Effort ladder: does more AI effort buy quality?",
    "description": "176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.",
    "question": "On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate for Sonnet, Opus and GPT-6.1 Sol, and what does it cost in time, tokens and list price per pass?",
    "answer": "Not on this set. Every one of the 11 configurations (Sonnet at low, medium, high and default, Opus at low, medium, high and default and GPT-6.1 Sol through Codex CLI at low, medium and high) passed all 16 calls strictly (95% interval 81% to 100% each), so pass rate does not separate any effort level. The hard set has a ceiling for these models: with 16 calls per cell it cannot rule out a difference of up to about 19 points. What more effort did change is output and time. Median total time per call by effort: Sonnet: low 5.8 s, medium 7.6 s, high 8.8 s, default 8.0 s (median output tokens 667 / 770 / 1,192 / 1,054); Opus: low 7.5 s, medium 9.7 s, high 10.1 s, default 9.2 s (median output tokens 594 / 853 / 1,052 / 945); GPT-6.1 Sol (Codex CLI): low 13.6 s, medium 13.1 s, high 18.1 s (median output tokens 284 / 335 / 436). Within each model, the fastest and slowest calls of every effort overlap, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on subscriptions), cost per strict pass went from $0.0122 at low to $0.0167 at high for Sonnet, $0.0212 at low to $0.0337 at high for Opus and $0.0128 at low to $0.0151 at high for GPT-6.1 Sol (Codex CLI).",
    "date": "2026-10-06",
    "updated": "2026-10-06",
    "tags": [
      "effort",
      "reasoning-effort",
      "hard-tasks",
      "claude-sonnet",
      "claude-opus",
      "gpt-6-1-sol",
      "claude-code",
      "codex-cli",
      "latency",
      "tokens"
    ],
    "method": [
      "Protocols declared before the first call, one per route. A follow-up to the hard head-to-head (/benchmarks/hard-model-head-to-head), on the same 8 tasks, validators, controls and CLI flags.",
      "New cells: Claude Sonnet 5.5 (low) · Claude Code; Claude Sonnet 5.5 (medium) · Claude Code; Claude Sonnet 5.5 (high) · Claude Code; Claude Opus 5.5 (low) · Claude Code; Claude Opus 5.5 (medium) · Claude Code; GPT-6.1 Sol (low) · Codex CLI. 96 calls (8 tasks × 2 repetitions per cell), one call at a time per account; order rep-major, then task, then configuration.",
      "Reference cells (5): Claude Sonnet 5.5 · Claude Code; Claude Opus 5.5 (high) · Claude Code; Claude Opus 5.5 · Claude Code; GPT-6.1 Sol (medium) · Codex CLI; GPT-6.1 Sol (high) · Codex CLI, from the hard head-to-head. Reused, not rerun. Claude reference cells keep repetitions 1-2 (n = 16) so every cell has the same design; all 24 of their calls passed in the hard study.",
      "Effort: \"--effort <level>\" for Claude Code and the thread effort for Codex CLI. \"default\" means the flag was not passed and the CLI chose the level.",
      "Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.",
      "Strict pass, format miss and wrong answer as in the hard head-to-head. Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call.",
      "Stop rules: stop at the first usage-limit or rate-limit text. No batch stopped early, nothing was trimmed or retried, and no call failed.",
      "Cost per strict pass: list price × reported tokens for every call in the cell, divided by its strict passes. A calculation."
    ],
    "caveats": [
      "Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.",
      "Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.",
      "Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.",
      "Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.",
      "List-price costs are calculations; the calls used flat subscriptions."
    ],
    "sourceIds": [
      "agent-effort-ladder",
      "calc-repricing",
      "price-anthropic",
      "price-openai"
    ],
    "stats": [
      {
        "id": "effort-ladder-pass-new",
        "label": "New effort-ladder calls that passed strictly",
        "value": 1,
        "unit": "rate",
        "display": "100% (96/96)",
        "n": 96,
        "ci": [
          0.9615,
          1
        ]
      },
      {
        "id": "effort-ladder-pass-all",
        "label": "All effort-ladder calls that passed strictly, reference cells included",
        "value": 1,
        "unit": "rate",
        "display": "100% (176/176)",
        "n": 176,
        "ci": [
          0.9786,
          1
        ]
      },
      {
        "id": "effort-ladder-cells",
        "label": "Configurations on the ladder (new + reference)",
        "value": 11,
        "unit": "count",
        "display": "11 (6 new, 5 reference)",
        "n": 176
      },
      {
        "id": "effort-ladder-format-misses",
        "label": "Format misses and wrong answers on the ladder",
        "value": 0,
        "unit": "count",
        "display": "0 format misses, 0 wrong answers",
        "n": 176
      }
    ],
    "charts": [
      {
        "id": "effort-ladder-pass-rate",
        "title": "Strict pass rate by effort on eight hard tasks",
        "subtitle": "Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals",
        "kind": "dot-range",
        "unit": "rate",
        "yLabel": "Passed",
        "series": [
          {
            "name": "Strict pass",
            "points": [
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 1,
                "lo": 0.8064,
                "hi": 1,
                "n": 16
              }
            ]
          }
        ],
        "note": "Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.",
        "whisker": "ci95",
        "sourceIds": [
          "agent-effort-ladder"
        ]
      },
      {
        "id": "effort-ladder-total-latency",
        "title": "Total time per call by effort on hard tasks",
        "subtitle": "Median per configuration; whiskers = fastest and slowest call",
        "kind": "dot-range",
        "unit": "seconds",
        "yLabel": "Seconds",
        "series": [
          {
            "name": "Total time per call by effort on hard tasks",
            "points": [
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 5.82,
                "lo": 2.78,
                "hi": 19.96,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code",
                "value": 7.63,
                "lo": 2.71,
                "hi": 24.01,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code",
                "value": 8.81,
                "lo": 2.93,
                "hi": 35.81,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 7.97,
                "lo": 2.26,
                "hi": 21.61,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code",
                "value": 7.5,
                "lo": 3.34,
                "hi": 15.82,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code",
                "value": 9.72,
                "lo": 4.78,
                "hi": 31.36,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 10.11,
                "lo": 3.63,
                "hi": 63,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 9.18,
                "lo": 4.24,
                "hi": 27.21,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 13.62,
                "lo": 7.94,
                "hi": 44.29,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 13.11,
                "lo": 8.54,
                "hi": 61.6,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 18.12,
                "lo": 11.67,
                "hi": 92.21,
                "n": 16
              }
            ]
          }
        ],
        "note": "Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-effort-ladder"
        ]
      },
      {
        "id": "effort-ladder-time-by-effort",
        "title": "Median total time per call, by effort",
        "subtitle": "One line per model and route; default = the effort flag was not passed",
        "kind": "line",
        "unit": "seconds",
        "xLabel": "Effort",
        "yLabel": "Seconds (median)",
        "series": [
          {
            "name": "Claude Sonnet 5.5 · Claude Code",
            "points": [
              {
                "label": "low",
                "value": 5.82,
                "n": 16
              },
              {
                "label": "medium",
                "value": 7.63,
                "n": 16
              },
              {
                "label": "high",
                "value": 8.81,
                "n": 16
              },
              {
                "label": "default",
                "value": 7.97,
                "n": 16
              }
            ]
          },
          {
            "name": "Claude Opus 5.5 · Claude Code",
            "points": [
              {
                "label": "low",
                "value": 7.5,
                "n": 16
              },
              {
                "label": "medium",
                "value": 9.72,
                "n": 16
              },
              {
                "label": "high",
                "value": 10.11,
                "n": 16
              },
              {
                "label": "default",
                "value": 9.18,
                "n": 16
              }
            ]
          },
          {
            "name": "GPT-6.1 Sol · Codex CLI",
            "points": [
              {
                "label": "low",
                "value": 13.62,
                "n": 16
              },
              {
                "label": "medium",
                "value": 13.11,
                "n": 16
              },
              {
                "label": "high",
                "value": 18.12,
                "n": 16
              }
            ]
          }
        ],
        "note": "Medians only; the per-call ranges are in the total-time chart and they overlap. \"default\" is placed last because its level is not known: the CLI chose it.",
        "sourceIds": [
          "agent-effort-ladder"
        ]
      },
      {
        "id": "effort-ladder-output-tokens",
        "title": "Output tokens per call by effort on hard tasks",
        "subtitle": "Median per configuration; reasoning tokens as the CLI reports them",
        "kind": "grouped-bar",
        "unit": "tokens",
        "yLabel": "Tokens",
        "series": [
          {
            "name": "Output tokens",
            "points": [
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 667,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code",
                "value": 770,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code",
                "value": 1192,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1054,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code",
                "value": 594,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code",
                "value": 853,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 1052,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 945,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 284,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 335,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 436,
                "n": 16
              }
            ]
          },
          {
            "name": "Reasoning tokens",
            "points": [
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 273,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code",
                "value": 422,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code",
                "value": 745,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 668,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code",
                "value": 87,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code",
                "value": 518,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 614,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 538,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 63,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 150,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 225,
                "n": 16
              }
            ]
          }
        ],
        "note": "Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean \"not reported\". Their content is never captured. More tokens is not better or worse by itself.",
        "sourceIds": [
          "agent-effort-ladder"
        ]
      },
      {
        "id": "effort-ladder-cost-per-pass",
        "title": "List-price cost per strict pass by effort (calculation)",
        "subtitle": "All calls in a configuration divided by its strict passes",
        "kind": "bar",
        "unit": "usd",
        "yLabel": "USD per strict pass",
        "series": [
          {
            "name": "Cost per strict pass",
            "points": [
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 0.01219,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code",
                "value": 0.01352,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code",
                "value": 0.01671,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.01398,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code",
                "value": 0.02115,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code",
                "value": 0.02947,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 0.03368,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.02893,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 0.01284,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.02564,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 0.01514,
                "n": 16
              }
            ]
          }
        ],
        "note": "Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.",
        "sourceIds": [
          "agent-effort-ladder",
          "calc-repricing",
          "price-anthropic",
          "price-openai"
        ]
      }
    ],
    "tables": [
      {
        "id": "effort-ladder-cells",
        "title": "Every effort-ladder cell",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "effort",
            "label": "Effort",
            "unit": "text"
          },
          {
            "key": "origin",
            "label": "Cell",
            "unit": "text"
          },
          {
            "key": "strict",
            "label": "Strict passes",
            "unit": "text"
          },
          {
            "key": "ci",
            "label": "95% interval",
            "unit": "text"
          },
          {
            "key": "formatMisses",
            "label": "Format misses",
            "unit": "count"
          },
          {
            "key": "wrong",
            "label": "Wrong answers",
            "unit": "count"
          },
          {
            "key": "medianTotal",
            "label": "Median total (s)",
            "unit": "seconds"
          },
          {
            "key": "rangeTotal",
            "label": "Fastest to slowest (s)",
            "unit": "text"
          },
          {
            "key": "medianOut",
            "label": "Median output tokens",
            "unit": "tokens"
          },
          {
            "key": "meanOut",
            "label": "Mean output tokens",
            "unit": "tokens"
          },
          {
            "key": "medianReasoning",
            "label": "Median reasoning tokens",
            "unit": "tokens"
          },
          {
            "key": "perPass",
            "label": "USD per strict pass (calculation)",
            "unit": "usd"
          }
        ],
        "rows": [
          {
            "config": "Claude Sonnet 5.5 (low) · Claude Code",
            "effort": "low",
            "origin": "new run",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 5.82,
            "rangeTotal": "2.8 to 20",
            "medianOut": 667,
            "meanOut": 810,
            "medianReasoning": 273,
            "perPass": 0.01219
          },
          {
            "config": "Claude Sonnet 5.5 (medium) · Claude Code",
            "effort": "medium",
            "origin": "new run",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 7.63,
            "rangeTotal": "2.7 to 24",
            "medianOut": 770,
            "meanOut": 966,
            "medianReasoning": 422,
            "perPass": 0.01352
          },
          {
            "config": "Claude Sonnet 5.5 (high) · Claude Code",
            "effort": "high",
            "origin": "new run",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 8.81,
            "rangeTotal": "2.9 to 35.8",
            "medianOut": 1192,
            "meanOut": 1284,
            "medianReasoning": 745,
            "perPass": 0.01671
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "effort": "default",
            "origin": "reference (hard head-to-head)",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 7.97,
            "rangeTotal": "2.3 to 21.6",
            "medianOut": 1054,
            "meanOut": 989,
            "medianReasoning": 668,
            "perPass": 0.01398
          },
          {
            "config": "Claude Opus 5.5 (low) · Claude Code",
            "effort": "low",
            "origin": "new run",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 7.5,
            "rangeTotal": "3.3 to 15.8",
            "medianOut": 594,
            "meanOut": 665,
            "medianReasoning": 87,
            "perPass": 0.02115
          },
          {
            "config": "Claude Opus 5.5 (medium) · Claude Code",
            "effort": "medium",
            "origin": "new run",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 9.72,
            "rangeTotal": "4.8 to 31.4",
            "medianOut": 853,
            "meanOut": 1104,
            "medianReasoning": 518,
            "perPass": 0.02947
          },
          {
            "config": "Claude Opus 5.5 (high) · Claude Code",
            "effort": "high",
            "origin": "reference (hard head-to-head)",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 10.11,
            "rangeTotal": "3.6 to 63",
            "medianOut": 1052,
            "meanOut": 1314,
            "medianReasoning": 614,
            "perPass": 0.03368
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "effort": "default",
            "origin": "reference (hard head-to-head)",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 9.18,
            "rangeTotal": "4.2 to 27.2",
            "medianOut": 945,
            "meanOut": 1053,
            "medianReasoning": 538,
            "perPass": 0.02893
          },
          {
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "effort": "low",
            "origin": "new run",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 13.62,
            "rangeTotal": "7.9 to 44.3",
            "medianOut": 284,
            "meanOut": 420,
            "medianReasoning": 63,
            "perPass": 0.01284
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "effort": "medium",
            "origin": "reference (hard head-to-head)",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 13.11,
            "rangeTotal": "8.5 to 61.6",
            "medianOut": 335,
            "meanOut": 530,
            "medianReasoning": 150,
            "perPass": 0.02564
          },
          {
            "config": "GPT-6.1 Sol (high) · Codex CLI",
            "effort": "high",
            "origin": "reference (hard head-to-head)",
            "strict": "16/16",
            "ci": "81% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "medianTotal": 18.12,
            "rangeTotal": "11.7 to 92.2",
            "medianOut": 436,
            "meanOut": 702,
            "medianReasoning": 225,
            "perPass": 0.01514
          }
        ]
      }
    ],
    "related": [
      "hard-model-head-to-head",
      "caching-consistency"
    ]
  },
  "sources": [
    {
      "id": "agent-effort-ladder",
      "title": "Effort ladder: the hard task set at each effort level",
      "kind": "run",
      "date": "2026-10-06",
      "note": "The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.",
      "data": [
        "/benchmarks/raw/effort-ladder/receipts.json",
        "/benchmarks/raw/provider-h2h-hard/receipts.json"
      ]
    },
    {
      "id": "price-anthropic",
      "title": "Anthropic list prices (Claude models)",
      "kind": "price-list",
      "date": "2026-09-21",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "note": "Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price."
    },
    {
      "id": "price-openai",
      "title": "OpenAI list prices",
      "kind": "price-list",
      "date": "2026-10-03",
      "url": "https://developers.openai.com/api/docs/pricing",
      "note": "Token prices as listed by the vendor on 2026-10-03."
    },
    {
      "id": "calc-repricing",
      "title": "Repricing calculation",
      "kind": "calculation",
      "date": "2026-10-05",
      "note": "Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes."
    }
  ]
}
