{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/coding-agents-head-to-head",
  "study": {
    "slug": "coding-agents-head-to-head",
    "title": "Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks",
    "seoTitle": "Claude Code vs Codex CLI: 6 coding tasks, hidden tests",
    "description": "36 graded sessions: Claude Code with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol. All passed every hidden test; time, tool calls and diffs differ.",
    "question": "On small real repository tasks graded by hidden tests, how do coding-agent CLIs compare when they run with their normal file and shell tools?",
    "answer": "All 3 agents passed every hidden check in every session (12/12, 12/12, 12/12; 95% Wilson 76–100% each), so this task set cannot separate them on quality. Sonnet 5.5 in Claude Code was fastest (median 23.1 s; its sessions took 18.7 s to 44.5 s), Opus 5.5 in Claude Code took a median 56.9 s, GPT-6.1 Sol in Codex CLI took a median 113.4 s; the run ranges of Sonnet 5.5 in Claude Code and GPT-6.1 Sol in Codex CLI do not overlap. Codex made more tool calls (median 12.5 vs 7.5 and 7.5) and larger diffs (median 94 lines vs 45 and 77.5, mostly added tests), and every Codex session also followed the tester’s global AGENTS.md (10 of 12 wrote a work log nobody asked for), so its time and diff include extra work. Opus used 1.9× the output tokens of Sonnet in the same CLI (medians). Gemini CLI was not run: it needed a browser login.",
    "date": "2026-10-06",
    "updated": "2026-10-06",
    "tags": [
      "claude-code",
      "codex-cli",
      "coding-agents",
      "hidden-tests",
      "head-to-head"
    ],
    "method": [
      "Six small Node.js repositories: a pagination bug across two files, a CLI flag with exact error text, a refactor that must keep every quirk, an LRU cache to a written spec, a queue race and strict TypeScript types. Each has 1 to 4 visible tests; the hidden checks live outside the repository and run only after the session ends.",
      "Controls before the first session (Node v25.2.1): every base repository fails its hidden checks and every reference patch passes all of them.",
      "Claude Code 2.1.286 headless with the sonnet and opus models at the default effort; tools Bash, Read, Edit, Write, Glob, Grep; no web tools, MCP servers or sub-agents; the OS sandbox on, writes limited to the repository, no network; setting sources off. Codex CLI exec with GPT-6.1 Sol at medium effort, workspace-write sandbox, no network, user config ignored (its version was not recorded).",
      "6 tasks × 2 repetitions × 3 agents = 36 sessions with a 10-minute timeout; nothing retried. The Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts.",
      "Pass = every hidden check passes; no partial credit. Recorded per session: wall time, tool calls, tokens as reported, files touched, lines changed against the base commit, commits and edits outside the repository.",
      "Gemini CLI 0.16.0 was probed first: it asked for a browser login, so it was not run and is not counted. The protocol was declared after the controls and before the first session."
    ],
    "caveats": [
      "Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.",
      "Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.",
      "Each run pairs a CLI with a model (Claude Code with Claude models, Codex CLI with GPT-6.1 Sol), so the results cannot separate the CLI from the model.",
      "n = 12 sessions per agent (2 per task). Time ranges are the fastest and slowest sessions, not confidence intervals; the two lanes shared one machine.",
      "Tokens are as each CLI reports them, with different tokenizers and context handling: compare tokens inside Claude Code (Sonnet vs Opus), not across vendors. Costs are list-price calculations on subscription sessions, not invoices."
    ],
    "sourceIds": [
      "agent-coding-agents",
      "calc-repricing",
      "price-anthropic",
      "price-openai"
    ],
    "stats": [
      {
        "id": "coding-agents-pass-all",
        "label": "Sessions that passed every hidden check, all three agents",
        "value": 1,
        "unit": "rate",
        "display": "100% (36/36)",
        "n": 36,
        "ci": [
          0.9036,
          1
        ],
        "note": "The ceiling: 36 of 36 means this task set cannot rank the agents on quality."
      },
      {
        "id": "coding-agents-sessions",
        "label": "Graded sessions (6 tasks × 2 repetitions × 3 agents)",
        "value": 36,
        "unit": "count",
        "display": "36",
        "n": 36,
        "note": "0 timeouts, 0 errors or usage-limit stops, nothing retried. Gemini CLI not run: it asked for a browser login."
      },
      {
        "id": "coding-agents-median-time-fastest",
        "label": "Median time per session, Sonnet 5.5 in Claude Code",
        "value": 23.1,
        "unit": "seconds",
        "display": "23.1 s",
        "n": 12,
        "note": "Fastest 18.7 s, slowest 44.5 s (a range of 12 sessions, not an interval)."
      },
      {
        "id": "coding-agents-time-ratio",
        "label": "Median time, GPT-6.1 Sol in Codex CLI vs Sonnet 5.5 in Claude Code (ratio of medians)",
        "value": 4.92,
        "unit": "ratio",
        "display": "4.9×",
        "n": 12,
        "note": "113.4 s vs 23.1 s. The run ranges do not overlap. Codex time includes the work its standing instructions asked for (see the caveats)."
      },
      {
        "id": "coding-agents-tester-notes",
        "label": "Codex sessions that read the tester’s global notes (standing instructions)",
        "value": 12,
        "unit": "count",
        "display": "12 of 12",
        "n": 12,
        "note": "10 of 12 also wrote a WORKLOG.md and 9 reported a commit attempt; no task asked for either. No Claude Code session did any of this."
      },
      {
        "id": "coding-agents-outside-edits",
        "label": "Edits that landed outside the task repository",
        "value": 0,
        "unit": "count",
        "display": "0",
        "n": 36,
        "note": "3 attempted writes to a temp folder (1 Sonnet 5.5, 2 Opus 5.5) were refused by the Claude Code permission check."
      },
      {
        "id": "coding-agents-cost-total",
        "label": "List-price estimate of all 36 sessions (calculation, not an invoice)",
        "value": 4.87,
        "unit": "usd",
        "display": "$4.87",
        "n": 36
      }
    ],
    "charts": [
      {
        "id": "coding-agents-pass-rate",
        "title": "Coding sessions that passed every hidden check",
        "subtitle": "A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals",
        "kind": "dot-range",
        "unit": "rate",
        "whisker": "ci95",
        "yLabel": "Passed",
        "viz": "IntervalDotPlot",
        "series": [
          {
            "name": "Passed every hidden check",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              }
            ]
          }
        ],
        "note": "6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.",
        "sourceIds": [
          "agent-coding-agents"
        ]
      },
      {
        "id": "coding-agents-wall-time",
        "title": "Time per coding session",
        "subtitle": "Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)",
        "kind": "dot-range",
        "unit": "seconds",
        "whisker": "minmax",
        "yLabel": "Seconds",
        "viz": "LatencyLanes",
        "series": [
          {
            "name": "Wall time per session",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 23.1,
                "lo": 18.7,
                "hi": 44.5,
                "n": 12,
                "highlight": true
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 56.9,
                "lo": 29.8,
                "hi": 185.8,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 113.4,
                "lo": 78.5,
                "hi": 221.9,
                "n": 12
              }
            ]
          }
        ],
        "note": "CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.",
        "sourceIds": [
          "agent-coding-agents"
        ]
      },
      {
        "id": "coding-agents-time-by-task",
        "title": "Time per coding task",
        "subtitle": "Median of 2 sessions per task and agent, seconds",
        "kind": "grouped-bar",
        "unit": "seconds",
        "xLabel": "Task",
        "yLabel": "Seconds",
        "viz": "SmallMultiplesBars",
        "series": [
          {
            "name": "Claude Sonnet 5.5 · Claude Code",
            "points": [
              {
                "label": "Pagination fix",
                "value": 19.1,
                "n": 2
              },
              {
                "label": "CLI --top flag",
                "value": 43.6,
                "n": 2
              },
              {
                "label": "Invoice refactor",
                "value": 24.1,
                "n": 2
              },
              {
                "label": "LRU cache",
                "value": 23.6,
                "n": 2
              },
              {
                "label": "Queue race",
                "value": 22.7,
                "n": 2
              },
              {
                "label": "Strict TypeScript types",
                "value": 27.6,
                "n": 2
              }
            ]
          },
          {
            "name": "Claude Opus 5.5 · Claude Code",
            "points": [
              {
                "label": "Pagination fix",
                "value": 39.3,
                "n": 2
              },
              {
                "label": "CLI --top flag",
                "value": 55.3,
                "n": 2
              },
              {
                "label": "Invoice refactor",
                "value": 58.4,
                "n": 2
              },
              {
                "label": "LRU cache",
                "value": 62.4,
                "n": 2
              },
              {
                "label": "Queue race",
                "value": 118.6,
                "n": 2
              },
              {
                "label": "Strict TypeScript types",
                "value": 58,
                "n": 2
              }
            ]
          },
          {
            "name": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
            "points": [
              {
                "label": "Pagination fix",
                "value": 82.7,
                "n": 2
              },
              {
                "label": "CLI --top flag",
                "value": 101.2,
                "n": 2
              },
              {
                "label": "Invoice refactor",
                "value": 129.1,
                "n": 2
              },
              {
                "label": "LRU cache",
                "value": 203.1,
                "n": 2
              },
              {
                "label": "Queue race",
                "value": 98.6,
                "n": 2
              },
              {
                "label": "Strict TypeScript types",
                "value": 143.6,
                "n": 2
              }
            ]
          }
        ],
        "note": "2 sessions per cell, so one slow session moves a bar: Opus 5.5 in Claude Code took 51.5 s and 185.8 s on Queue race.",
        "sourceIds": [
          "agent-coding-agents"
        ]
      },
      {
        "id": "coding-agents-tool-calls",
        "title": "Tool calls per coding session",
        "subtitle": "Median; whiskers = fewest and most of 12 sessions (not an interval)",
        "kind": "dot-range",
        "unit": "calls",
        "whisker": "minmax",
        "yLabel": "Tool calls",
        "viz": "IntervalDotPlot",
        "series": [
          {
            "name": "Tool calls per session",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 7.5,
                "lo": 3,
                "hi": 14,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 7.5,
                "lo": 5,
                "hi": 14,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 12.5,
                "lo": 8,
                "hi": 18,
                "n": 12
              }
            ]
          }
        ],
        "note": "Claude Code counts its tool calls (Bash, Read, Edit, Write, Glob, Grep). Codex CLI counts shell commands and file changes; it has no separate read tool, so it reads files with shell commands. Turns are not compared: Codex reports one turn per run.",
        "sourceIds": [
          "agent-coding-agents"
        ]
      },
      {
        "id": "coding-agents-token-mix",
        "title": "Tokens per session, as each CLI reports them",
        "subtitle": "Mean per session by kind · not comparable across vendors",
        "kind": "stacked-bar",
        "unit": "tokens",
        "yLabel": "Tokens",
        "viz": "TokenStack",
        "series": [
          {
            "name": "Cache read",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 73846,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 96964,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 174560,
                "n": 12
              }
            ]
          },
          {
            "name": "Cache write",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 9163,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 11378,
                "n": 12
              }
            ]
          },
          {
            "name": "Uncached input",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 13,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 16,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 19811,
                "n": 12
              }
            ]
          },
          {
            "name": "Output",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 3359,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 5621,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 4072,
                "n": 12
              }
            ]
          }
        ],
        "note": "Claude Code reports uncached input, cache reads, cache writes and output. Codex CLI reports input (its cached part included), the cached part and output (reasoning included), and no cache writes. Different tokenizers and context handling: compare Sonnet with Opus here, not Claude with Codex.",
        "sourceIds": [
          "agent-coding-agents"
        ]
      },
      {
        "id": "coding-agents-diff-lines",
        "title": "Lines changed per session, by kind of file",
        "subtitle": "Mean lines added plus deleted per session, against the base commit",
        "kind": "stacked-bar",
        "unit": "count",
        "yLabel": "Lines",
        "viz": "TokenStack",
        "series": [
          {
            "name": "Code the task is about",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 48.4,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 56.8,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 52.7,
                "n": 12
              }
            ]
          },
          {
            "name": "Tests",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 2.4,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 19.4,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 70.4,
                "n": 12
              }
            ]
          },
          {
            "name": "README and package.json",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.8,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 3.3,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 4.6,
                "n": 12
              }
            ]
          },
          {
            "name": "Work log and notes nobody asked for",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 9.4,
                "n": 12
              }
            ]
          }
        ],
        "note": "Counted from each session’s patch, new files included. Median lines per session: Sonnet 5.5 in Claude Code 45, Opus 5.5 in Claude Code 77.5, GPT-6.1 Sol in Codex CLI 94. Codex wrote a WORKLOG.md in 10 of 12 sessions; no task asked for one.",
        "sourceIds": [
          "agent-coding-agents"
        ]
      },
      {
        "id": "coding-agents-cost-per-pass",
        "title": "List-price cost per passing coding session (calculation)",
        "subtitle": "Reported tokens of all 12 sessions × list price, divided by the passes",
        "kind": "bar",
        "unit": "usd",
        "yLabel": "USD per pass",
        "viz": "CostBars",
        "series": [
          {
            "name": "List-price cost per pass",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.085,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.2229,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI",
                "value": 0.0978,
                "n": 12
              }
            ]
          }
        ],
        "note": "Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.",
        "sourceIds": [
          "agent-coding-agents",
          "calc-repricing",
          "price-anthropic",
          "price-openai"
        ]
      }
    ],
    "tables": [
      {
        "id": "coding-agents-cells",
        "title": "Every task and agent: sessions that passed (of 2)",
        "viz": "HeatMatrix",
        "columns": [
          {
            "key": "task",
            "label": "Task",
            "unit": "text"
          },
          {
            "key": "a0",
            "label": "Sonnet 5.5 in Claude Code",
            "unit": "text"
          },
          {
            "key": "a1",
            "label": "Opus 5.5 in Claude Code",
            "unit": "text"
          },
          {
            "key": "a2",
            "label": "GPT-6.1 Sol in Codex CLI",
            "unit": "text"
          },
          {
            "key": "checks",
            "label": "Hidden checks",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "task": "Pagination fix",
            "a0": "2/2",
            "a1": "2/2",
            "a2": "2/2",
            "checks": "13 hidden tests"
          },
          {
            "task": "CLI --top flag",
            "a0": "2/2",
            "a1": "2/2",
            "a2": "2/2",
            "checks": "11 hidden tests that run the CLI"
          },
          {
            "task": "Invoice refactor",
            "a0": "2/2",
            "a1": "2/2",
            "a2": "2/2",
            "checks": "20 golden behaviour tests and 4 structure tests"
          },
          {
            "task": "LRU cache",
            "a0": "2/2",
            "a1": "2/2",
            "a2": "2/2",
            "checks": "16 hidden tests"
          },
          {
            "task": "Queue race",
            "a0": "2/2",
            "a1": "2/2",
            "a2": "2/2",
            "checks": "10 hidden tests"
          },
          {
            "task": "Strict TypeScript types",
            "a0": "2/2",
            "a1": "2/2",
            "a2": "2/2",
            "checks": "static checks, tsc with a hidden usage file (11 @ts-expect-error probes) and 4 runtime tests"
          }
        ]
      },
      {
        "id": "coding-agents-tasks",
        "title": "The six tasks and their controls",
        "columns": [
          {
            "key": "task",
            "label": "Task",
            "unit": "text"
          },
          {
            "key": "what",
            "label": "What the agent had to do",
            "unit": "text"
          },
          {
            "key": "checks",
            "label": "Hidden checks",
            "unit": "text"
          },
          {
            "key": "base",
            "label": "Base repository",
            "unit": "text"
          },
          {
            "key": "reference",
            "label": "Reference patch",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "task": "Pagination fix",
            "what": "Fix a pagination bug that spans two files (offset, page count, next and previous flags, validation) to the README rules",
            "checks": "13 hidden tests",
            "base": "fails (tests 1/13)",
            "reference": "passes (tests 13/13)"
          },
          {
            "task": "CLI --top flag",
            "what": "Add a --top <n> flag with strict validation and exact error text to a small CLI",
            "checks": "11 hidden tests that run the CLI",
            "base": "fails (tests 2/11)",
            "reference": "passes (tests 11/11)"
          },
          {
            "task": "Invoice refactor",
            "what": "Remove duplication without a change in behaviour, floating-point and validation-order quirks included",
            "checks": "20 golden behaviour tests and 4 structure tests",
            "base": "fails (tests 21/24)",
            "reference": "passes (tests 24/24)"
          },
          {
            "task": "LRU cache",
            "what": "Implement an LRU cache to a written spec (recency, eviction callback, TTL with an injected clock, SameValueZero keys, O(1) get and set)",
            "checks": "16 hidden tests",
            "base": "fails (tests 0/16)",
            "reference": "passes (tests 16/16)"
          },
          {
            "task": "Queue race",
            "what": "Fix a concurrency race and an onIdle hang in an async task queue",
            "checks": "10 hidden tests",
            "base": "fails (tests 4/10)",
            "reference": "passes (tests 10/10)"
          },
          {
            "task": "Strict TypeScript types",
            "what": "Make tsc pass on strict settings with no any and no suppressions; the usage file and the config stay unchanged",
            "checks": "static checks, tsc with a hidden usage file (11 @ts-expect-error probes) and 4 runtime tests",
            "base": "fails (tests 4/4; static checks fail)",
            "reference": "passes (tests 4/4)"
          }
        ]
      }
    ],
    "related": [
      "hard-model-head-to-head",
      "model-head-to-head",
      "swe-bench-opus-vs-sonnet",
      "effort-ladder"
    ],
    "hero": {
      "statIds": [
        "coding-agents-time-ratio"
      ]
    }
  },
  "sources": [
    {
      "id": "agent-coding-agents",
      "title": "Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks",
      "kind": "run",
      "date": "2026-10-06",
      "note": "Six small Node.js repositories with hidden tests; controls before the first session (every base fails, every reference passes). Claude Code 2.1.286 with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol at medium effort, 2 repetitions per task, OS sandboxes without network, protocol declared before the first session, every session kept. Gemini CLI was probed and not run (browser login).",
      "data": [
        "/benchmarks/raw/coding-agents/sessions.json"
      ]
    },
    {
      "id": "price-anthropic",
      "title": "Anthropic list prices (Claude models)",
      "kind": "price-list",
      "date": "2026-09-21",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "note": "Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price."
    },
    {
      "id": "price-openai",
      "title": "OpenAI list prices",
      "kind": "price-list",
      "date": "2026-10-03",
      "url": "https://developers.openai.com/api/docs/pricing",
      "note": "Token prices as listed by the vendor on 2026-10-03."
    },
    {
      "id": "calc-repricing",
      "title": "Repricing calculation",
      "kind": "calculation",
      "date": "2026-10-05",
      "note": "Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes."
    }
  ]
}
