{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/thinking-token-bill",
  "study": {
    "slug": "thinking-token-bill",
    "title": "How much of an AI bill is thinking? Reasoning tokens by model and effort",
    "seoTitle": "Thinking tokens by model and effort: share and cost",
    "description": "Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.",
    "question": "Across 378 recorded calls of Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol, how many output tokens come from reasoning (thinking)? What do they cost at list price per call and per strict pass, and do they track time? The calls cover eight hard tasks, five short tasks and several efforts.",
    "answer": "On eight hard tasks, reasoning tokens were 46% to 92% of the output tokens of a median call (16 to 24 calls per configuration). Haiku 4.5 had the highest median (92%), GPT-6.1 Sol (medium) had the lowest (46%) and the other 5 sat between 54% and 64%. Per-call ranges are wide (Sonnet 5.5 0% to 96%). The ranges of every pair overlap, so this run ranks no configuration. Output share is not bill share. At list price (a calculation) the reasoning part of a call cost a mean $0.0023 (GPT-6.1 Sol (medium)) to $0.0537 (Fable 5.1). That was 9% (GPT-6.1 Sol (medium)) to 80% (Haiku 4.5) of the total list-price cost in each configuration. This pooled share divides summed reasoning cost by summed total cost. Input tokens cost money too. The calls ran on flat subscriptions, so this is not a bill. Fable 5.1's thinking cost 8.1x Sonnet 5.5's per call. The output price explains a factor of 5.0 ($50 against $10 per million tokens). More reasoning tokens explain the rest, a factor of 1.6 (means). Higher effort settings had higher mean recorded reasoning counts in these batches. We paired the same eight tasks. At high effort, mean reasoning tokens per call were higher than at low effort on these tasks: Sonnet 5.5 7 of 8, Opus 5.5 8 of 8, GPT-6.1 Sol 8 of 8. The pooled reasoning share rose from low to high effort. Sonnet 5.5 53% to 73%. Opus 5.5 38% to 69%. GPT-6.1 Sol 29% to 59%. It rose at each step (low, medium, high) in all 3 ladders. The reasoning cost per call rose 2.2x for Sonnet 5.5, 3.6x for Opus 5.5 and 3.4x for GPT-6.1 Sol (a calculation). Every effort cell passed 16/16 strictly. The pass count stayed the same on this set, which has a ceiling (95% Wilson 81% to 100% per cell). Recorded reasoning counts correlated with total time. Within each Claude model, the rank correlation between reasoning tokens and total time was 0.85 to 0.98 (Spearman, 24 to 80 calls each). On the same task, 1,000 more reasoning tokens went with 7.6 to 14.5 s more time (a calculation). For GPT-6.1 Sol in the Codex CLI, the recorded rank correlation was: 0.56 (48 calls). Thinking costs money whether or not the call passes. Haiku 4.5 passed 11/24 strictly (95% Wilson 28% to 65%). The 13 calls that did not pass held 61% of its reasoning cost (a calculation).",
    "date": "2026-10-06",
    "updated": "2026-10-06",
    "tags": [
      "thought-experiment",
      "calculation",
      "reasoning-tokens",
      "thinking-tokens",
      "effort",
      "llm-pricing",
      "claude-haiku",
      "claude-sonnet",
      "claude-opus",
      "claude-fable",
      "gpt-6-1-sol",
      "claude-code",
      "codex-cli",
      "latency"
    ],
    "method": [
      "This study is a calculation over receipts that already exist. We made no new model call. The inputs are three recorded runs. The hard head-to-head gives 152 completed calls. We leave its 30 pre-inference blocked attempts out of token calculations. They remain failures of the original route attempt, not model answers. The effort ladder gives 96 new calls. The five-task head-to-head gives 130 calls. All current counted calls completed. Recorded errors count as failures when their usage is known. We keep the separate blocked batch on record. The original Codex route was blocked. A later batch ran after a successful uncounted probe. It resumed after two completed calls and skipped them.",
      "Sum check. Before we compute any share, we test one point. Do the output tokens include the reasoning tokens? We use this accounting assumption and check its consistency with the receipts. We run three tests on each CLI route. Test 1: reasoning never exceeds output. Test 2: within one task and model, one more reasoning token adds about one output token. A slope near 1 supports the assumption. Correlation alone cannot prove the counter semantics. Test 3: we compare characters per remaining token with characters per output token on calls whose reasoning counter is zero. Claude Code: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 290 calls). Within one task and model, each extra reasoning token went with 0.99 extra output tokens (leave-one-task-out range 0.98 to 1.01, 290 calls). Output minus reasoning has 1.94 characters per token. Calls with 0 reasoning have 1.98. For comparison, output including reasoning has only 0.45. Codex CLI: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 88 calls). Within one task and model, each extra reasoning token went with 0.97 extra output tokens (leave-one-task-out range 0.97 to 0.99, 88 calls). Output minus reasoning has 2.99 characters per token. Calls with 0 reasoning have 2.76. For comparison, output including reasoning has only 1.36.",
      "Not reported is not zero. A route reports reasoning when at least one of its calls reports more than 0. Both routes do (Claude Code 212 of 290 calls, Codex CLI 68 of 88 calls). So a reported 0 is the recorded counter, not proof of no internal reasoning. 98 calls reported 0, and their replies look like visible text only (test 3). If any model attempt lacks usable token counts, this builder returns no study. It never prices unknown usage as zero. 0 calls had no usable count.",
      "Reasoning share of one call = reasoning tokens ÷ output tokens. A configuration gets two numbers. One is the median of its per-call shares. The other is the pooled share (all reasoning tokens ÷ all output tokens). Ranges are the lowest and highest call. They are not intervals.",
      "Hard tasks: we use all calls of each configuration in the hard head-to-head. That is 16 to 24 calls: 8 tasks × 3 repetitions for Claude Code and 8 × 2 for Codex CLI. Effort ladder: 11 cells of 8 tasks × 2 repetitions, the design of the effort-ladder study. We reuse its reference cells from the hard head-to-head (Claude repetitions 1-2 only). Short tasks: the five validated tasks (10 to 15 calls per configuration).",
      "Cost is a calculation at list price, not a bill. The calls ran on flat subscriptions. Reasoning cost = reasoning tokens × the model's output price. Remaining output cost = (output tokens − reasoning tokens) × the same price. We use the rest as a visible-answer estimate. Input cost covers the whole prompt. Cache writes use the one-hour list-price assumption; the receipts do not state the cache lifetime. Prices per million output tokens: Claude Haiku 4.5 $5, Claude Sonnet 5.5 $10, Claude Opus 5.5 $20, Claude Fable 5.1 $50, GPT-6.1 Sol $10. Sources: Anthropic list prices of 2026-09-21, OpenAI of 2026-10-03.",
      "Cost per strict pass = the list-price cost of all calls in the cell, failures included, ÷ the cell's strict passes. Hard and effort cells use the hard-set strict rule. Short-task cells use their original rule, which can strip a wrapping fence.",
      "Time link: we use all calls of the hard head-to-head and the effort ladder. For each model and route, we compute the Spearman rank correlation of reasoning tokens with total time (with a leave-one-task-out sensitivity range, not a confidence interval). We also compute the least-squares slope in seconds per 1,000 reasoning tokens. Then we compute the slope within task. We centre each task on its own mean. This controls for differences in task means, not effort or batch effects.",
      "Isolation, tasks, validators and flags follow the hard head-to-head and the five-task head-to-head. Each call ran in a fresh empty folder with tools off and one turn. The timeout was 300 s per call (180 s for the short tasks). One call ran at a time per account.",
      "Claude Code calls ran with an output cap of 16,000 tokens in all three runs. The highest output of any analysed Claude call was 9,321, so no call hit the cap. The Codex CLI has no cap setting. Its highest output was 2,569."
    ],
    "caveats": [
      "The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.",
      "The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.",
      "These are hand-built, tuned case sets, not random workload samples. The short set and most hard-set configurations hit a pass ceiling. Repeats of the same tasks are not independent. Wilson pass intervals describe counted calls under an independence assumption, not performance on new tasks.",
      "Each CLI reports its own reasoning counter, and we never see the reasoning text. So we cannot check what each vendor counts. The sum check shows only that the counters are consistent with inclusion in output; it does not prove their semantics. Shares from Claude Code and Codex CLI do not measure like for like how much each model thinks.",
      "Every prompt asks for a short reply in a strict format (code, JSON, one line or a regex, with no explanation). So the visible answer is short and the reasoning share is high. Longer visible replies could change the share. We did not measure that workload.",
      "Small cells: 16 to 24 calls per configuration over 8 tasks. Per-call ranges are wide and overlap for every pair of configurations. So the medians describe this run and rank nothing. A sentence says one side is ahead only when its ranges do not overlap.",
      "Input includes CLI context that these receipts do not separately count. It moves with cache hits. So the part of the call cost that goes to reasoning depends on the CLI and on the cache, not only on the model. GPT-6.1 Sol at medium and at high effort ran in different batches and show very different input cost per call.",
      "Every effort-ladder cell passed 16/16, so the set has a ceiling. Higher effort had higher mean recorded reasoning cost in these batches. It does not show that more thinking never helps on harder work.",
      "The time link is a correlation, not a cause. Task, effort and batch can affect reasoning, visible output and time together. Total time includes CLI start-up. The within-task slope controls for differences in task means. It does not remove effort or batch effects. Output tokens (reasoning plus visible answer) track time at least as closely as reasoning alone. The Codex CLI slope (36.4 s per 1,000 reasoning tokens) has a calculated ratio of 2.5 to 4.8 times the Claude Code slopes. We did not test why. Calls are repeats of 8 tasks, so they are not independent. The ranges omit one whole task at a time. They show sensitivity to the task mix, not 95% coverage.",
      "Effort levels are not the same scale across vendors. \"Default\" means we did not pass the effort flag, and the CLI chose. The effort-ladder reference cells ran in a different batch and hour than the new cells.",
      "List-price costs are calculations, because the calls used flat subscriptions. A price change moves every cost here. It leaves every token count unchanged."
    ],
    "sourceIds": [
      "calc-thinking-bill",
      "agent-provider-h2h-hard",
      "agent-effort-ladder",
      "agent-provider-h2h",
      "price-anthropic",
      "price-openai"
    ],
    "hero": {
      "statIds": [
        "thinking-bill-share-highest",
        "thinking-bill-share-lowest"
      ]
    },
    "stats": [
      {
        "id": "thinking-bill-calls",
        "label": "Recorded calls analysed (no new calls)",
        "value": 378,
        "unit": "count",
        "display": "378 (152 hard head-to-head, 96 new effort-ladder, 130 five-task head-to-head)",
        "n": 378
      },
      {
        "id": "thinking-bill-zero-reasoning",
        "label": "Calls that reported 0 reasoning tokens (kept as 0, not \"not reported\")",
        "value": 98,
        "unit": "count",
        "display": "98 of 378; 0 calls had no usable count",
        "n": 378
      },
      {
        "id": "thinking-bill-sum-claude",
        "label": "Sum check, Claude Code: output tokens gained per extra reasoning token, same task and model (calculation)",
        "value": 0.993,
        "unit": "ratio",
        "display": "0.99 (leave-one-task-out range 0.98 to 1.01; 290 calls)",
        "n": 290,
        "note": "About 1 is consistent with reasoning inside output. This correlation does not prove how a CLI counts tokens."
      },
      {
        "id": "thinking-bill-sum-codex",
        "label": "Sum check, Codex CLI: output tokens gained per extra reasoning token, same task and model (calculation)",
        "value": 0.967,
        "unit": "ratio",
        "display": "0.97 (leave-one-task-out range 0.97 to 0.99; 88 calls)",
        "n": 88,
        "note": "About 1 is consistent with reasoning inside output. This correlation does not prove how a CLI counts tokens."
      },
      {
        "id": "thinking-bill-share-highest",
        "label": "Highest median reasoning share of output tokens, hard tasks (calculation)",
        "value": 0.9168,
        "unit": "rate",
        "display": "92% (Claude Haiku 4.5 · Claude Code; range 76% to 99%)",
        "n": 24
      },
      {
        "id": "thinking-bill-share-lowest",
        "label": "Lowest median reasoning share of output tokens, hard tasks (calculation)",
        "value": 0.4633,
        "unit": "rate",
        "display": "46% (GPT-6.1 Sol (medium) · Codex CLI; range 11% to 87%)",
        "n": 16
      },
      {
        "id": "thinking-bill-reasoning-cost-highest",
        "label": "Highest mean list-price cost of reasoning per call, hard tasks (calculation)",
        "value": 0.053696,
        "unit": "usd",
        "display": "$0.0537 (Claude Fable 5.1 · Claude Code; call range $0.0037 to $0.2944)",
        "n": 24
      },
      {
        "id": "thinking-bill-reasoning-cost-lowest",
        "label": "Lowest mean list-price cost of reasoning per call, hard tasks (calculation)",
        "value": 0.002273,
        "unit": "usd",
        "display": "$0.0023 (GPT-6.1 Sol (medium) · Codex CLI; call range $0.0006 to $0.0084)",
        "n": 16
      },
      {
        "id": "thinking-bill-cost-share-highest",
        "label": "Highest pooled reasoning share of total list-price cost, hard tasks (calculation)",
        "value": 0.7953,
        "unit": "rate",
        "display": "80% (Claude Haiku 4.5 · Claude Code; call range 40% to 90%)",
        "n": 24
      },
      {
        "id": "thinking-bill-cost-share-lowest",
        "label": "Lowest pooled reasoning share of total list-price cost, hard tasks (calculation)",
        "value": 0.0887,
        "unit": "rate",
        "display": "9% (GPT-6.1 Sol (medium) · Codex CLI; call range 2% to 20%)",
        "n": 16
      },
      {
        "id": "thinking-bill-fable-vs-sonnet",
        "label": "Reasoning cost per call, Fable 5.1 ÷ Sonnet 5.5, hard tasks (calculation, means)",
        "value": 8.056,
        "unit": "ratio",
        "display": "8.1x ($0.0537 vs $0.0067)",
        "n": 24
      },
      {
        "id": "thinking-bill-haiku-failed-share",
        "label": "Share of Haiku 4.5 reasoning cost spent on calls that did not pass (calculation)",
        "value": 0.6114,
        "unit": "rate",
        "display": "61% (13 of 24 calls did not pass strictly)",
        "n": 24
      },
      {
        "id": "thinking-bill-effort-tasks-sonnet-5-5",
        "label": "Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Sonnet 5.5 · Claude Code (paired, same tasks; calculation)",
        "value": 7,
        "unit": "count",
        "display": "7 of 8 tasks",
        "n": 8,
        "note": "Each task ran 2 times at each effort. Eight tasks is a small sample; this counts tasks, it is not an interval."
      },
      {
        "id": "thinking-bill-effort-tasks-opus-5-5",
        "label": "Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Opus 5.5 · Claude Code (paired, same tasks; calculation)",
        "value": 8,
        "unit": "count",
        "display": "8 of 8 tasks",
        "n": 8,
        "note": "Each task ran 2 times at each effort. Eight tasks is a small sample; this counts tasks, it is not an interval."
      },
      {
        "id": "thinking-bill-effort-tasks-gpt-6-1-sol",
        "label": "Tasks where mean reasoning tokens per call were higher at high than at low effort, GPT-6.1 Sol · Codex CLI (paired, same tasks; calculation)",
        "value": 8,
        "unit": "count",
        "display": "8 of 8 tasks",
        "n": 8,
        "note": "Each task ran 2 times at each effort. Eight tasks is a small sample; this counts tasks, it is not an interval."
      },
      {
        "id": "thinking-bill-effort-ratio-sonnet-5-5",
        "label": "Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means)",
        "value": 2.17,
        "unit": "ratio",
        "display": "2.2x ($0.0043 at low, $0.0094 at high)",
        "n": 16
      },
      {
        "id": "thinking-bill-effort-ratio-opus-5-5",
        "label": "Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means)",
        "value": 3.584,
        "unit": "ratio",
        "display": "3.6x ($0.0050 at low, $0.0180 at high)",
        "n": 16
      },
      {
        "id": "thinking-bill-effort-ratio-gpt-6-1-sol",
        "label": "Reasoning cost per call, high ÷ low effort, GPT-6.1 Sol · Codex CLI (calculation, means)",
        "value": 3.366,
        "unit": "ratio",
        "display": "3.4x ($0.0012 at low, $0.0041 at high)",
        "n": 16
      },
      {
        "id": "thinking-bill-time-rho-claude",
        "label": "Spearman, reasoning tokens vs total time, per Claude model: lowest (calculation)",
        "value": 0.852,
        "unit": "score",
        "display": "0.85 to 0.98 across 4 Claude models",
        "n": 200,
        "note": "The value is the lowest of the per-model correlations; the display gives the range."
      },
      {
        "id": "thinking-bill-time-slope-claude",
        "label": "Seconds per 1,000 reasoning tokens on the same task, per Claude model: lowest (calculation)",
        "value": 7.649,
        "unit": "seconds",
        "display": "7.6 to 14.5 s across 4 Claude models",
        "n": 200,
        "note": "The value is the lowest of the per-model slopes; the display gives the range."
      },
      {
        "id": "thinking-bill-time-rho-codex",
        "label": "Spearman, reasoning tokens vs total time, GPT-6.1 Sol in Codex CLI (calculation)",
        "value": 0.556,
        "unit": "score",
        "display": "0.56 (leave-one-task-out range 0.34 to 0.65; 48 calls)",
        "n": 48,
        "note": "The range omits one whole task at a time. It is a sensitivity check, not a confidence interval."
      }
    ],
    "charts": [
      {
        "id": "thinking-bill-share",
        "title": "Reasoning share of output tokens per call on hard tasks (calculation)",
        "subtitle": "Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)",
        "kind": "bar",
        "unit": "percent",
        "yLabel": "Reasoning share of output tokens (%)",
        "series": [
          {
            "name": "Median call",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 91.68,
                "lo": 76.46,
                "hi": 99.27,
                "n": 24
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 64.24,
                "lo": 23.44,
                "hi": 97.19,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 57.01,
                "lo": 29.19,
                "hi": 90.8,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 54.79,
                "lo": 29.92,
                "hi": 95.6,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 54.54,
                "lo": 0,
                "hi": 95.91,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 54.43,
                "lo": 36.14,
                "hi": 96.23,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 46.33,
                "lo": 11.42,
                "hi": 86.85,
                "n": 16
              }
            ]
          }
        ],
        "note": "Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.",
        "whisker": "minmax",
        "polarity": "none",
        "sourceIds": [
          "calc-thinking-bill",
          "agent-provider-h2h-hard",
          "agent-effort-ladder",
          "agent-provider-h2h",
          "price-anthropic",
          "price-openai"
        ]
      },
      {
        "id": "thinking-bill-cost-per-call",
        "title": "List-price cost per call: reasoning, remaining output and input (calculation)",
        "subtitle": "Mean per call on the hard tasks; the three parts add up to the call",
        "kind": "stacked-bar",
        "unit": "usd",
        "yLabel": "USD per call (list price)",
        "series": [
          {
            "name": "Reasoning (output tokens)",
            "points": [
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 0.053696,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 0.017969,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0.024492,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.012528,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.002273,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 0.004114,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.006665,
                "n": 24
              }
            ]
          },
          {
            "name": "Remaining output (visible-answer estimate)",
            "points": [
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 0.018702,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 0.00799,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0.001795,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.008003,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.003025,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 0.002905,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.003672,
                "n": 24
              }
            ]
          },
          {
            "name": "Input (prompt, cache priced)",
            "points": [
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 0.02091,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 0.007407,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0.00451,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.00771,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.020339,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 0.008117,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.004012,
                "n": 24
              }
            ]
          }
        ],
        "note": "Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.",
        "polarity": "none",
        "sourceIds": [
          "calc-thinking-bill",
          "agent-provider-h2h-hard",
          "agent-effort-ladder",
          "agent-provider-h2h",
          "price-anthropic",
          "price-openai"
        ]
      },
      {
        "id": "thinking-bill-by-effort",
        "title": "Reasoning cost per strict pass by effort, with the total (calculation)",
        "subtitle": "List price ÷ strict passes; every cell is 8 tasks × 2 repetitions",
        "kind": "grouped-bar",
        "unit": "usd",
        "yLabel": "USD per strict pass (list price)",
        "series": [
          {
            "name": "Reasoning cost per strict pass",
            "points": [
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 0.004317,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code",
                "value": 0.005946,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code",
                "value": 0.009369,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.006299,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code",
                "value": 0.005031,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code",
                "value": 0.01344,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 0.018034,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.013104,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 0.001223,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.002273,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 0.004114,
                "n": 16
              }
            ]
          },
          {
            "name": "Total cost per strict pass",
            "points": [
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 0.012191,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code",
                "value": 0.01352,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code",
                "value": 0.016705,
                "n": 16
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.013978,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code",
                "value": 0.021152,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code",
                "value": 0.029475,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 0.033677,
                "n": 16
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.028925,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 0.012837,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 0.025637,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 0.015137,
                "n": 16
              }
            ]
          }
        ],
        "note": "Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). \"Default\" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.",
        "polarity": "none",
        "sourceIds": [
          "calc-thinking-bill",
          "agent-provider-h2h-hard",
          "agent-effort-ladder",
          "agent-provider-h2h",
          "price-anthropic",
          "price-openai"
        ]
      },
      {
        "id": "thinking-bill-vs-time",
        "title": "Reasoning tokens vs total time per call (calculation)",
        "subtitle": "One point per call: 248 calls from the hard head-to-head and the effort ladder",
        "kind": "scatter",
        "unit": "seconds",
        "xLabel": "Reasoning tokens per call",
        "yLabel": "Total time per call (seconds)",
        "series": [
          {
            "name": "Claude Haiku 4.5 · Claude Code",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code · merge-ranges · rep 1",
                "x": 1977,
                "value": 18.53
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · day-hours · rep 1",
                "x": 4372,
                "value": 39
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · csv-parse · rep 1",
                "x": 3590,
                "value": 33.81
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · event-loop-order · rep 1",
                "x": 3781,
                "value": 25.96
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · talk-schedule · rep 1",
                "x": 6236,
                "value": 54.64
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · semver-regex · rep 1",
                "x": 7498,
                "value": 64.48
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · money-refactor · rep 1",
                "x": 1452,
                "value": 15.91
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · q1-sql · rep 1",
                "x": 4305,
                "value": 38.89
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · merge-ranges · rep 2",
                "x": 2606,
                "value": 24.71
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · day-hours · rep 2",
                "x": 3515,
                "value": 26.15
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · csv-parse · rep 2",
                "x": 4497,
                "value": 39.01
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · event-loop-order · rep 2",
                "x": 6575,
                "value": 54.02
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · talk-schedule · rep 2",
                "x": 7954,
                "value": 68.77
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · semver-regex · rep 2",
                "x": 6538,
                "value": 54.49
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · money-refactor · rep 2",
                "x": 1955,
                "value": 15.27
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · q1-sql · rep 2",
                "x": 6691,
                "value": 56.13
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · merge-ranges · rep 3",
                "x": 2703,
                "value": 21.22
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · day-hours · rep 3",
                "x": 4614,
                "value": 40.56
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · csv-parse · rep 3",
                "x": 8569,
                "value": 75.13
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · event-loop-order · rep 3",
                "x": 7065,
                "value": 56.65
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · talk-schedule · rep 3",
                "x": 4922,
                "value": 37.97
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · semver-regex · rep 3",
                "x": 7994,
                "value": 67.03
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · money-refactor · rep 3",
                "x": 2221,
                "value": 17.16
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code · q1-sql · rep 3",
                "x": 5934,
                "value": 47.11
              }
            ]
          },
          {
            "name": "Claude Sonnet 5.5 · Claude Code",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code · merge-ranges · rep 1",
                "x": 0,
                "value": 2.93
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · day-hours · rep 1",
                "x": 1249,
                "value": 21.61
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · csv-parse · rep 1",
                "x": 551,
                "value": 8.85
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · event-loop-order · rep 1",
                "x": 1064,
                "value": 8.17
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · talk-schedule · rep 1",
                "x": 716,
                "value": 7.36
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · semver-regex · rep 1",
                "x": 0,
                "value": 2.26
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · money-refactor · rep 1",
                "x": 0,
                "value": 3.57
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · q1-sql · rep 1",
                "x": 799,
                "value": 8.91
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · merge-ranges · rep 2",
                "x": 0,
                "value": 3.06
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · day-hours · rep 2",
                "x": 1614,
                "value": 19.62
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · csv-parse · rep 2",
                "x": 722,
                "value": 9.56
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · event-loop-order · rep 2",
                "x": 1197,
                "value": 9.87
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · talk-schedule · rep 2",
                "x": 736,
                "value": 7.76
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · semver-regex · rep 2",
                "x": 267,
                "value": 3.67
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · money-refactor · rep 2",
                "x": 545,
                "value": 7.18
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · q1-sql · rep 2",
                "x": 619,
                "value": 8.47
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · merge-ranges · rep 3",
                "x": 0,
                "value": 2.38
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · day-hours · rep 3",
                "x": 3060,
                "value": 34.79
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · csv-parse · rep 3",
                "x": 547,
                "value": 10.61
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · event-loop-order · rep 3",
                "x": 1177,
                "value": 9.57
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · talk-schedule · rep 3",
                "x": 657,
                "value": 7.74
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · semver-regex · rep 3",
                "x": 0,
                "value": 2.49
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · money-refactor · rep 3",
                "x": 0,
                "value": 3.44
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code · q1-sql · rep 3",
                "x": 477,
                "value": 7.56
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · merge-ranges · rep 1",
                "x": 0,
                "value": 2.79
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · merge-ranges · rep 1",
                "x": 0,
                "value": 2.71
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · merge-ranges · rep 1",
                "x": 0,
                "value": 2.93
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · day-hours · rep 1",
                "x": 1489,
                "value": 19.96
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · day-hours · rep 1",
                "x": 1941,
                "value": 22.48
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · day-hours · rep 1",
                "x": 2597,
                "value": 27.18
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · csv-parse · rep 1",
                "x": 0,
                "value": 4.32
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · csv-parse · rep 1",
                "x": 423,
                "value": 7.83
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · csv-parse · rep 1",
                "x": 770,
                "value": 10.01
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · event-loop-order · rep 1",
                "x": 995,
                "value": 8.8
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · event-loop-order · rep 1",
                "x": 1200,
                "value": 9.98
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · event-loop-order · rep 1",
                "x": 1295,
                "value": 10.89
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · talk-schedule · rep 1",
                "x": 598,
                "value": 6.49
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · talk-schedule · rep 1",
                "x": 651,
                "value": 7.9
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · talk-schedule · rep 1",
                "x": 734,
                "value": 9.07
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · semver-regex · rep 1",
                "x": 0,
                "value": 3.51
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · semver-regex · rep 1",
                "x": 251,
                "value": 4.12
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · semver-regex · rep 1",
                "x": 252,
                "value": 4.01
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · money-refactor · rep 1",
                "x": 0,
                "value": 3.72
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · money-refactor · rep 1",
                "x": 0,
                "value": 3.89
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · money-refactor · rep 1",
                "x": 693,
                "value": 8.09
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · q1-sql · rep 1",
                "x": 545,
                "value": 7.92
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · q1-sql · rep 1",
                "x": 421,
                "value": 7.43
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · q1-sql · rep 1",
                "x": 826,
                "value": 10.79
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · merge-ranges · rep 2",
                "x": 0,
                "value": 2.78
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · merge-ranges · rep 2",
                "x": 0,
                "value": 2.93
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · merge-ranges · rep 2",
                "x": 0,
                "value": 3.49
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · day-hours · rep 2",
                "x": 1091,
                "value": 16.28
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · day-hours · rep 2",
                "x": 2093,
                "value": 24.01
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · day-hours · rep 2",
                "x": 3610,
                "value": 35.81
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · csv-parse · rep 2",
                "x": 0,
                "value": 4.4
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · csv-parse · rep 2",
                "x": 0,
                "value": 4.65
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · csv-parse · rep 2",
                "x": 783,
                "value": 16.1
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · event-loop-order · rep 2",
                "x": 974,
                "value": 9.73
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · event-loop-order · rep 2",
                "x": 1098,
                "value": 8.72
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · event-loop-order · rep 2",
                "x": 1226,
                "value": 10.77
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · talk-schedule · rep 2",
                "x": 624,
                "value": 6.71
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · talk-schedule · rep 2",
                "x": 624,
                "value": 9.61
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · talk-schedule · rep 2",
                "x": 755,
                "value": 8.1
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · semver-regex · rep 2",
                "x": 0,
                "value": 2.86
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · semver-regex · rep 2",
                "x": 245,
                "value": 3.82
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · semver-regex · rep 2",
                "x": 243,
                "value": 3.71
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · money-refactor · rep 2",
                "x": 0,
                "value": 5.16
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · money-refactor · rep 2",
                "x": 0,
                "value": 3.99
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · money-refactor · rep 2",
                "x": 580,
                "value": 7.61
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code · q1-sql · rep 2",
                "x": 591,
                "value": 7.68
              },
              {
                "label": "Claude Sonnet 5.5 (medium) · Claude Code · q1-sql · rep 2",
                "x": 566,
                "value": 9.18
              },
              {
                "label": "Claude Sonnet 5.5 (high) · Claude Code · q1-sql · rep 2",
                "x": 626,
                "value": 8.55
              }
            ]
          },
          {
            "name": "Claude Opus 5.5 · Claude Code",
            "points": [
              {
                "label": "Claude Opus 5.5 · Claude Code · merge-ranges · rep 1",
                "x": 100,
                "value": 4.24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · merge-ranges · rep 1",
                "x": 188,
                "value": 5.36
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · day-hours · rep 1",
                "x": 1580,
                "value": 26.3
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · day-hours · rep 1",
                "x": 3301,
                "value": 63
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · csv-parse · rep 1",
                "x": 543,
                "value": 10.98
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · csv-parse · rep 1",
                "x": 528,
                "value": 11.1
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · event-loop-order · rep 1",
                "x": 958,
                "value": 11.91
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · event-loop-order · rep 1",
                "x": 1274,
                "value": 13.6
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · talk-schedule · rep 1",
                "x": 574,
                "value": 7.82
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · talk-schedule · rep 1",
                "x": 656,
                "value": 8.9
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · semver-regex · rep 1",
                "x": 300,
                "value": 5.35
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · semver-regex · rep 1",
                "x": 293,
                "value": 5
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · money-refactor · rep 1",
                "x": 429,
                "value": 8.15
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · money-refactor · rep 1",
                "x": 528,
                "value": 9.12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · q1-sql · rep 1",
                "x": 796,
                "value": 12.78
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · q1-sql · rep 1",
                "x": 920,
                "value": 13.66
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · merge-ranges · rep 2",
                "x": 114,
                "value": 4.65
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · merge-ranges · rep 2",
                "x": 182,
                "value": 5.52
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · day-hours · rep 2",
                "x": 1778,
                "value": 27.21
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · day-hours · rep 2",
                "x": 2756,
                "value": 36.44
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · csv-parse · rep 2",
                "x": 524,
                "value": 11.47
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · csv-parse · rep 2",
                "x": 573,
                "value": 12.22
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · event-loop-order · rep 2",
                "x": 1015,
                "value": 10.2
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · event-loop-order · rep 2",
                "x": 1300,
                "value": 13.45
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · talk-schedule · rep 2",
                "x": 533,
                "value": 7.63
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · talk-schedule · rep 2",
                "x": 655,
                "value": 8.54
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · semver-regex · rep 2",
                "x": 244,
                "value": 4.75
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · semver-regex · rep 2",
                "x": 103,
                "value": 3.63
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · money-refactor · rep 2",
                "x": 214,
                "value": 7.2
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · money-refactor · rep 2",
                "x": 431,
                "value": 8.84
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · q1-sql · rep 2",
                "x": 781,
                "value": 12.44
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · q1-sql · rep 2",
                "x": 739,
                "value": 12.46
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · merge-ranges · rep 3",
                "x": 122,
                "value": 4.45
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · merge-ranges · rep 3",
                "x": 221,
                "value": 5.73
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · day-hours · rep 3",
                "x": 1136,
                "value": 19.74
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · day-hours · rep 3",
                "x": 3027,
                "value": 38.79
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · csv-parse · rep 3",
                "x": 440,
                "value": 11.67
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · csv-parse · rep 3",
                "x": 800,
                "value": 12.62
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · event-loop-order · rep 3",
                "x": 1109,
                "value": 11.48
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · event-loop-order · rep 3",
                "x": 1198,
                "value": 13.5
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · talk-schedule · rep 3",
                "x": 444,
                "value": 6.88
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · talk-schedule · rep 3",
                "x": 553,
                "value": 7.46
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · semver-regex · rep 3",
                "x": 293,
                "value": 5.06
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · semver-regex · rep 3",
                "x": 122,
                "value": 3.98
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · money-refactor · rep 3",
                "x": 199,
                "value": 6.19
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · money-refactor · rep 3",
                "x": 304,
                "value": 10.95
              },
              {
                "label": "Claude Opus 5.5 · Claude Code · q1-sql · rep 3",
                "x": 808,
                "value": 12.96
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code · q1-sql · rep 3",
                "x": 911,
                "value": 12.91
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · merge-ranges · rep 1",
                "x": 0,
                "value": 3.34
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · merge-ranges · rep 1",
                "x": 109,
                "value": 5.51
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · day-hours · rep 1",
                "x": 394,
                "value": 13.02
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · day-hours · rep 1",
                "x": 1993,
                "value": 31.12
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · csv-parse · rep 1",
                "x": 471,
                "value": 9.88
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · csv-parse · rep 1",
                "x": 444,
                "value": 18.26
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · event-loop-order · rep 1",
                "x": 684,
                "value": 8.7
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · event-loop-order · rep 1",
                "x": 1120,
                "value": 12.42
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · talk-schedule · rep 1",
                "x": 498,
                "value": 7.72
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · talk-schedule · rep 1",
                "x": 471,
                "value": 6.99
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · semver-regex · rep 1",
                "x": 88,
                "value": 4.21
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · semver-regex · rep 1",
                "x": 119,
                "value": 6.8
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · money-refactor · rep 1",
                "x": 0,
                "value": 5.65
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · money-refactor · rep 1",
                "x": 183,
                "value": 6.87
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · q1-sql · rep 1",
                "x": 0,
                "value": 10.64
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · q1-sql · rep 1",
                "x": 793,
                "value": 13.13
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · merge-ranges · rep 2",
                "x": 0,
                "value": 3.44
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · merge-ranges · rep 2",
                "x": 115,
                "value": 5.29
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · day-hours · rep 2",
                "x": 587,
                "value": 15.82
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · day-hours · rep 2",
                "x": 1846,
                "value": 31.36
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · csv-parse · rep 2",
                "x": 0,
                "value": 6.21
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · csv-parse · rep 2",
                "x": 655,
                "value": 10.91
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · event-loop-order · rep 2",
                "x": 791,
                "value": 9.16
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · event-loop-order · rep 2",
                "x": 1014,
                "value": 11.19
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · talk-schedule · rep 2",
                "x": 426,
                "value": 7.29
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · talk-schedule · rep 2",
                "x": 565,
                "value": 8.53
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · semver-regex · rep 2",
                "x": 86,
                "value": 3.79
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · semver-regex · rep 2",
                "x": 243,
                "value": 4.78
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · money-refactor · rep 2",
                "x": 0,
                "value": 4.93
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · money-refactor · rep 2",
                "x": 253,
                "value": 7.07
              },
              {
                "label": "Claude Opus 5.5 (low) · Claude Code · q1-sql · rep 2",
                "x": 0,
                "value": 9.23
              },
              {
                "label": "Claude Opus 5.5 (medium) · Claude Code · q1-sql · rep 2",
                "x": 829,
                "value": 13.8
              }
            ]
          },
          {
            "name": "Claude Fable 5.1 · Claude Code",
            "points": [
              {
                "label": "Claude Fable 5.1 · Claude Code · merge-ranges · rep 1",
                "x": 85,
                "value": 4.76
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · day-hours · rep 1",
                "x": 1815,
                "value": 33.71
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · csv-parse · rep 1",
                "x": 903,
                "value": 16.2
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · event-loop-order · rep 1",
                "x": 1765,
                "value": 23.9
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · talk-schedule · rep 1",
                "x": 1090,
                "value": 17.27
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · semver-regex · rep 1",
                "x": 261,
                "value": 9.6
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · money-refactor · rep 1",
                "x": 646,
                "value": 12.91
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · q1-sql · rep 1",
                "x": 875,
                "value": 16.07
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · merge-ranges · rep 2",
                "x": 75,
                "value": 4.46
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · day-hours · rep 2",
                "x": 5889,
                "value": 90
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · csv-parse · rep 2",
                "x": 1088,
                "value": 21.38
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · event-loop-order · rep 2",
                "x": 1228,
                "value": 16.78
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · talk-schedule · rep 2",
                "x": 786,
                "value": 11.23
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · semver-regex · rep 2",
                "x": 236,
                "value": 7.52
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · money-refactor · rep 2",
                "x": 796,
                "value": 13.68
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · q1-sql · rep 2",
                "x": 865,
                "value": 20.31
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · merge-ranges · rep 3",
                "x": 197,
                "value": 15.21
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · day-hours · rep 3",
                "x": 1427,
                "value": 25.09
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · csv-parse · rep 3",
                "x": 1000,
                "value": 15.88
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · event-loop-order · rep 3",
                "x": 1606,
                "value": 26.91
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · talk-schedule · rep 3",
                "x": 836,
                "value": 12.14
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · semver-regex · rep 3",
                "x": 266,
                "value": 5.72
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · money-refactor · rep 3",
                "x": 948,
                "value": 21.85
              },
              {
                "label": "Claude Fable 5.1 · Claude Code · q1-sql · rep 3",
                "x": 1091,
                "value": 24.58
              }
            ]
          },
          {
            "name": "GPT-6.1 Sol · Codex CLI",
            "points": [
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · merge-ranges · rep 1",
                "x": 163,
                "value": 13.2
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · merge-ranges · rep 1",
                "x": 192,
                "value": 14.4
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · day-hours · rep 1",
                "x": 830,
                "value": 61.6
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · day-hours · rep 1",
                "x": 1533,
                "value": 82.99
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · csv-parse · rep 1",
                "x": 86,
                "value": 14.79
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · csv-parse · rep 1",
                "x": 256,
                "value": 22.91
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · event-loop-order · rep 1",
                "x": 199,
                "value": 8.54
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · event-loop-order · rep 1",
                "x": 347,
                "value": 15.44
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · talk-schedule · rep 1",
                "x": 189,
                "value": 12.63
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · talk-schedule · rep 1",
                "x": 189,
                "value": 13.05
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · semver-regex · rep 1",
                "x": 101,
                "value": 9.63
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · semver-regex · rep 1",
                "x": 306,
                "value": 17.84
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · money-refactor · rep 1",
                "x": 89,
                "value": 11.74
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · money-refactor · rep 1",
                "x": 224,
                "value": 19.39
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · q1-sql · rep 1",
                "x": 61,
                "value": 13.76
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · q1-sql · rep 1",
                "x": 181,
                "value": 18.4
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · merge-ranges · rep 2",
                "x": 137,
                "value": 14.63
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · merge-ranges · rep 2",
                "x": 226,
                "value": 15.12
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · day-hours · rep 2",
                "x": 839,
                "value": 56.39
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · day-hours · rep 2",
                "x": 1750,
                "value": 92.21
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · csv-parse · rep 2",
                "x": 112,
                "value": 16.64
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · csv-parse · rep 2",
                "x": 220,
                "value": 20.56
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · event-loop-order · rep 2",
                "x": 251,
                "value": 12.5
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · event-loop-order · rep 2",
                "x": 375,
                "value": 22.51
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · talk-schedule · rep 2",
                "x": 197,
                "value": 11.81
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · talk-schedule · rep 2",
                "x": 229,
                "value": 12.21
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · semver-regex · rep 2",
                "x": 193,
                "value": 13.01
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · semver-regex · rep 2",
                "x": 189,
                "value": 11.67
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · money-refactor · rep 2",
                "x": 71,
                "value": 11.85
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · money-refactor · rep 2",
                "x": 144,
                "value": 15.43
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI · q1-sql · rep 2",
                "x": 119,
                "value": 17.81
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI · q1-sql · rep 2",
                "x": 222,
                "value": 19.62
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · merge-ranges · rep 1",
                "x": 63,
                "value": 17.2
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · day-hours · rep 1",
                "x": 458,
                "value": 44.29
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · csv-parse · rep 1",
                "x": 78,
                "value": 17.66
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · event-loop-order · rep 1",
                "x": 198,
                "value": 14.61
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · talk-schedule · rep 1",
                "x": 189,
                "value": 13.6
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · semver-regex · rep 1",
                "x": 44,
                "value": 8.57
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · money-refactor · rep 1",
                "x": 62,
                "value": 13.12
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · q1-sql · rep 1",
                "x": 0,
                "value": 13.64
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · merge-ranges · rep 2",
                "x": 0,
                "value": 7.94
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · day-hours · rep 2",
                "x": 400,
                "value": 40.71
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · csv-parse · rep 2",
                "x": 0,
                "value": 17.85
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · event-loop-order · rep 2",
                "x": 185,
                "value": 11.1
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · talk-schedule · rep 2",
                "x": 189,
                "value": 12.41
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · semver-regex · rep 2",
                "x": 50,
                "value": 8.9
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · money-refactor · rep 2",
                "x": 0,
                "value": 9.54
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI · q1-sql · rep 2",
                "x": 40,
                "value": 14.42
              }
            ]
          }
        ],
        "note": "Calculation, not a run: each point is one recorded call. Spearman rank correlations and slopes are in the table. Total time includes CLI start-up and the visible answer. Task, effort and batch can affect both counts and time. The plot shows association, not cause. 248 calls over 8 tasks; calls are not independent.",
        "polarity": "none",
        "sourceIds": [
          "calc-thinking-bill",
          "agent-provider-h2h-hard",
          "agent-effort-ladder",
          "agent-provider-h2h",
          "price-anthropic",
          "price-openai"
        ]
      },
      {
        "id": "thinking-bill-short-vs-hard",
        "title": "Reasoning share on short tasks vs hard tasks (calculation)",
        "subtitle": "Median call per configuration; five short tasks and eight hard tasks",
        "kind": "grouped-bar",
        "unit": "percent",
        "yLabel": "Reasoning share of output tokens (median call, %)",
        "series": [
          {
            "name": "Eight hard tasks",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 91.68,
                "lo": 76.46,
                "hi": 99.27,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 54.54,
                "lo": 0,
                "hi": 95.91,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 54.79,
                "lo": 29.92,
                "hi": 95.6,
                "n": 24
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 54.43,
                "lo": 36.14,
                "hi": 96.23,
                "n": 24
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 64.24,
                "lo": 23.44,
                "hi": 97.19,
                "n": 24
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 46.33,
                "lo": 11.42,
                "hi": 86.85,
                "n": 16
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 57.01,
                "lo": 29.19,
                "hi": 90.8,
                "n": 16
              }
            ]
          },
          {
            "name": "Five short tasks",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 90.19,
                "lo": 73.1,
                "hi": 97.59,
                "n": 15
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0,
                "lo": 0,
                "hi": 72.75,
                "n": 15
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0,
                "lo": 0,
                "hi": 93.33,
                "n": 15
              },
              {
                "label": "Claude Opus 5.5 (high) · Claude Code",
                "value": 43.59,
                "lo": 0,
                "hi": 93.33,
                "n": 15
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 0,
                "lo": 0,
                "hi": 74.01,
                "n": 15
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 41.05,
                "lo": 0,
                "hi": 71.43,
                "n": 15
              },
              {
                "label": "GPT-6.1 Sol (high) · Codex CLI",
                "value": 58.06,
                "lo": 0,
                "hi": 75.76,
                "n": 15
              }
            ]
          }
        ],
        "note": "Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.",
        "whisker": "minmax",
        "polarity": "none",
        "sourceIds": [
          "calc-thinking-bill",
          "agent-provider-h2h-hard",
          "agent-effort-ladder",
          "agent-provider-h2h",
          "price-anthropic",
          "price-openai"
        ]
      }
    ],
    "tables": [
      {
        "id": "thinking-bill-hard-cells",
        "title": "Reasoning tokens and list-price cost per call, hard tasks (calculation)",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "calls",
            "label": "Calls",
            "unit": "count"
          },
          {
            "key": "strict",
            "label": "Strict passes",
            "unit": "text"
          },
          {
            "key": "medOut",
            "label": "Median output tokens",
            "unit": "tokens"
          },
          {
            "key": "medRea",
            "label": "Median reasoning tokens",
            "unit": "tokens"
          },
          {
            "key": "tokenRange",
            "label": "Output / reasoning tokens, lowest to highest call",
            "unit": "text"
          },
          {
            "key": "costRange",
            "label": "Reasoning / total cost, lowest to highest call (USD, calculation)",
            "unit": "text"
          },
          {
            "key": "costShareRange",
            "label": "Reasoning cost share, lowest to highest call (calculation)",
            "unit": "text"
          },
          {
            "key": "zero",
            "label": "Calls with 0 reasoning tokens",
            "unit": "count"
          },
          {
            "key": "medShare",
            "label": "Median call: reasoning share of output",
            "unit": "rate"
          },
          {
            "key": "shareRange",
            "label": "Share, lowest to highest call",
            "unit": "text"
          },
          {
            "key": "pooled",
            "label": "Pooled share (all reasoning ÷ all output)",
            "unit": "rate"
          },
          {
            "key": "meanReasoning",
            "label": "Reasoning cost per call (USD, calculation)",
            "unit": "usd"
          },
          {
            "key": "meanVisible",
            "label": "Remaining output cost per call (USD, calculation)",
            "unit": "usd"
          },
          {
            "key": "meanInput",
            "label": "Input cost per call (USD)",
            "unit": "usd"
          },
          {
            "key": "meanTotal",
            "label": "Total cost per call (USD)",
            "unit": "usd"
          },
          {
            "key": "costShare",
            "label": "Pooled reasoning share of list-price cost (calculation)",
            "unit": "rate"
          },
          {
            "key": "reasoningPerPass",
            "label": "Reasoning cost per strict pass (USD)",
            "unit": "usd"
          },
          {
            "key": "totalPerPass",
            "label": "Total cost per strict pass (USD)",
            "unit": "usd"
          }
        ],
        "rows": [
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "calls": 24,
            "strict": "11/24 (95% Wilson 28% to 65%)",
            "medOut": 5064,
            "medRea": 4556,
            "tokenRange": "1899 to 9321 / 1452 to 8569",
            "costRange": "$0.0073 to $0.0428 / $0.0148 to $0.0505",
            "costShareRange": "40% to 90%",
            "zero": 0,
            "medShare": 0.9168,
            "shareRange": "76% to 99%",
            "pooled": 0.9317,
            "meanReasoning": 0.024492,
            "meanVisible": 0.001795,
            "meanInput": 0.00451,
            "meanTotal": 0.030798,
            "costShare": 0.7953,
            "reasoningPerPass": 0.053438,
            "totalPerPass": 0.067196
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "calls": 24,
            "strict": "24/24 (95% Wilson 86% to 100%)",
            "medOut": 1050,
            "medRea": 585,
            "tokenRange": "176 to 3895 / 0 to 3060",
            "costRange": "$0.0000 to $0.0306 / $0.0051 to $0.0424",
            "costShareRange": "0% to 74%",
            "zero": 7,
            "medShare": 0.5454,
            "shareRange": "0% to 96%",
            "pooled": 0.6448,
            "meanReasoning": 0.006665,
            "meanVisible": 0.003672,
            "meanInput": 0.004012,
            "meanTotal": 0.014349,
            "costShare": 0.4645,
            "reasoningPerPass": 0.006665,
            "totalPerPass": 0.014349
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "calls": 24,
            "strict": "24/24 (95% Wilson 86% to 100%)",
            "medOut": 945,
            "medRea": 529,
            "tokenRange": "323 to 2531 / 100 to 1778",
            "costRange": "$0.0020 to $0.0356 / $0.0131 to $0.0572",
            "costShareRange": "10% to 74%",
            "zero": 0,
            "medShare": 0.5479,
            "shareRange": "30% to 96%",
            "pooled": 0.6102,
            "meanReasoning": 0.012528,
            "meanVisible": 0.008003,
            "meanInput": 0.00771,
            "meanTotal": 0.028241,
            "costShare": 0.4436,
            "reasoningPerPass": 0.012528,
            "totalPerPass": 0.028241
          },
          {
            "config": "Claude Opus 5.5 (high) · Claude Code",
            "calls": 24,
            "strict": "24/24 (95% Wilson 86% to 100%)",
            "medOut": 1052,
            "medRea": 614,
            "tokenRange": "285 to 4052 / 103 to 3301",
            "costRange": "$0.0021 to $0.0660 / $0.0121 to $0.0877",
            "costShareRange": "17% to 77%",
            "zero": 0,
            "medShare": 0.5443,
            "shareRange": "36% to 96%",
            "pooled": 0.6922,
            "meanReasoning": 0.017969,
            "meanVisible": 0.00799,
            "meanInput": 0.007407,
            "meanTotal": 0.033366,
            "costShare": 0.5385,
            "reasoningPerPass": 0.017969,
            "totalPerPass": 0.033366
          },
          {
            "config": "Claude Fable 5.1 · Claude Code",
            "calls": 24,
            "strict": "24/24 (95% Wilson 86% to 100%)",
            "medOut": 1366,
            "medRea": 889,
            "tokenRange": "318 to 6465 / 75 to 5889",
            "costRange": "$0.0037 to $0.2944 / $0.0322 to $0.3399",
            "costShareRange": "5% to 87%",
            "zero": 0,
            "medShare": 0.6424,
            "shareRange": "23% to 97%",
            "pooled": 0.7417,
            "meanReasoning": 0.053696,
            "meanVisible": 0.018702,
            "meanInput": 0.02091,
            "meanTotal": 0.093308,
            "costShare": 0.5755,
            "reasoningPerPass": 0.053696,
            "totalPerPass": 0.093308
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 335,
            "medRea": 150,
            "tokenRange": "237 to 1766 / 61 to 839",
            "costRange": "$0.0006 to $0.0084 / $0.0099 to $0.0422",
            "costShareRange": "2% to 20%",
            "zero": 0,
            "medShare": 0.4633,
            "shareRange": "11% to 87%",
            "pooled": 0.429,
            "meanReasoning": 0.002273,
            "meanVisible": 0.003025,
            "meanInput": 0.020339,
            "meanTotal": 0.025637,
            "costShare": 0.0887,
            "reasoningPerPass": 0.002273,
            "totalPerPass": 0.025637
          },
          {
            "config": "GPT-6.1 Sol (high) · Codex CLI",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 436,
            "medRea": 225,
            "tokenRange": "284 to 2569 / 144 to 1750",
            "costRange": "$0.0014 to $0.0175 / $0.0092 to $0.0307",
            "costShareRange": "8% to 60%",
            "zero": 0,
            "medShare": 0.5701,
            "shareRange": "29% to 91%",
            "pooled": 0.5861,
            "meanReasoning": 0.004114,
            "meanVisible": 0.002905,
            "meanInput": 0.008117,
            "meanTotal": 0.015137,
            "costShare": 0.2718,
            "reasoningPerPass": 0.004114,
            "totalPerPass": 0.015137
          }
        ]
      },
      {
        "id": "thinking-bill-effort-cells",
        "title": "Reasoning tokens and list-price cost by effort, every effort-ladder cell (calculation)",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "calls",
            "label": "Calls",
            "unit": "count"
          },
          {
            "key": "strict",
            "label": "Strict passes",
            "unit": "text"
          },
          {
            "key": "medOut",
            "label": "Median output tokens",
            "unit": "tokens"
          },
          {
            "key": "medRea",
            "label": "Median reasoning tokens",
            "unit": "tokens"
          },
          {
            "key": "tokenRange",
            "label": "Output / reasoning tokens, lowest to highest call",
            "unit": "text"
          },
          {
            "key": "costRange",
            "label": "Reasoning / total cost, lowest to highest call (USD, calculation)",
            "unit": "text"
          },
          {
            "key": "costShareRange",
            "label": "Reasoning cost share, lowest to highest call (calculation)",
            "unit": "text"
          },
          {
            "key": "zero",
            "label": "Calls with 0 reasoning tokens",
            "unit": "count"
          },
          {
            "key": "medShare",
            "label": "Median call: reasoning share of output",
            "unit": "rate"
          },
          {
            "key": "shareRange",
            "label": "Share, lowest to highest call",
            "unit": "text"
          },
          {
            "key": "pooled",
            "label": "Pooled share (all reasoning ÷ all output)",
            "unit": "rate"
          },
          {
            "key": "meanReasoning",
            "label": "Reasoning cost per call (USD, calculation)",
            "unit": "usd"
          },
          {
            "key": "meanVisible",
            "label": "Remaining output cost per call (USD, calculation)",
            "unit": "usd"
          },
          {
            "key": "meanInput",
            "label": "Input cost per call (USD)",
            "unit": "usd"
          },
          {
            "key": "meanTotal",
            "label": "Total cost per call (USD)",
            "unit": "usd"
          },
          {
            "key": "costShare",
            "label": "Pooled reasoning share of list-price cost (calculation)",
            "unit": "rate"
          },
          {
            "key": "reasoningPerPass",
            "label": "Reasoning cost per strict pass (USD)",
            "unit": "usd"
          },
          {
            "key": "totalPerPass",
            "label": "Total cost per strict pass (USD)",
            "unit": "usd"
          }
        ],
        "rows": [
          {
            "config": "Claude Sonnet 5.5 (low) · Claude Code",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 667,
            "medRea": 273,
            "tokenRange": "176 to 2263 / 0 to 1489",
            "costRange": "$0.0000 to $0.0149 / $0.0051 to $0.0261",
            "costShareRange": "0% to 71%",
            "zero": 8,
            "medShare": 0.2142,
            "shareRange": "0% to 95%",
            "pooled": 0.5327,
            "meanReasoning": 0.004317,
            "meanVisible": 0.003787,
            "meanInput": 0.004088,
            "meanTotal": 0.012191,
            "costShare": 0.3541,
            "reasoningPerPass": 0.004317,
            "totalPerPass": 0.012191
          },
          {
            "config": "Claude Sonnet 5.5 (medium) · Claude Code",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 770,
            "medRea": 422,
            "tokenRange": "224 to 2836 / 0 to 2093",
            "costRange": "$0.0000 to $0.0209 / $0.0056 to $0.0318",
            "costShareRange": "0% to 75%",
            "zero": 5,
            "medShare": 0.5162,
            "shareRange": "0% to 96%",
            "pooled": 0.6158,
            "meanReasoning": 0.005946,
            "meanVisible": 0.003709,
            "meanInput": 0.003865,
            "meanTotal": 0.01352,
            "costShare": 0.4398,
            "reasoningPerPass": 0.005946,
            "totalPerPass": 0.01352
          },
          {
            "config": "Claude Sonnet 5.5 (high) · Claude Code",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 1192,
            "medRea": 745,
            "tokenRange": "220 to 4187 / 0 to 3610",
            "costRange": "$0.0000 to $0.0361 / $0.0056 to $0.0453",
            "costShareRange": "0% to 80%",
            "zero": 2,
            "medShare": 0.6131,
            "shareRange": "0% to 96%",
            "pooled": 0.7297,
            "meanReasoning": 0.009369,
            "meanVisible": 0.003471,
            "meanInput": 0.003866,
            "meanTotal": 0.016705,
            "costShare": 0.5608,
            "reasoningPerPass": 0.009369,
            "totalPerPass": 0.016705
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 1054,
            "medRea": 668,
            "tokenRange": "176 to 2185 / 0 to 1614",
            "costRange": "$0.0000 to $0.0161 / $0.0051 to $0.0253",
            "costShareRange": "0% to 74%",
            "zero": 4,
            "medShare": 0.5601,
            "shareRange": "0% to 96%",
            "pooled": 0.6368,
            "meanReasoning": 0.006299,
            "meanVisible": 0.003593,
            "meanInput": 0.004086,
            "meanTotal": 0.013978,
            "costShare": 0.4507,
            "reasoningPerPass": 0.006299,
            "totalPerPass": 0.013978
          },
          {
            "config": "Claude Opus 5.5 (low) · Claude Code",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 594,
            "medRea": 87,
            "tokenRange": "220 to 1569 / 0 to 791",
            "costRange": "$0.0000 to $0.0158 / $0.0108 to $0.0380",
            "costShareRange": "0% to 67%",
            "zero": 7,
            "medShare": 0.3065,
            "shareRange": "0% to 94%",
            "pooled": 0.3785,
            "meanReasoning": 0.005031,
            "meanVisible": 0.008261,
            "meanInput": 0.00786,
            "meanTotal": 0.021152,
            "costShare": 0.2379,
            "reasoningPerPass": 0.005031,
            "totalPerPass": 0.021152
          },
          {
            "config": "Claude Opus 5.5 (medium) · Claude Code",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 853,
            "medRea": 518,
            "tokenRange": "301 to 3075 / 109 to 1993",
            "costRange": "$0.0022 to $0.0399 / $0.0125 to $0.0681",
            "costShareRange": "16% to 74%",
            "zero": 0,
            "medShare": 0.5353,
            "shareRange": "28% to 96%",
            "pooled": 0.609,
            "meanReasoning": 0.01344,
            "meanVisible": 0.00863,
            "meanInput": 0.007405,
            "meanTotal": 0.029475,
            "costShare": 0.456,
            "reasoningPerPass": 0.01344,
            "totalPerPass": 0.029475
          },
          {
            "config": "Claude Opus 5.5 (high) · Claude Code",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 1052,
            "medRea": 614,
            "tokenRange": "285 to 4052 / 103 to 3301",
            "costRange": "$0.0021 to $0.0660 / $0.0121 to $0.0877",
            "costShareRange": "17% to 77%",
            "zero": 0,
            "medShare": 0.5369,
            "shareRange": "36% to 96%",
            "pooled": 0.6865,
            "meanReasoning": 0.018034,
            "meanVisible": 0.008236,
            "meanInput": 0.007407,
            "meanTotal": 0.033677,
            "costShare": 0.5355,
            "reasoningPerPass": 0.018034,
            "totalPerPass": 0.033677
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 945,
            "medRea": 538,
            "tokenRange": "323 to 2531 / 100 to 1778",
            "costRange": "$0.0020 to $0.0356 / $0.0131 to $0.0572",
            "costShareRange": "10% to 72%",
            "zero": 0,
            "medShare": 0.5435,
            "shareRange": "31% to 95%",
            "pooled": 0.6221,
            "meanReasoning": 0.013104,
            "meanVisible": 0.007961,
            "meanInput": 0.00786,
            "meanTotal": 0.028925,
            "costShare": 0.453,
            "reasoningPerPass": 0.013104,
            "totalPerPass": 0.028925
          },
          {
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 284,
            "medRea": 63,
            "tokenRange": "172 to 1310 / 0 to 458",
            "costRange": "$0.0000 to $0.0046 / $0.0092 to $0.0264",
            "costShareRange": "0% to 22%",
            "zero": 4,
            "medShare": 0.2351,
            "shareRange": "0% to 84%",
            "pooled": 0.2909,
            "meanReasoning": 0.001223,
            "meanVisible": 0.00298,
            "meanInput": 0.008635,
            "meanTotal": 0.012837,
            "costShare": 0.0952,
            "reasoningPerPass": 0.001223,
            "totalPerPass": 0.012837
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 335,
            "medRea": 150,
            "tokenRange": "237 to 1766 / 61 to 839",
            "costRange": "$0.0006 to $0.0084 / $0.0099 to $0.0422",
            "costShareRange": "2% to 20%",
            "zero": 0,
            "medShare": 0.4633,
            "shareRange": "11% to 87%",
            "pooled": 0.429,
            "meanReasoning": 0.002273,
            "meanVisible": 0.003025,
            "meanInput": 0.020339,
            "meanTotal": 0.025637,
            "costShare": 0.0887,
            "reasoningPerPass": 0.002273,
            "totalPerPass": 0.025637
          },
          {
            "config": "GPT-6.1 Sol (high) · Codex CLI",
            "calls": 16,
            "strict": "16/16 (95% Wilson 81% to 100%)",
            "medOut": 436,
            "medRea": 225,
            "tokenRange": "284 to 2569 / 144 to 1750",
            "costRange": "$0.0014 to $0.0175 / $0.0092 to $0.0307",
            "costShareRange": "8% to 60%",
            "zero": 0,
            "medShare": 0.5701,
            "shareRange": "29% to 91%",
            "pooled": 0.5861,
            "meanReasoning": 0.004114,
            "meanVisible": 0.002905,
            "meanInput": 0.008117,
            "meanTotal": 0.015137,
            "costShare": 0.2718,
            "reasoningPerPass": 0.004114,
            "totalPerPass": 0.015137
          }
        ]
      },
      {
        "id": "thinking-bill-short-cells",
        "title": "Reasoning tokens and list-price cost per call, five short tasks (calculation)",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "calls",
            "label": "Calls",
            "unit": "count"
          },
          {
            "key": "strict",
            "label": "Strict passes",
            "unit": "text"
          },
          {
            "key": "medOut",
            "label": "Median output tokens",
            "unit": "tokens"
          },
          {
            "key": "medRea",
            "label": "Median reasoning tokens",
            "unit": "tokens"
          },
          {
            "key": "tokenRange",
            "label": "Output / reasoning tokens, lowest to highest call",
            "unit": "text"
          },
          {
            "key": "costRange",
            "label": "Reasoning / total cost, lowest to highest call (USD, calculation)",
            "unit": "text"
          },
          {
            "key": "costShareRange",
            "label": "Reasoning cost share, lowest to highest call (calculation)",
            "unit": "text"
          },
          {
            "key": "zero",
            "label": "Calls with 0 reasoning tokens",
            "unit": "count"
          },
          {
            "key": "medShare",
            "label": "Median call: reasoning share of output",
            "unit": "rate"
          },
          {
            "key": "shareRange",
            "label": "Share, lowest to highest call",
            "unit": "text"
          },
          {
            "key": "pooled",
            "label": "Pooled share (all reasoning ÷ all output)",
            "unit": "rate"
          },
          {
            "key": "meanReasoning",
            "label": "Reasoning cost per call (USD, calculation)",
            "unit": "usd"
          },
          {
            "key": "meanVisible",
            "label": "Remaining output cost per call (USD, calculation)",
            "unit": "usd"
          },
          {
            "key": "meanInput",
            "label": "Input cost per call (USD)",
            "unit": "usd"
          },
          {
            "key": "meanTotal",
            "label": "Total cost per call (USD)",
            "unit": "usd"
          },
          {
            "key": "costShare",
            "label": "Pooled reasoning share of list-price cost (calculation)",
            "unit": "rate"
          },
          {
            "key": "reasoningPerPass",
            "label": "Reasoning cost per strict pass (USD)",
            "unit": "usd"
          },
          {
            "key": "totalPerPass",
            "label": "Total cost per strict pass (USD)",
            "unit": "usd"
          }
        ],
        "rows": [
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "calls": 15,
            "strict": "15/15 (95% Wilson 80% to 100%)",
            "medOut": 367,
            "medRea": 297,
            "tokenRange": "275 to 2851 / 212 to 2643",
            "costRange": "$0.0011 to $0.0132 / $0.0051 to $0.0180",
            "costShareRange": "20% to 74%",
            "zero": 0,
            "medShare": 0.9019,
            "shareRange": "73% to 98%",
            "pooled": 0.9049,
            "meanReasoning": 0.004133,
            "meanVisible": 0.000434,
            "meanInput": 0.00379,
            "meanTotal": 0.008358,
            "costShare": 0.4945,
            "reasoningPerPass": 0.004133,
            "totalPerPass": 0.008358
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "calls": 15,
            "strict": "12/15 (95% Wilson 55% to 93%)",
            "medOut": 107,
            "medRea": 0,
            "tokenRange": "44 to 745 / 0 to 542",
            "costRange": "$0.0000 to $0.0054 / $0.0034 to $0.0102",
            "costShareRange": "0% to 53%",
            "zero": 12,
            "medShare": 0,
            "shareRange": "0% to 73%",
            "pooled": 0.4468,
            "meanReasoning": 0.000882,
            "meanVisible": 0.001092,
            "meanInput": 0.003015,
            "meanTotal": 0.004989,
            "costShare": 0.1768,
            "reasoningPerPass": 0.001103,
            "totalPerPass": 0.006236
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "calls": 15,
            "strict": "15/15 (95% Wilson 80% to 100%)",
            "medOut": 64,
            "medRea": 0,
            "tokenRange": "44 to 853 / 0 to 616",
            "costRange": "$0.0000 to $0.0123 / $0.0059 to $0.0223",
            "costShareRange": "0% to 55%",
            "zero": 9,
            "medShare": 0,
            "shareRange": "0% to 93%",
            "pooled": 0.5912,
            "meanReasoning": 0.002585,
            "meanVisible": 0.001788,
            "meanInput": 0.005714,
            "meanTotal": 0.010088,
            "costShare": 0.2563,
            "reasoningPerPass": 0.002585,
            "totalPerPass": 0.010088
          },
          {
            "config": "Claude Opus 5.5 (low) · Claude Code",
            "calls": 15,
            "strict": "15/15 (95% Wilson 80% to 100%)",
            "medOut": 64,
            "medRea": 0,
            "tokenRange": "44 to 637 / 0 to 405",
            "costRange": "$0.0000 to $0.0081 / $0.0058 to $0.0179",
            "costShareRange": "0% to 45%",
            "zero": 9,
            "medShare": 0,
            "shareRange": "0% to 93%",
            "pooled": 0.4098,
            "meanReasoning": 0.001253,
            "meanVisible": 0.001805,
            "meanInput": 0.005232,
            "meanTotal": 0.00829,
            "costShare": 0.1512,
            "reasoningPerPass": 0.001253,
            "totalPerPass": 0.00829
          },
          {
            "config": "Claude Opus 5.5 (high) · Claude Code",
            "calls": 15,
            "strict": "15/15 (95% Wilson 80% to 100%)",
            "medOut": 78,
            "medRea": 34,
            "tokenRange": "44 to 1094 / 0 to 857",
            "costRange": "$0.0000 to $0.0171 / $0.0059 to $0.0271",
            "costShareRange": "0% to 63%",
            "zero": 7,
            "medShare": 0.4359,
            "shareRange": "0% to 93%",
            "pooled": 0.6562,
            "meanReasoning": 0.003448,
            "meanVisible": 0.001807,
            "meanInput": 0.005233,
            "meanTotal": 0.010488,
            "costShare": 0.3288,
            "reasoningPerPass": 0.003448,
            "totalPerPass": 0.010488
          },
          {
            "config": "Claude Fable 5.1 · Claude Code",
            "calls": 15,
            "strict": "15/15 (95% Wilson 80% to 100%)",
            "medOut": 64,
            "medRea": 0,
            "tokenRange": "4 to 908 / 0 to 672",
            "costRange": "$0.0000 to $0.0336 / $0.0049 to $0.0584",
            "costShareRange": "0% to 67%",
            "zero": 12,
            "medShare": 0,
            "shareRange": "0% to 74%",
            "pooled": 0.5724,
            "meanReasoning": 0.005957,
            "meanVisible": 0.00445,
            "meanInput": 0.010133,
            "meanTotal": 0.020539,
            "costShare": 0.29,
            "reasoningPerPass": 0.005957,
            "totalPerPass": 0.020539
          },
          {
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "calls": 10,
            "strict": "10/10 (95% Wilson 72% to 100%)",
            "medOut": 42,
            "medRea": 20,
            "tokenRange": "28 to 225 / 0 to 103",
            "costRange": "$0.0000 to $0.0010 / $0.0074 to $0.0265",
            "costShareRange": "0% to 11%",
            "zero": 4,
            "medShare": 0.2528,
            "shareRange": "0% to 71%",
            "pooled": 0.3132,
            "meanReasoning": 0.000332,
            "meanVisible": 0.000728,
            "meanInput": 0.008924,
            "meanTotal": 0.009984,
            "costShare": 0.0333,
            "reasoningPerPass": 0.000332,
            "totalPerPass": 0.009984
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "calls": 15,
            "strict": "15/15 (95% Wilson 80% to 100%)",
            "medOut": 42,
            "medRea": 20,
            "tokenRange": "27 to 295 / 0 to 137",
            "costRange": "$0.0000 to $0.0014 / $0.0054 to $0.0269",
            "costShareRange": "0% to 23%",
            "zero": 6,
            "medShare": 0.4105,
            "shareRange": "0% to 71%",
            "pooled": 0.4142,
            "meanReasoning": 0.000512,
            "meanVisible": 0.000724,
            "meanInput": 0.014404,
            "meanTotal": 0.01564,
            "costShare": 0.0327,
            "reasoningPerPass": 0.000512,
            "totalPerPass": 0.01564
          },
          {
            "config": "GPT-6.1 Sol (high) · Codex CLI",
            "calls": 15,
            "strict": "15/15 (95% Wilson 80% to 100%)",
            "medOut": 42,
            "medRea": 21,
            "tokenRange": "28 to 457 / 0 to 289",
            "costRange": "$0.0000 to $0.0029 / $0.0066 to $0.0281",
            "costShareRange": "0% to 29%",
            "zero": 6,
            "medShare": 0.5806,
            "shareRange": "0% to 76%",
            "pooled": 0.5814,
            "meanReasoning": 0.001007,
            "meanVisible": 0.000725,
            "meanInput": 0.011484,
            "meanTotal": 0.013216,
            "costShare": 0.0762,
            "reasoningPerPass": 0.001007,
            "totalPerPass": 0.013216
          }
        ]
      },
      {
        "id": "thinking-bill-time-link",
        "title": "Does reasoning track time? Rank correlation and slope per model (calculation)",
        "columns": [
          {
            "key": "model",
            "label": "Model and route",
            "unit": "text"
          },
          {
            "key": "calls",
            "label": "Calls",
            "unit": "count"
          },
          {
            "key": "rho",
            "label": "Spearman: reasoning tokens vs time",
            "unit": "score"
          },
          {
            "key": "rhoCi",
            "label": "Leave-one-task-out range (not a 95% interval)",
            "unit": "text"
          },
          {
            "key": "rhoOut",
            "label": "Spearman: output tokens vs time",
            "unit": "score"
          },
          {
            "key": "perK",
            "label": "Seconds per 1,000 reasoning tokens (all calls)",
            "unit": "seconds"
          },
          {
            "key": "withinK",
            "label": "Seconds per 1,000 reasoning tokens (within task)",
            "unit": "seconds"
          },
          {
            "key": "withinCi",
            "label": "Within-task slope, leave-one-task-out range",
            "unit": "text"
          },
          {
            "key": "medRea",
            "label": "Median reasoning tokens",
            "unit": "tokens"
          },
          {
            "key": "medTime",
            "label": "Median total time (s)",
            "unit": "seconds"
          },
          {
            "key": "timeRange",
            "label": "Total time, lowest to highest call (s; not an interval)",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "model": "Claude Haiku 4.5 · Claude Code",
            "calls": 24,
            "rho": 0.982,
            "rhoCi": "0.98 to 0.99",
            "rhoOut": 0.986,
            "perK": 8.408,
            "withinK": 8.817,
            "withinCi": "8.6 to 9.1",
            "medRea": 4556,
            "medTime": 39.01,
            "timeRange": "15.27 to 75.13"
          },
          {
            "model": "Claude Sonnet 5.5 · Claude Code",
            "calls": 72,
            "rho": 0.921,
            "rhoCi": "0.88 to 0.94",
            "rhoOut": 0.94,
            "perK": 9.28,
            "withinK": 7.649,
            "withinCi": "7.4 to 7.8",
            "medRea": 586,
            "medTime": 7.75,
            "timeRange": "2.26 to 35.81"
          },
          {
            "model": "Claude Opus 5.5 · Claude Code",
            "calls": 80,
            "rho": 0.852,
            "rhoCi": "0.82 to 0.89",
            "rhoOut": 0.97,
            "perK": 13.091,
            "withinK": 11.754,
            "withinCi": "5.7 to 12.7",
            "medRea": 511,
            "medTime": 9.01,
            "timeRange": "3.34 to 63.00"
          },
          {
            "model": "Claude Fable 5.1 · Claude Code",
            "calls": 24,
            "rho": 0.924,
            "rhoCi": "0.89 to 0.93",
            "rhoOut": 0.941,
            "perK": 14.39,
            "withinK": 14.454,
            "withinCi": "14.4 to 22.6",
            "medRea": 889,
            "medTime": 16.13,
            "timeRange": "4.46 to 90.00"
          },
          {
            "model": "GPT-6.1 Sol · Codex CLI",
            "calls": 48,
            "rho": 0.556,
            "rhoCi": "0.34 to 0.65",
            "rhoOut": 0.85,
            "perK": 50.374,
            "withinK": 36.351,
            "withinCi": "31.8 to 36.7",
            "medRea": 189,
            "medTime": 14.51,
            "timeRange": "7.94 to 92.20"
          }
        ]
      },
      {
        "id": "thinking-bill-sum-check",
        "title": "Consistency check for reasoning within output tokens (calculation)",
        "columns": [
          {
            "key": "route",
            "label": "CLI route",
            "unit": "text"
          },
          {
            "key": "calls",
            "label": "Calls",
            "unit": "count"
          },
          {
            "key": "withReasoning",
            "label": "Calls with more than 0 reasoning tokens",
            "unit": "count"
          },
          {
            "key": "zero",
            "label": "Calls that reported 0 (kept as 0)",
            "unit": "count"
          },
          {
            "key": "over",
            "label": "Calls with reasoning above output",
            "unit": "count"
          },
          {
            "key": "slope",
            "label": "Output tokens gained per extra reasoning token (within task)",
            "unit": "ratio"
          },
          {
            "key": "slopeCi",
            "label": "Leave-one-task-out range (not a 95% interval)",
            "unit": "text"
          },
          {
            "key": "cpvisible",
            "label": "Characters per token of output minus reasoning",
            "unit": "score"
          },
          {
            "key": "cpoutZero",
            "label": "Characters per output token, calls with 0 reasoning",
            "unit": "score"
          },
          {
            "key": "cpoutWith",
            "label": "Characters per output token, calls with reasoning",
            "unit": "score"
          },
          {
            "key": "verdict",
            "label": "Reasoning is part of output",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "route": "Claude Code",
            "calls": 290,
            "withReasoning": 212,
            "zero": 78,
            "over": 0,
            "slope": 0.993,
            "slopeCi": "0.982 to 1.009",
            "cpvisible": 1.94,
            "cpoutZero": 1.98,
            "cpoutWith": 0.45,
            "verdict": "consistent with inclusion; not proof"
          },
          {
            "route": "Codex CLI",
            "calls": 88,
            "withReasoning": 68,
            "zero": 20,
            "over": 0,
            "slope": 0.967,
            "slopeCi": "0.965 to 0.986",
            "cpvisible": 2.99,
            "cpoutZero": 2.76,
            "cpoutWith": 1.36,
            "verdict": "consistent with inclusion; not proof"
          }
        ]
      }
    ],
    "related": [
      "effort-ladder",
      "hard-model-head-to-head"
    ]
  },
  "sources": [
    {
      "id": "agent-provider-h2h",
      "title": "Provider head-to-head: Claude Code models vs Codex efforts",
      "kind": "run",
      "date": "2026-10-05",
      "note": "Five short tasks with deterministic validators, declared protocol, every attempt kept.",
      "data": [
        "/benchmarks/raw/provider-h2h/receipts.json"
      ]
    },
    {
      "id": "agent-provider-h2h-hard",
      "title": "Provider head-to-head, hard set: eight hard tasks with strict validators",
      "kind": "run",
      "date": "2026-10-06",
      "note": "Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.",
      "data": [
        "/benchmarks/raw/provider-h2h-hard/receipts.json"
      ]
    },
    {
      "id": "agent-effort-ladder",
      "title": "Effort ladder: the hard task set at each effort level",
      "kind": "run",
      "date": "2026-10-06",
      "note": "The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.",
      "data": [
        "/benchmarks/raw/effort-ladder/receipts.json",
        "/benchmarks/raw/provider-h2h-hard/receipts.json"
      ]
    },
    {
      "id": "price-anthropic",
      "title": "Anthropic list prices (Claude models)",
      "kind": "price-list",
      "date": "2026-09-21",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "note": "Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price."
    },
    {
      "id": "price-openai",
      "title": "OpenAI list prices",
      "kind": "price-list",
      "date": "2026-10-03",
      "url": "https://developers.openai.com/api/docs/pricing",
      "note": "Token prices as listed by the vendor on 2026-10-03."
    },
    {
      "id": "calc-thinking-bill",
      "title": "Reasoning token bill (calculation)",
      "kind": "calculation",
      "date": "2026-10-07",
      "note": "Reported reasoning tokens priced at the recorded list prices. A calculation, not a new run.",
      "data": [
        "/benchmarks/raw/provider-h2h-hard/receipts.json",
        "/benchmarks/raw/effort-ladder/receipts.json",
        "/benchmarks/raw/provider-h2h/receipts.json"
      ]
    }
  ]
}
