{"i":22,"slug":"thinking-token-bill","chart":{"id":"thinking-bill-short-vs-hard","title":"Reasoning share on short tasks vs hard tasks (calculation)","subtitle":"Median call per configuration; five short tasks and eight hard tasks","kind":"grouped-bar","unit":"percent","yLabel":"Reasoning share of output tokens (median call, %)","series":[{"name":"Eight hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16]]}},{"name":"Five short tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",90.19,73.1,97.59,15],["Claude Sonnet 5.5 · Claude Code",0,0,72.75,15],["Claude Opus 5.5 · Claude Code",0,0,93.33,15],["Claude Opus 5.5 (high) · Claude Code",43.59,0,93.33,15],["Claude Fable 5.1 · Claude Code",0,0,74.01,15],["GPT-6.1 Sol (medium) · Codex CLI",41.05,0,71.43,15],["GPT-6.1 Sol (high) · Codex CLI",58.06,0,75.76,15]]}}],"note":"Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.","whisker":"minmax","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}}