{"i":22,"slug":"thinking-token-bill","chart":{"id":"thinking-bill-share","title":"Reasoning share of output tokens per call on hard tasks (calculation)","subtitle":"Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)","kind":"bar","unit":"percent","yLabel":"Reasoning share of output tokens (%)","series":[{"name":"Median call","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16]]}}],"note":"Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.","whisker":"minmax","polarity":"none","sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"]}}