{"method":["This study is a calculation over receipts that already exist. We made no new model call. The inputs are three recorded runs. The hard head-to-head gives 152 completed calls. We leave its 30 pre-inference blocked attempts out of token calculations. They remain failures of the original route attempt, not model answers. The effort ladder gives 96 new calls. The five-task head-to-head gives 130 calls. All current counted calls completed. Recorded errors count as failures when their usage is known. We keep the separate blocked batch on record. The original Codex route was blocked. A later batch ran after a successful uncounted probe. It resumed after two completed calls and skipped them.","Sum check. Before we compute any share, we test one point. Do the output tokens include the reasoning tokens? We use this accounting assumption and check its consistency with the receipts. We run three tests on each CLI route. Test 1: reasoning never exceeds output. Test 2: within one task and model, one more reasoning token adds about one output token. A slope near 1 supports the assumption. Correlation alone cannot prove the counter semantics. Test 3: we compare characters per remaining token with characters per output token on calls whose reasoning counter is zero. Claude Code: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 290 calls). Within one task and model, each extra reasoning token went with 0.99 extra output tokens (leave-one-task-out range 0.98 to 1.01, 290 calls). Output minus reasoning has 1.94 characters per token. Calls with 0 reasoning have 1.98. For comparison, output including reasoning has only 0.45. Codex CLI: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 88 calls). Within one task and model, each extra reasoning token went with 0.97 extra output tokens (leave-one-task-out range 0.97 to 0.99, 88 calls). Output minus reasoning has 2.99 characters per token. Calls with 0 reasoning have 2.76. For comparison, output including reasoning has only 1.36.","Not reported is not zero. A route reports reasoning when at least one of its calls reports more than 0. Both routes do (Claude Code 212 of 290 calls, Codex CLI 68 of 88 calls). So a reported 0 is the recorded counter, not proof of no internal reasoning. 98 calls reported 0, and their replies look like visible text only (test 3). If any model attempt lacks usable token counts, this builder returns no study. It never prices unknown usage as zero. 0 calls had no usable count.","Reasoning share of one call = reasoning tokens ÷ output tokens. A configuration gets two numbers. One is the median of its per-call shares. The other is the pooled share (all reasoning tokens ÷ all output tokens). Ranges are the lowest and highest call. They are not intervals.","Hard tasks: we use all calls of each configuration in the hard head-to-head. That is 16 to 24 calls: 8 tasks × 3 repetitions for Claude Code and 8 × 2 for Codex CLI. Effort ladder: 11 cells of 8 tasks × 2 repetitions, the design of the effort-ladder study. We reuse its reference cells from the hard head-to-head (Claude repetitions 1-2 only). Short tasks: the five validated tasks (10 to 15 calls per configuration).","Cost is a calculation at list price, not a bill. The calls ran on flat subscriptions. Reasoning cost = reasoning tokens × the model's output price. Remaining output cost = (output tokens − reasoning tokens) × the same price. We use the rest as a visible-answer estimate. Input cost covers the whole prompt. Cache writes use the one-hour list-price assumption; the receipts do not state the cache lifetime. Prices per million output tokens: Claude Haiku 4.5 $5, Claude Sonnet 5.5 $10, Claude Opus 5.5 $20, Claude Fable 5.1 $50, GPT-6.1 Sol $10. Sources: Anthropic list prices of 2026-09-21, OpenAI of 2026-10-03.","Cost per strict pass = the list-price cost of all calls in the cell, failures included, ÷ the cell's strict passes. Hard and effort cells use the hard-set strict rule. Short-task cells use their original rule, which can strip a wrapping fence.","Time link: we use all calls of the hard head-to-head and the effort ladder. For each model and route, we compute the Spearman rank correlation of reasoning tokens with total time (with a leave-one-task-out sensitivity range, not a confidence interval). We also compute the least-squares slope in seconds per 1,000 reasoning tokens. Then we compute the slope within task. We centre each task on its own mean. This controls for differences in task means, not effort or batch effects.","Isolation, tasks, validators and flags follow the hard head-to-head and the five-task head-to-head. Each call ran in a fresh empty folder with tools off and one turn. The timeout was 300 s per call (180 s for the short tasks). One call ran at a time per account.","Claude Code calls ran with an output cap of 16,000 tokens in all three runs. The highest output of any analysed Claude call was 9,321, so no call hit the cap. The Codex CLI has no cap setting. Its highest output was 2,569."]}