• Claude Haiku
  • Extended Thinking
  • Routing
  • Hard Tasks
  • Latency
  • Cost

Does thinking pay for Claude Haiku 4.5? Thinking on vs off

Does extended thinking pay for Claude Haiku 4.5: what does it buy in accuracy, and what does it cost in time and money, on typed routing decisions and on hard tasks?

Published · 6 charts · Download the data or a carousel

87%

95% CI 78%–92% · n = 82

71/82 · Claude Haiku 4.5 (thinking off): exact routing decisions

89%

95% CI 80%–94% · n = 82

73/82 · Claude Haiku 4.5 (thinking on): exact routing decisions

0.754 · Paired exact test, thinking off vs on (exact McNemar p) · n = 82

p = 0.754 (6 only on, 4 only off, 82 cases)

The answer

Routing accuracy is unresolved on this sample of 82 paired decisions. Thinking off answered 87% (71/82; 95% interval 78% to 92%) exactly. Thinking on answered 89% (73/82; 95% interval 80% to 94%) exactly. Only thinking on was right in 6 cases; only thinking off in 4. Both were right in 67 cases and wrong in 5. The exact McNemar test gives p = 0.754. The test does not show a difference. This does not establish equal accuracy. The two 95% intervals overlap. Observed median wall time: 4.66 s with thinking off and 12.54 s with thinking on. The thinking-on median is 2.7 times the thinking-off median (a calculation). Wall-time ranges: 2.20 s to 11.57 s off; 5.86 s to 51.28 s on. These ranges overlap and are not confidence intervals. The 95th percentiles are 8.18 s off and 34.48 s on. Mean thinking tokens per decision: 0 off and 1,101 on. List-price cost per 1,000 decisions: $3.36 off and $8.92 on (a calculation). The recorded Sonnet 5.5 low-effort reference answered 94% (77/82; 95% interval 87% to 97%) exactly. Its median wall time was 2.60 s, with range 1.99 s to 5.58 s (not an interval). Only Sonnet was right in 7 cases; only thinking-off Haiku in 1. Exact McNemar p = 0.07 (secondary test). Hard tasks: 8 tasks, each repeated 3 times per arm. Thinking off passed 17% (4/24; 95% interval 7% to 36%) strictly. Thinking on passed 46% (11/24; 95% interval 28% to 65%) strictly. Format misses: 0 off and 5 on. These are right answers in the wrong wrapping, never strict passes. The strict 95% intervals overlap, so strict pass rate does not separate the arms. Lenient passes count format misses: 17% (4/24; 95% interval 7% to 36%) off; 67% (16/24; 95% interval 47% to 82%) on. The lenient intervals do not overlap: thinking on is ahead on this reading. Median total time per call: 2.9 s off and 39.0 s on. Single-call ranges: 1.7 s to 13.0 s off; 15.3 s to 75.1 s on. These ranges do not overlap and are not confidence intervals. List-price cost per strict pass: $0.0365 off and $0.0672 on (a calculation). The calculation divides all call costs by 4 and 11 strict passes. It has no interval. Thinking off was cheaper per strict pass in this sample, but it passed fewer calls. Both thinking-on arms reuse earlier recorded sessions, so the timing comparison is not concurrent.

Key numbers

91% (177/194)

Claude Haiku 4.5 (thinking off): per-question routing accuracy

95% CI 86%–94% · n = 194

94% (183/194)

Claude Haiku 4.5 (thinking on): per-question routing accuracy

95% CI 90%–97% · n = 194

4.66s

Claude Haiku 4.5 (thinking off): median wall time per routing decision

n = 82

12.54s

Claude Haiku 4.5 (thinking on): median wall time per routing decision

n = 82

2.7x

Median wall time, thinking on ÷ thinking off (calculation)

(12.54 s ÷ 4.66 s) · n = 82

0

Claude Haiku 4.5 (thinking off): thinking tokens per routing decision (mean)

n = 82

1,101

Claude Haiku 4.5 (thinking on): thinking tokens per routing decision (mean)

(range 285 to 4,443) · n = 82

$3.364

Claude Haiku 4.5 (thinking off): list-price cost per 1,000 routing decisions (calculation)

n = 82

$8.924

Claude Haiku 4.5 (thinking on): list-price cost per 1,000 routing decisions (calculation)

n = 82

17% (4/24)

Claude Haiku 4.5 (thinking off): strict passes on hard tasks

95% CI 6.7%–36% · n = 24

46% (11/24)

Claude Haiku 4.5 (thinking on): strict passes on hard tasks

95% CI 28%–65% · n = 24

2.9s

Claude Haiku 4.5 (thinking off): median total time per hard-task call

n = 24

39.0s

Claude Haiku 4.5 (thinking on): median total time per hard-task call

n = 24

4,556

Claude Haiku 4.5 (thinking on): median reasoning tokens per hard-task call

n = 24

$0.0365

Claude Haiku 4.5 (thinking off): list-price cost per strict pass on hard tasks (calculation)

n = 24

$0.0672

Claude Haiku 4.5 (thinking on): list-price cost per strict pass on hard tasks (calculation)

n = 24

106

Counted model calls made for this study (thinking off)

(82 routing, 24 hard tasks; 2 more uncounted probes) · n = 106

0 of 106

Thinking-off calls that reported any thinking tokens

n = 106

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

  • Exact decisions (every scored question right)
  • Per-question accuracy
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

3 rows, 2 series: Exact decisions (every scored question right), Per-question accuracy. Exact decisions (every scored question right): highest Claude Sonnet 5.5 (low) · Claude Code 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 87% (95% interval 78%–92%, n 82). All intervals overlap. Per-question accuracy: highest Claude Sonnet 5.5 (low) · Claude Code 97% (95% interval 94%–99%, n 194). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 91% (95% interval 86%–94%, n 194). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 82–194 per row

Claude Haiku 4.5 and Claude Sonnet 5.5 (low effort); 82 decisions, the same cases for every arm

Whiskers are 95% Wilson intervals. Questions within a decision are related; per-question intervals are descriptive, not an independent-question test. The thinking-on and Sonnet arms are the recorded 2026-10-05 routing run, reused, not rerun; the thinking-off arm ran later on another account. An unanswered question counts as wrong.

Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing

Share card (PNG)
  • Wall time (CLI)
  • Model time (API)
Entrance: medians race at 9× real timeMotion reduced: press Replay to animateThe slowest median is 12.5 s. The clock runs at the recorded speed.
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

3 rows, 2 series: Wall time (CLI), Model time (API). Wall time (CLI): slowest Claude Haiku 4.5 (thinking on) · Claude Code 12.5 s (median to p95 12.5 s–34.5 s, n 82). Fastest Claude Sonnet 5.5 (low) · Claude Code 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap. Model time (API): slowest Claude Haiku 4.5 (thinking on) · Claude Code 10.5 s (median to p95 10.5 s–32.1 s, n 82). Fastest Claude Sonnet 5.5 (low) · Claude Code 1.6 s (median to p95 1.6 s–2.6 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n = 82 per row

Median wall time and model (API) time; whisker to the 95th percentile

Whiskers run from p50 to p95, not a confidence interval. Percentiles use the nearest-rank rule of the routing-overhead study, so the thinking-on and Sonnet medians match that study (the Jev-vs-LLM page interpolates between ranks and shows slightly different values). One call at a time through the Claude Code CLI; the thinking-on and Sonnet arms ran on another day.

Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing

Share card (PNG)
  • Thinking tokens
  • Visible output tokens
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

3 rows, 2 series: Thinking tokens, Visible output tokens. Thinking tokens: highest Claude Haiku 4.5 (thinking on) · Claude Code 1,101 (n 82). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 0 (n 82). Visible output tokens: highest Claude Haiku 4.5 (thinking off) · Claude Code 366 (n 82). Lowest Claude Sonnet 5.5 (low) · Claude Code 105 (n 82).

Notesn = 82 per row

Mean per decision, as the Claude Code CLI reports them

Visible output = output tokens minus thinking tokens. It includes the structured answer the CLI asks for. Thinking tokens are counted by the CLI; their content is never captured. More tokens is not better or worse by itself.

Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing

Share card (PNG)
Calculation
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Hover or focus a bar for its ratio to Claude Haiku 4.5 (thinking of… (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 (thinking on) · Claude Code $8.92 (n 82). Lowest Claude Haiku 4.5 (thinking off) · Claude Code $3.36 (n 82).

Notesn = 82 per row

Reported tokens × list price, per 1,000 decisions

Calculation, not a bill: the calls ran on a flat subscription. Reported input, cache and output tokens (thinking tokens are part of output) × list price, with the price table of the Jev-vs-LLM study. Sonnet cache writes use the one-hour rate ($4 per million tokens), as recorded in its receipts. The calculation matches the CLI-reported cost.

Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)

Share card (PNG)
  • Strict pass
  • Lenient (format misses counted)
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code

2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Haiku 4.5 (thinking on) · Claude Code 46% (95% interval 28%–65%, n 24). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 17% (95% interval 6.7%–36%, n 24). All intervals overlap. Lenient (format misses counted): highest Claude Haiku 4.5 (thinking on) · Claude Code 67% (95% interval 47%–82%, n 24). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 17% (95% interval 6.7%–36%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 24 per row

Claude Haiku 4.5 in Claude Code; strict and lenient

Whiskers are 95% Wilson intervals over calls. Three repeats per task are related, so these intervals do not measure uncertainty across unseen tasks. A format miss (right answer in the wrong wrapping) never counts as a strict pass. The thinking-on cell is the 24 recorded Haiku receipts of the hard head-to-head, reused; the thinking-off cell ran later on another account.

Sources: Haiku thinking on vs off, Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
Entrance: medians race at 28× real time
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

2 rows. Slowest Claude Haiku 4.5 (thinking on) · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Haiku 4.5 (thinking off) · Claude Code 3 s (range 1.7 s–13 s, n 24). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 24 per row

Median per cell; whiskers = fastest and slowest call

Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network; the thinking-on calls ran in an earlier session. Timings include CLI start-up.

Sources: Haiku thinking on vs off, Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)

Tables

Same cases, two arms: paired test on exact decisions

Pair (first arm = thinking off)CasesBoth rightOnly thinking off rightOnly the other arm rightBoth wrongExact McNemar p (two-sided)
Thinking off vs thinking on (the same Haiku 4.5)82674650.754
Thinking off vs Sonnet 5.5 at low effort (secondary)82701740.07

Routing: every arm

ArmThinking settingCalls (errors)ExactPer-questionWall p50 / p95 (s)API p50 / p95 (s)Wall min to max (s; not an interval)API min to max (s; not an interval)Thinking tokens per decision (range)Visible output tokensUSD per 1,000 (calculation)
Claude Haiku 4.5 (thinking off) · Claude CodeMAX_THINKING_TOKENS=082 (0)87% (71/82; 95% interval 78% to 92%)91% (177/194; 95% interval 86% to 94%)4.66 / 8.183.79 / 7.432.2 to 11.571.37 to 10.830 (0 to 0)366$3.36
Claude Haiku 4.5 (thinking on) · Claude CodeCLI default82 (0)89% (73/82; 95% interval 80% to 94%)94% (183/194; 95% interval 90% to 97%)12.54 / 34.4810.51 / 32.135.86 to 51.284.33 to 49.491101 (285 to 4443)318$8.92
Claude Sonnet 5.5 (low) · Claude Code--effort low82 (0)94% (77/82; 95% interval 87% to 97%)97% (189/194; 95% interval 94% to 99%)2.6 / 4.31.6 / 2.581.99 to 5.581.06 to 4.762 (0 to 63)105$7.32

Routing: exact decisions by decision type

Decision type (cases)Claude Haiku 4.5 (thinking off) · Claude CodeClaude Haiku 4.5 (thinking on) · Claude CodeClaude Sonnet 5.5 (low) · Claude Code
Failure class (18)83% (15/18; 95% interval 61% to 94%)94% (17/18; 95% interval 74% to 99%)100% (18/18; 95% interval 82% to 100%)
Message intent (20)100% (20/20; 95% interval 84% to 100%)100% (20/20; 95% interval 84% to 100%)100% (20/20; 95% interval 84% to 100%)
Is it a rule? (12)100% (12/12; 95% interval 76% to 100%)100% (12/12; 95% interval 76% to 100%)100% (12/12; 95% interval 76% to 100%)
Context shape (32)75% (24/32; 95% interval 58% to 87%)75% (24/32; 95% interval 58% to 87%)84% (27/32; 95% interval 68% to 93%)

Hard tasks: both cells

CellCallsStrict passFormat missesWrong answersErrors or incomplete calls (failures)Total time median (min to max, s)Reasoning tokens medianOutput tokens medianUSD per strict pass (calculation)
Claude Haiku 4.5 (thinking off) · Claude Code2417% (4/24; 95% interval 7% to 36%)02002.9 (1.7 to 13)0291$0.037
Claude Haiku 4.5 (thinking on) · Claude Code2446% (11/24; 95% interval 28% to 65%)58039 (15.3 to 75.1)4,5565,064$0.067

Hard tasks: strict passes per task

TaskClaude Haiku 4.5 (thinking off) · Claude CodeClaude Haiku 4.5 (thinking on) · Claude Code
Fix an interval-merge function (off-by-one and edge cases)100% (3/3; 95% interval 44% to 100%)100% (3/3; 95% interval 44% to 100%)
Fix a time-zone day-length function (DST)0% (0/3; 95% interval 0% to 56%)33% (1/3; 95% interval 6% to 79%)
Write a CSV parser (quoted newlines, strict errors)0% (0/3; 95% interval 0% to 56%)67% (2/3; 95% interval 21% to 94%) (+1 format miss)
Predict JavaScript event-loop output order0% (0/3; 95% interval 0% to 56%)0% (0/3; 95% interval 0% to 56%)
Solve a multi-constraint room schedule0% (0/3; 95% interval 0% to 56%)0% (0/3; 95% interval 0% to 56%) (+2 format misses)
Write a strict SemVer 2.0.0 regex0% (0/3; 95% interval 0% to 56%)100% (3/3; 95% interval 44% to 100%)
Refactor to remove duplication, keep 20 tests green33% (1/3; 95% interval 6% to 79%)67% (2/3; 95% interval 21% to 94%)
Write a SQLite reporting query (fan-out, ties, boundaries)0% (0/3; 95% interval 0% to 56%)0% (0/3; 95% interval 0% to 56%) (+2 format misses)

Method

  1. Routing uses the platform’s labelled decision suites: Failure class (failure-class-2026-10-05.1, 18 cases); Message intent (frontdoor-intent-2026-10-04.2, 20 cases); Is it a rule? (memory-is-rule-2026-10-04.2, 12 cases); Context shape (context-shape-2026-10-05.1, 32 cases).
  2. Production rules select the cases and leave the scored keys open. This gives 82 decisions and 194 scored questions.
  3. Each decision uses one Claude Code CLI call. Calls use JSON output, tools off, no MCP and no saved session. They use the production system text, prompt and JSON schema. Calls run one at a time; temperature cannot be set.
  4. The decision-eval runner scores each reply. Exact means every scored question is acceptable. Per-question accuracy counts each scored question. Errors and unanswered questions count as wrong.
  5. Thinking off sets MAX_THINKING_TOKENS=0 on the claude process. Two uncounted probes check the setting, one per route. The probes and all new counted calls reported zero thinking tokens.
  6. Thinking off uses new calls. Thinking on uses the CLI default in the recorded routing run. Sonnet 5.5 uses low effort in that recorded run. Both reference arms are reused; they were not rerun.
  7. The exact two-sided McNemar test uses case-level pairs that differ (binomial, p = 0.5). Rate intervals are 95% Wilson intervals.
  8. Timing uses wall time and the CLI’s reported API time. The same code computes nearest-rank p50 and p95 from each arm’s call log. Cost uses reported input, cache and output tokens times list price (a calculation).
  9. Hard tasks use the 8 unchanged tasks and sandboxed deterministic validators from the hard head-to-head.
  10. Controls ran before the first new model call, but before protocol creation. All 8 reference answers passed (8 tested). All 26 planted wrong answers failed. All 8 wrapped references were flagged as format misses.
  11. Strict pass means the whole reply passes. A format miss means a lenient extractor finds a passing answer; it never counts as a strict pass. Errors and incomplete attempts count as failures.
  12. Thinking off used 24 calls: 8 tasks, three repetitions each. Calls ran one at a time, with no effort flag, tools off, no MCP and one turn. The timeout was 300 s; the output-token cap was 16,000. Thinking on reuses 24 recorded default Haiku receipts.
  13. Cost per strict pass divides all call costs in a cell by its strict passes. Each call cost uses list price times reported tokens (a calculation).
  14. Recorded order: controls at 2026-10-06T21:27:56.463Z; protocol creation at 2026-10-06T21:29:01.607477Z; then the two probes. The first new counted call started at 2026-10-06T21:29:57.270Z. Reused reference calls predate both the controls and this protocol.
  15. Declared caps: 106 new counted calls and 112 calls including probes. Observed: 106 counted calls plus 2 probes; the total stayed within both call caps.
  16. Attempts: 106 counted calls, 2 uncounted probes. Nothing was trimmed or retried, and no run hit a usage limit. Every attempt is in the published extract.

Caveats

  • The thinking-on and Sonnet arms are the recorded routing run of 2026-10-05, reused. The thinking-off arm ran on a different day and on a different Claude subscription account, so this is a confounded comparison. Day, account, CLI version and input-token differences can affect the results. The data cannot isolate the effect of thinking.
  • 82 decisions in four small hand-labelled sets: intervals are wide, and a paired test with few discordant cases has little power. The case sets and question wording were tuned in fix waves against Jev answers (2026-10-04 to 2026-10-05).
  • Thinking off here means MAX_THINKING_TOKENS=0 on the Claude Code process, checked by the CLI’s own thinking-token counters. It says nothing about the API’s thinking settings or about other models.
  • Claude Code adds its own start-up time and tool-schema tokens to every call; a direct API router would skip them. CLI timings include that overhead.
  • List-price costs are calculations from reported tokens; the calls used a flat subscription. Sonnet cache writes use the one-hour rate ($4 per million tokens), as recorded in its receipts. The calculation matches the CLI-reported cost.
  • The Haiku routing arms have the same prompt character count for every case, but thinking-on calls report about 297 more input tokens per call (1,831 against 1,534, a calculation from the two means). The cause is unknown, and the CLI version of the thinking-on run is not recorded. At Haiku’s list price that is about $0.30 of the $5.56 cost gap per 1,000 decisions (a calculation).
  • Hard tasks: 24 calls per cell (3 per task); a 4/24 result has a 95% interval of 7% to 36%. The thinking-on cell is the recorded hard head-to-head receipts (batch of 2026-10-06, effort flag not passed), reused; the thinking-off cell ran later, on another account.
  • These are eight distinct hard tasks, repeated three times each. Repeats on the same task are related. Call-level Wilson intervals are descriptive; they do not measure uncertainty across unseen tasks.
  • Strict format rules decide part of the hard result: a reply in a code fence fails strictly. The lenient reading is shown next to it so the two can be told apart.
  • Cost per strict pass divides a cell’s cost by its strict passes (4 and 11). It has no interval, and the strict pass-rate intervals overlap, so the order of the two costs per pass is not settled.
  • Both arms hit a strict-pass ceiling on Fix an interval-merge function (off-by-one and edge cases): 100% (3/3; 95% interval 44% to 100%) in each arm. Repeats on these tasks cannot establish equal accuracy on harder variants.
  • Controls and reused calls predate protocol creation. The file predates the new probes and counted calls, but later amendments changed it. No frozen initial copy verifies the original wording.
  • The run folder has no saved pre-batch usage-gate readings. Later corroboration does not prove that each pre-batch gate ran.
  • The runs used a shared Mac. Host load was not recorded; the run does not prove isolation. Timing gaps do not establish a cause.
  • Routing has a ceiling on these subsets: Message intent, 100% (20/20; 95% interval 84% to 100%) in every arm; Is it a rule?, 100% (12/12; 95% interval 76% to 100%) in every arm. These results cannot establish equal accuracy on harder cases.

Sources

  • Haiku thinking on vs off

    Our recorded runs ·

    Receipts of the Haiku thinking study: Claude Haiku 4.5 through Claude Code with thinking off (MAX_THINKING_TOKENS=0) against the recorded thinking-on arms. Routing arms carry per-arm totals computed from each arm’s per-call log and per-case results; hard-task receipts are every attempt of the thinking-off run. Every attempt is kept, failures included. The thinking-on hard-task receipts live in the hard head-to-head extract.

    Raw data: haiku-thinking/receipts.json

  • Routing runs: Jev router vs LLM routing

    Our recorded runs ·

    Routing decisions recorded per case and arm.

    Raw data: routing/receipts.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • Provider head-to-head, hard set: eight hard tasks with strict validators

    Our recorded runs ·

    Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.

    Raw data: provider-h2h-hard/receipts.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Does thinking pay for Claude Haiku 4.5? Thinking on vs off”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/haiku-thinking-on-off.

Models and comparisons in this study

More studies

All benchmarks
Live story
  • Head to head
  • Hard Tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.

91% (139/152)Calls that passed strictly (hard set) · n = 152

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.