• Claude Code
  • Hooks
  • Stop Hook
  • Agent Memory

Claude Code Stop hook, tested: it enforced the rules, used 1.6x the median input and lacked a fact

Claude Code Stop hook: 60/60 Sonnet code-rule checks (95% interval 94–100%). Median input was 1.6x no memory (calculation). Limits and raw data.

TL;DR

  • A Stop hook is a script that Claude Code runs when the agent tries to finish. Our memory study graded 200 sessions (Sonnet 5.5, Haiku 4.5). 50 used a hook.
  • It enforced the code rules: 60 of 60 checks on Sonnet (95% interval 94.0% to 100%), against 57 of 60 with no memory (86.3% to 98.3%). The intervals overlap, so there is no clear gain.
  • It used more median input: 1.6x on Sonnet and 1.8x on Haiku (calculations). Token and time ranges overlap.
  • It lacked the late-fee fact: with the hook alone, 0 of 3 Sonnet and 0 of 2 Haiku sessions used 1.25%. The 95% intervals are 0% to 56.2% and 0% to 65.8%.
  • Hook plus an 11-line file had the top full-pass count on both models: Sonnet 15/15 (95% interval 79.6% to 100%; tied). Haiku passed 9/10 (59.6% to 98.2%) against 2/10 (5.7% to 51.0%) with no memory.
  • Our recommendation: use a hook for checks you can write as code. Keep facts in a short file.
Live story · 62 sDoes memory help Claude Code? 8 kinds of agent memory, tested

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.

Transcript
  1. Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
  2. Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  3. Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  4. With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
  5. Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
  6. The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  7. The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
  8. Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  9. Write down what the repo cannot show. Enforce what a script can check.

What a Claude Code Stop hook does

A Stop hook runs when the agent tries to finish. If its check fails, the hook exits with code 2 and prints a message. Claude Code sends the message to the agent, and the agent keeps working.

Our hook checked the agent's changes for seven kinds of problem, such as float money and a missing changelog line. It blocked at most three times. The main post covers all 8 kinds of memory we tested and links the script.

What the hook enforced

Rules the code already shows

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Rules the folders hint at

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Team knowledge only

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 3 series: Rules the code already shows, Rules the folders hint at, Team knowledge only. Rules the code already shows: highest Curated, 11 lines 100% (95% interval 94%–100%, n 60). Lowest /init CLAUDE.md 95% (95% interval 86%–98%, n 60). All intervals overlap. Rules the folders hint at: all at 100%.

NotesWhiskers: 95% Wilson intervaln 15–60 per row6 of 8 (Rules the code already shows) at 100%: this task set cannot separate them.

Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Sonnet with the hook alone passed 60 of 60 code-rule checks (94.0% to 100%; four rules in each of 15 sessions). With no memory it passed 57 of 60 (86.3% to 98.3%), and all three misses broke the money rule. The intervals overlap, so we do not call the hook ahead. The code shows these rules. Sonnet mostly followed them without a memory file. Haiku passed 40 of 40 both ways (95% interval 91.2% to 100%).

A hook can also enforce a team rule that a script can check, such as the changelog line. On the four tasks that need one, Haiku met it in 8 of 8 sessions with the hook and 0 of 8 without. The intervals are 67.6% to 100% and 0% to 32.4%, so that is a clear difference. On Sonnet it was 10 of 12 (55.2% to 95.3%) against 6 of 12 (25.4% to 74.6%). Those 95% intervals overlap.

What the hook cost in tokens and time

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows. Highest Stop hook only 125,674 (n 15). Lowest Curated, 11 lines 65,045 (n 15).

Notesn = 15 per row

Median input tokens per session, cache reads included (Claude Sonnet 5.5)

Input = uncached input + cache reads + cache writes over every turn, as the CLI reports it. Most of it is read from the prompt cache. The hook adds turns: each block sends the agent back to work.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Sonnet used a median 125,674 input tokens per session with the hook alone and 78,455 with no memory: 1.6x (calculation). Haiku used 530,935 against 301,233: 1.8x (calculation). Most of the input is cache reads, which cost less (prompt caching, measured).

The session ranges overlap (calculation from the published session data). Sonnet ran 48,251 to 173,108 with no memory and 53,420 to 176,503 with the hook. Haiku ran 148,407 to 529,443 and 286,447 to 733,979. These are observed min–max ranges, not confidence intervals; n = 15 Sonnet and 10 Haiku sessions per condition.

Time per session

Median wall time in seconds

  • Claude Sonnet 5.5
  • Claude Haiku 4.5 (square)
Sorted by gap, largest first.
No memory
/init CLAUDE.md
Stop hook only
Curated + hook
Curated, 11 lines
Handbook, 210 lines
Dreamed notes
Raw notes, 60 lines

Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: slowest Stop hook only 27.3 s (n 15). Fastest No memory 18 s (n 15). Claude Haiku 4.5: slowest Stop hook only 68.5 s (n 10). Fastest Handbook, 210 lines 49.9 s (n 10).

Notesn 10–15 per row

Up to four sessions ran at a time on one machine. Sessions without memory were often shorter because they stopped to ask or skipped the changelog.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Sonnet median time was 18.0 s without memory (range 12.4 to 44.8 s) and 27.3 s with the hook (14.4 to 38.3 s). Haiku medians were 54.2 s (30.2 to 66.2 s) and 68.5 s (42.3 to 94.4 s). These min–max ranges overlap; they are not confidence intervals. Samples are 15 Sonnet and 10 Haiku sessions per condition. Up to four sessions ran at once, so time is rough.

Blocks do not explain all of it. The hook blocked 5 times in 15 Sonnet sessions, so at most 5 sessions had a block. Yet 8 sit at or above the median, so at least 3 unblocked sessions used 125,674 tokens or more (calculation). We did not find the cause. No-memory sessions often stopped to ask or skipped the changelog. This test does not isolate the causes of shorter sessions.

What the hook could not do

  • Used the current rate (1.25%)
  • Asked for the rate, wrote no code
  • Guessed another rate
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One square per count; counts at the right are exact and in legend order.

8 rows, 3 series: Used the current rate (1.25%), Asked for the rate, wrote no code, Guessed another rate. Used the current rate (1.25%): highest Curated, 11 lines 3 (n 3). Lowest Stop hook only 0 (n 3). Asked for the rate, wrote no code: highest No memory 3 (n 3). Lowest Curated + hook 0 (n 3).

Notesn = 3 per row

The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%

Late-fee task, 3 sessions per condition. "Asked" means the agent searched the repository, found no rate and stopped with a question instead of code.

Source: Agent memory study: 8 kinds of project memory on Claude Code

The late-fee rate (1.25%) is in no file of the repository. This is one task: n is 3 Sonnet and 2 Haiku sessions per condition. With the hook alone, 0 of 3 Sonnet sessions used the rate: 2 asked and 1 guessed. 0 of 2 Haiku sessions used it: both guessed. No memory gave the same result. The 95% intervals for using the rate are 0% to 56.2% on Sonnet and 0% to 65.8% on Haiku. Where a memory file held the rate, Sonnet used it in 15 of 15 sessions (79.6% to 100%). That pools five file conditions, with three late-fee sessions each.

A hook checks a rule. A rate is a fact. A script cannot know the rate unless someone writes it in, and then the script is a memory file in code. That is our reading, not a measured result.

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest /init CLAUDE.md 87% (95% interval 62%–96%, n 15). Lowest Curated + hook 0% (95% interval 0%–20%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest No memory 100% (95% interval 72%–100%, n 10). Lowest Curated + hook 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row

Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals

The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.

Source: Agent memory study: 8 kinds of project memory on Claude Code

The README test command fails on Node 25. With the hook alone, Sonnet ran it in 9 of 15 sessions (35.8% to 80.2%). With no memory it ran it in 12 of 15 (54.8% to 93.0%), so the intervals overlap. Haiku ran it in 10 of 10 sessions both ways (95% interval 72.2% to 100%). The hook checks changed code, not commands. With the 11-line file, 0 of 15 Sonnet and 0 of 10 Haiku sessions ran it. Their 95% intervals are 0% to 20.4% and 0% to 27.8%.

Hook plus a short curated file

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest /init CLAUDE.md 60% (95% interval 36%–80%, n 15). All intervals overlap. Claude Haiku 4.5: highest Curated + hook 90% (95% interval 60%–98%, n 10). Lowest /init CLAUDE.md 20% (95% interval 5.7%–51%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row4 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Hidden tests pass and every convention check passes · 95% Wilson intervals

Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Hook alone first. Sonnet passed 12/15 in full (54.8% to 93.0%) against 9/15 with no memory (35.8% to 80.2%), and the intervals overlap. All three extra passes came from one task, the report formatting bug (3 of 3 with the hook, 0 of 3 without; calculation). Those task-level 95% intervals are 43.8% to 100% and 0% to 56.2%, so they overlap. Haiku passed 8/10 (49.0% to 94.3%) against 2/10 (5.7% to 51.0%). The intervals overlap by 2.0 points (calculation), so we do not call the hook ahead.

With the 11-line file added, Sonnet passed 15/15 (79.6% to 100%). Three other conditions also passed 15/15, so this task set hits a ceiling. Haiku passed 9/10 (59.6% to 98.2%) against 2/10 with no memory, and those intervals do not overlap. The file alone, at 7/10 (39.7% to 89.2%), overlaps both, so the hook added no clear gain over the file.

The pair used 45% less median input than the hook alone on Sonnet and 36% less on Haiku (calculations). With the pair, Sonnet input ranged from 43,396 to 157,870 tokens; Haiku input ranged from 176,184 to 490,181. Both ranges overlap the hook-only ranges. Blocks fell from 5 to 0 across 15 Sonnet sessions and from 10 to 3 across 10 Haiku sessions. The file may reduce blocks, but this test does not prove that cause.

Claude Code hooks vs CLAUDE.md, side by side

Cells read Sonnet 5.5 · Haiku 4.5. Brackets show 95% Wilson intervals for outcomes and min–max ranges for tokens. Each condition has 15 Sonnet and 10 Haiku sessions. Team-knowledge counts pool changelog checks and late-fee outcomes; they are not independent sessions.

ResultNo memoryCurated fileHook onlyCurated + hook
Full pass9/15 [35.8–80.2%] · 2/10 [5.7–51.0%]15/15 [79.6–100%] · 7/10 [39.7–89.2%]12/15 [54.8–93.0%] · 8/10 [49.0–94.3%]15/15 [79.6–100%] · 9/10 [59.6–98.2%]
Team-knowledge checks6/15 [19.8–64.3%] · 0/10 [0–27.8%]15/15 [79.6–100%] · 8/10 [49.0–94.3%]10/15 [41.7–84.8%] · 8/10 [49.0–94.3%]15/15 [79.6–100%] · 10/10 [72.2–100%]
Median input tokens (min–max)78,455 [48,251–173,108] · 301,233 [148,407–529,443]65,045 [42,802–136,178] · 284,212 [85,101–512,338]125,674 [53,420–176,503] · 530,935 [286,447–733,979]69,224 [43,396–157,870] · 340,491 [176,184–490,181]

On rules the code already shows, a short file matched the hook. Only the conditions with a rate-bearing memory file used the late-fee rate.

When to use a Stop hook: our recommendation

  1. Use a hook for rules a script can check: banned calls, money format, generated files, a changelog line.
  2. Put facts in a short file: dated decisions and gotchas, such as the real test command.
  3. Consider both. The pair had fewer blocks than the hook alone, but no clear full-pass gain over the file alone.
  4. Cap the blocks. Make each message name the file, the rule and the fix.
  5. Test the hook on your own tasks first.

How we measured

Sources: Sonnet session data, Haiku session data and memory files and hook script.

  • One small Node.js repository, five tasks, eight conditions. Two held the hook.
  • Claude Code 2.1.286, headless. Sonnet ran 3 repetitions per cell (n = 15 per condition) and Haiku 2 (n = 10).
  • Hidden tests and rule checks graded each session. The protocol file was created before the first counted session. Controls, /init and dreaming ran before it; the file later received a Haiku amendment and a note-count erratum. We counted every matrix attempt.
  • Input tokens count uncached input, cache reads and cache writes. Intervals are Wilson 95%. A side is ahead only when intervals do not overlap.

Caveats

  • The hook checks the same code rules as the grader, so treat the result as an upper bound for what this hook can enforce.
  • Code-rule intervals pool four checks per session. Team-knowledge intervals also pool checks within sessions. These checks are not independent trials.
  • One small synthetic repository and one author. We wrote the curated file with knowledge of the tasks.
  • Most full-pass intervals overlap, so most differences are not clear.
  • The agent can read the hook script, and its messages name the fix, so the hook may also act as a rule file. This test does not separate the two.
  • Headless sessions cannot get an answer, so asking counts as a failure.

We build Agent, which carries work and corrections across AI tools, so we have a stake in memory. Try Agent.

The data behind this post

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.