Claude Code Stop hook, tested: it enforced the rules, used 1.6x the median input and lacked a fact
Claude Code Stop hook: 60/60 Sonnet code-rule checks (95% interval 94–100%). Median input was 1.6x no memory (calculation). Limits and raw data.
TL;DR
- A Stop hook is a script that Claude Code runs when the agent tries to finish. Our memory study graded 200 sessions (Sonnet 5.5, Haiku 4.5). 50 used a hook.
- It enforced the code rules: 60 of 60 checks on Sonnet (95% interval 94.0% to 100%), against 57 of 60 with no memory (86.3% to 98.3%). The intervals overlap, so there is no clear gain.
- It used more median input: 1.6x on Sonnet and 1.8x on Haiku (calculations). Token and time ranges overlap.
- It lacked the late-fee fact: with the hook alone, 0 of 3 Sonnet and 0 of 2 Haiku sessions used 1.25%. The 95% intervals are 0% to 56.2% and 0% to 65.8%.
- Hook plus an 11-line file had the top full-pass count on both models: Sonnet 15/15 (95% interval 79.6% to 100%; tied). Haiku passed 9/10 (59.6% to 98.2%) against 2/10 (5.7% to 51.0%) with no memory.
- Our recommendation: use a hook for checks you can write as code. Keep facts in a short file.
Does memory help Claude Code? 8 kinds of agent memory, tested
200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.
Transcript
- Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
- Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
- Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
- Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
- The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
- Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- Write down what the repo cannot show. Enforce what a script can check.
What a Claude Code Stop hook does
A Stop hook runs when the agent tries to finish. If its check fails, the hook exits with code 2 and prints a message. Claude Code sends the message to the agent, and the agent keeps working.
Our hook checked the agent's changes for seven kinds of problem, such as float money and a missing changelog line. It blocked at most three times. The main post covers all 8 kinds of memory we tested and links the script.
What the hook enforced
Sonnet with the hook alone passed 60 of 60 code-rule checks (94.0% to 100%; four rules in each of 15 sessions). With no memory it passed 57 of 60 (86.3% to 98.3%), and all three misses broke the money rule. The intervals overlap, so we do not call the hook ahead. The code shows these rules. Sonnet mostly followed them without a memory file. Haiku passed 40 of 40 both ways (95% interval 91.2% to 100%).
A hook can also enforce a team rule that a script can check, such as the changelog line. On the four tasks that need one, Haiku met it in 8 of 8 sessions with the hook and 0 of 8 without. The intervals are 67.6% to 100% and 0% to 32.4%, so that is a clear difference. On Sonnet it was 10 of 12 (55.2% to 95.3%) against 6 of 12 (25.4% to 74.6%). Those 95% intervals overlap.
What the hook cost in tokens and time
Sonnet used a median 125,674 input tokens per session with the hook alone and 78,455 with no memory: 1.6x (calculation). Haiku used 530,935 against 301,233: 1.8x (calculation). Most of the input is cache reads, which cost less (prompt caching, measured).
The session ranges overlap (calculation from the published session data). Sonnet ran 48,251 to 173,108 with no memory and 53,420 to 176,503 with the hook. Haiku ran 148,407 to 529,443 and 286,447 to 733,979. These are observed min–max ranges, not confidence intervals; n = 15 Sonnet and 10 Haiku sessions per condition.
Sonnet median time was 18.0 s without memory (range 12.4 to 44.8 s) and 27.3 s with the hook (14.4 to 38.3 s). Haiku medians were 54.2 s (30.2 to 66.2 s) and 68.5 s (42.3 to 94.4 s). These min–max ranges overlap; they are not confidence intervals. Samples are 15 Sonnet and 10 Haiku sessions per condition. Up to four sessions ran at once, so time is rough.
Blocks do not explain all of it. The hook blocked 5 times in 15 Sonnet sessions, so at most 5 sessions had a block. Yet 8 sit at or above the median, so at least 3 unblocked sessions used 125,674 tokens or more (calculation). We did not find the cause. No-memory sessions often stopped to ask or skipped the changelog. This test does not isolate the causes of shorter sessions.
What the hook could not do
The late-fee rate (1.25%) is in no file of the repository. This is one task: n is 3 Sonnet and 2 Haiku sessions per condition. With the hook alone, 0 of 3 Sonnet sessions used the rate: 2 asked and 1 guessed. 0 of 2 Haiku sessions used it: both guessed. No memory gave the same result. The 95% intervals for using the rate are 0% to 56.2% on Sonnet and 0% to 65.8% on Haiku. Where a memory file held the rate, Sonnet used it in 15 of 15 sessions (79.6% to 100%). That pools five file conditions, with three late-fee sessions each.
A hook checks a rule. A rate is a fact. A script cannot know the rate unless someone writes it in, and then the script is a memory file in code. That is our reading, not a measured result.
The README test command fails on Node 25. With the hook alone, Sonnet ran it in 9 of 15 sessions (35.8% to 80.2%). With no memory it ran it in 12 of 15 (54.8% to 93.0%), so the intervals overlap. Haiku ran it in 10 of 10 sessions both ways (95% interval 72.2% to 100%). The hook checks changed code, not commands. With the 11-line file, 0 of 15 Sonnet and 0 of 10 Haiku sessions ran it. Their 95% intervals are 0% to 20.4% and 0% to 27.8%.
Hook plus a short curated file
Hook alone first. Sonnet passed 12/15 in full (54.8% to 93.0%) against 9/15 with no memory (35.8% to 80.2%), and the intervals overlap. All three extra passes came from one task, the report formatting bug (3 of 3 with the hook, 0 of 3 without; calculation). Those task-level 95% intervals are 43.8% to 100% and 0% to 56.2%, so they overlap. Haiku passed 8/10 (49.0% to 94.3%) against 2/10 (5.7% to 51.0%). The intervals overlap by 2.0 points (calculation), so we do not call the hook ahead.
With the 11-line file added, Sonnet passed 15/15 (79.6% to 100%). Three other conditions also passed 15/15, so this task set hits a ceiling. Haiku passed 9/10 (59.6% to 98.2%) against 2/10 with no memory, and those intervals do not overlap. The file alone, at 7/10 (39.7% to 89.2%), overlaps both, so the hook added no clear gain over the file.
The pair used 45% less median input than the hook alone on Sonnet and 36% less on Haiku (calculations). With the pair, Sonnet input ranged from 43,396 to 157,870 tokens; Haiku input ranged from 176,184 to 490,181. Both ranges overlap the hook-only ranges. Blocks fell from 5 to 0 across 15 Sonnet sessions and from 10 to 3 across 10 Haiku sessions. The file may reduce blocks, but this test does not prove that cause.
Claude Code hooks vs CLAUDE.md, side by side
Cells read Sonnet 5.5 · Haiku 4.5. Brackets show 95% Wilson intervals for outcomes and min–max ranges for tokens. Each condition has 15 Sonnet and 10 Haiku sessions. Team-knowledge counts pool changelog checks and late-fee outcomes; they are not independent sessions.
| Result | No memory | Curated file | Hook only | Curated + hook |
|---|---|---|---|---|
| Full pass | 9/15 [35.8–80.2%] · 2/10 [5.7–51.0%] | 15/15 [79.6–100%] · 7/10 [39.7–89.2%] | 12/15 [54.8–93.0%] · 8/10 [49.0–94.3%] | 15/15 [79.6–100%] · 9/10 [59.6–98.2%] |
| Team-knowledge checks | 6/15 [19.8–64.3%] · 0/10 [0–27.8%] | 15/15 [79.6–100%] · 8/10 [49.0–94.3%] | 10/15 [41.7–84.8%] · 8/10 [49.0–94.3%] | 15/15 [79.6–100%] · 10/10 [72.2–100%] |
| Median input tokens (min–max) | 78,455 [48,251–173,108] · 301,233 [148,407–529,443] | 65,045 [42,802–136,178] · 284,212 [85,101–512,338] | 125,674 [53,420–176,503] · 530,935 [286,447–733,979] | 69,224 [43,396–157,870] · 340,491 [176,184–490,181] |
On rules the code already shows, a short file matched the hook. Only the conditions with a rate-bearing memory file used the late-fee rate.
When to use a Stop hook: our recommendation
- Use a hook for rules a script can check: banned calls, money format, generated files, a changelog line.
- Put facts in a short file: dated decisions and gotchas, such as the real test command.
- Consider both. The pair had fewer blocks than the hook alone, but no clear full-pass gain over the file alone.
- Cap the blocks. Make each message name the file, the rule and the fix.
- Test the hook on your own tasks first.
How we measured
Sources: Sonnet session data, Haiku session data and memory files and hook script.
- One small Node.js repository, five tasks, eight conditions. Two held the hook.
- Claude Code 2.1.286, headless. Sonnet ran 3 repetitions per cell (n = 15 per condition) and Haiku 2 (n = 10).
- Hidden tests and rule checks graded each session. The protocol file was created before the first counted session. Controls, /init and dreaming ran before it; the file later received a Haiku amendment and a note-count erratum. We counted every matrix attempt.
- Input tokens count uncached input, cache reads and cache writes. Intervals are Wilson 95%. A side is ahead only when intervals do not overlap.
Caveats
- The hook checks the same code rules as the grader, so treat the result as an upper bound for what this hook can enforce.
- Code-rule intervals pool four checks per session. Team-knowledge intervals also pool checks within sessions. These checks are not independent trials.
- One small synthetic repository and one author. We wrote the curated file with knowledge of the tasks.
- Most full-pass intervals overlap, so most differences are not clear.
- The agent can read the hook script, and its messages name the fix, so the hook may also act as a rule file. This test does not separate the two.
- Headless sessions cannot get an answer, so asking counts as a failure.
What to read next
- Study page, charts and raw data
- All 8 kinds of memory: Does CLAUDE.md help?
- The two models side by side: Haiku 4.5 vs Sonnet 5.5
We build Agent, which carries work and corrections across AI tools, so we have a stake in memory. Try Agent.