Explainer · Claude Code hook
Claude Code hooks, explained with a measured Stop hook
Definition
A Claude Code hook runs a command at an event, such as a tool call or session stop. Our Stop hook blocked finishing and sent standard-error feedback with exit code 2.
Agent team · · 5 min read · Every number is from the public studies
- Claude Sonnet 5.5
- Claude Haiku 4.5
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | 95% interval | n |
|---|---|---|---|---|
| No memory | 60% | 20% | Claude Sonnet 5.5: 36%–80%; Claude Haiku 4.5: 5.7%–51% | 15 |
| /init CLAUDE.md | 60% | 20% | Claude Sonnet 5.5: 36%–80%; Claude Haiku 4.5: 5.7%–51% | 15 |
| Curated, 11 lines | 100% | 70% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 40%–89% | 15 |
| Raw notes, 60 lines | 93% | 60% | Claude Sonnet 5.5: 70%–99%; Claude Haiku 4.5: 31%–83% | 15 |
| Dreamed notes | 100% | 70% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 40%–89% | 15 |
| Handbook, 210 lines | 100% | 30% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 11%–60% | 15 |
| Stop hook only | 80% | 80% | Claude Sonnet 5.5: 55%–93%; Claude Haiku 4.5: 49%–94% | 15 |
| Curated + hook | 100% | 90% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 60%–98% | 15 |
8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest /init CLAUDE.md 60% (95% interval 36%–80%, n 15). All intervals overlap. Claude Haiku 4.5: highest Curated + hook 90% (95% interval 60%–98%, n 10). Lowest /init CLAUDE.md 20% (95% interval 5.7%–51%, n 10). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 10–15 per row4 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.
Hidden tests pass and every convention check passes · 95% Wilson intervals
Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.
Source: Agent memory study: 8 kinds of project memory on Claude Code
Hook sessions passed every code-rule check. The hook did not supply the late-fee rate. Median input rose; ranges overlapped no-memory sessions.
How a Claude Code hook works
Set an event and command in .claude/settings.json. PreToolUse runs before a tool call; Stop runs when Claude finishes. Exit-code behavior depends on the event.
We tested only Stop. The published hook files include this settings file:
{
"hooks": {
"Stop": [
{
"hooks": [
{ "type": "command", "command": "node \"$CLAUDE_PROJECT_DIR/.claude/hooks/check-conventions.mjs\"" }
]
}
]
}
}The script checked float money, plain errors, system-clock use, console output, edited migrations, stale generated files and missing changelog lines. It printed failures and fixes and exited with code 2. After 3 blocks, it allowed stopping.
Our agent memory study graded 200 headless Claude Code CLI sessions (version 2.1.286) on one small repository. Per condition: n = 15 for Sonnet 5.5; n = 10 for Haiku 4.5.
What a Stop hook enforced, and what it missed
The hook and grader checked the same four code-rule categories with different file filters and patterns. Grades do not prove the hook caused compliance. We tested one hook on one repository.
On Sonnet 5.5, sessions with the hook alone passed 60 of 60 code-rule checks (95% interval 94.0% to 100%). With no memory, Sonnet passed 57 of 60 (86.3% to 98.3%). The intervals overlap: no clear lead. On Haiku 4.5, both passed 40 of 40 (95% interval 91.2% to 100%). The code rules hit a ceiling there.
Totals pool checks from 15 Sonnet and 10 Haiku sessions. Within-session checks are not independent.
Team knowledge covers the changelog and late-fee rate. All intervals are 95% Wilson intervals. Pooled results count checks; full passes count sessions.
Sonnet with the hook alone passed 10 of 15 team-knowledge checks (41.7% to 84.8%) against 6 of 15 with no memory (19.8% to 64.3%). Those intervals overlap: the gap is unclear. Haiku 4.5 passed 8 of 10 (49.0% to 94.3%) against 0 of 10 (0% to 27.8%). Those intervals do not overlap: the hook is ahead on Haiku here.
Split-check intervals are computed from the pass and fail outcomes in the Sonnet receipts and Haiku receipts. With the hook, Haiku passed the changelog check 8 of 8 times (67.6% to 100%) against 0 of 8 with no memory (0% to 32.4%). On Sonnet it was 10 of 12 (55.2% to 95.3%) against 6 of 12 (25.4% to 74.6%), and those overlap. Both Sonnet misses asked and wrote no code; the hook had nothing to check.
The script did not supply the late-fee rate (1.25%). Hooks could supply facts; we did not test that. With the hook alone, 0 of 3 Sonnet sessions used it (0% to 56.2%; direct receipt calculation). Two asked and wrote no code; one guessed. Across five memory-file conditions holding the rate, Sonnet used it 15 of 15 times (79.6% to 100%). Each had three late-fee sessions.
Unanswered headless questions count as failures. A person could answer live.
What the hook cost
A block adds turns. On Sonnet, median input per session was 125,674 tokens with the hook alone and 78,455 with no memory: 1.6x (calculation). On Haiku it was 530,935 against 301,233: 1.8x (receipt calculation). The hook blocked 5 times in 15 Sonnet sessions and 10 times in 10 Haiku sessions.
The 10 unblocked Sonnet hook sessions still had a median 126,604 input tokens (range 53,420 to 148,139; receipt calculation). We did not isolate why.
Input sums uncached input, cache reads and writes across turns (CLI-counter calculation). Most came from the prompt cache, including CLI context.
Session input ranges overlap; they are not confidence intervals. Sonnet: 53,420 to 176,503 with the hook, 48,251 to 173,108 with none (n = 15 each). Haiku: 286,447 to 733,979 with the hook, 148,407 to 529,443 with none (n = 10 each; range calculation from receipts). These ratios describe this run.
On Sonnet (n = 15), median time was 27.3 s with the hook alone (range 14.4 s to 38.3 s). With no memory it was 18.0 s (range 12.4 s to 44.8 s). On Haiku (n = 10) it was 68.5 s (range 42.3 s to 94.4 s) against 54.2 s (range 30.2 s to 66.2 s). Ranges overlap and are not confidence intervals; medians and ranges are linked-receipt calculations. Up to four sessions shared one machine. No-memory sessions often stopped to ask or skipped the changelog.
A hook plus a short memory file
A full pass requires passing hidden tests and every applicable check. Table intervals are direct receipt calculations (95% Wilson, z = 1.96).
| Condition | Sonnet 5.5 (n 15) | Haiku 4.5 (n 10) |
|---|---|---|
| No memory | 9/15 (35.7% to 80.2%) | 2/10 (5.7% to 51.0%) |
| Stop hook only | 12/15 (54.8% to 93.0%) | 8/10 (49.0% to 94.3%) |
| 11-line file plus hook | 15/15 (79.6% to 100%) | 9/10 (59.6% to 98.2%) |
On Sonnet, all interval pairs overlap: no condition is ahead. The 11-line file alone also passed 15 of 15 (95% interval 79.6% to 100%): a ceiling for this task set.
On Haiku, the hook alone and no memory overlap (49.0% to 51.0%), so that gap is unclear. The file plus hook leads no memory: intervals do not overlap. The file alone (7 of 10; 39.7% to 89.2%) overlaps both, so we cannot claim the hook added to the file.
How to use hooks well
These suggestions concern one synthetic repository. The curated file’s author knew the tasks. We did not compare CLAUDE.md, skills and hooks as rule locations.
- Put checkable rules in a hook, such as a changelog line.
- Put facts in a short memory file, such as the late-fee rate.
- Keep the hook message short: the file, the rule and the fix. Our reading; we did not vary message length.
- Test on a few real tasks with and without the hook, and compare pass rate, time and input tokens.
Frequently asked questions
What is a Stop hook in Claude Code?
A Stop hook runs when Claude finishes. In the tested version, exit code 2 blocked finishing and sent standard-error feedback to Claude.
Do hooks replace CLAUDE.md?
Our hook checked code; it did not supply the late-fee rate. A hook could supply facts.
Do hooks cost tokens?
Feedback can add turns and tokens. Sonnet’s median input ratio was 1.6x (calculation; n = 15 per condition). The ranges above overlap. We did not isolate the gap’s cause.
When should I use a hook?
Use one for a rule that a script can check, and keep facts in a short file. We tested one Stop hook on one repository.