Explainer · Claude Code hook

Claude Code hooks, explained with a measured Stop hook

Definition

A Claude Code hook runs a command at an event, such as a tool call or session stop. Our Stop hook blocked finishing and sent standard-error feedback with exit code 2.

Agent team · · 5 min read · Every number is from the public studies

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest /init CLAUDE.md 60% (95% interval 36%–80%, n 15). All intervals overlap. Claude Haiku 4.5: highest Curated + hook 90% (95% interval 60%–98%, n 10). Lowest /init CLAUDE.md 20% (95% interval 5.7%–51%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row4 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Hidden tests pass and every convention check passes · 95% Wilson intervals

Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Hook sessions passed every code-rule check. The hook did not supply the late-fee rate. Median input rose; ranges overlapped no-memory sessions.

How a Claude Code hook works

Set an event and command in .claude/settings.json. PreToolUse runs before a tool call; Stop runs when Claude finishes. Exit-code behavior depends on the event.

We tested only Stop. The published hook files include this settings file:

{
  "hooks": {
    "Stop": [
      {
        "hooks": [
          { "type": "command", "command": "node \"$CLAUDE_PROJECT_DIR/.claude/hooks/check-conventions.mjs\"" }
        ]
      }
    ]
  }
}

The script checked float money, plain errors, system-clock use, console output, edited migrations, stale generated files and missing changelog lines. It printed failures and fixes and exited with code 2. After 3 blocks, it allowed stopping.

Our agent memory study graded 200 headless Claude Code CLI sessions (version 2.1.286) on one small repository. Per condition: n = 15 for Sonnet 5.5; n = 10 for Haiku 4.5.

What a Stop hook enforced, and what it missed

The hook and grader checked the same four code-rule categories with different file filters and patterns. Grades do not prove the hook caused compliance. We tested one hook on one repository.

On Sonnet 5.5, sessions with the hook alone passed 60 of 60 code-rule checks (95% interval 94.0% to 100%). With no memory, Sonnet passed 57 of 60 (86.3% to 98.3%). The intervals overlap: no clear lead. On Haiku 4.5, both passed 40 of 40 (95% interval 91.2% to 100%). The code rules hit a ceiling there.

Totals pool checks from 15 Sonnet and 10 Haiku sessions. Within-session checks are not independent.

Team knowledge covers the changelog and late-fee rate. All intervals are 95% Wilson intervals. Pooled results count checks; full passes count sessions.

Sonnet with the hook alone passed 10 of 15 team-knowledge checks (41.7% to 84.8%) against 6 of 15 with no memory (19.8% to 64.3%). Those intervals overlap: the gap is unclear. Haiku 4.5 passed 8 of 10 (49.0% to 94.3%) against 0 of 10 (0% to 27.8%). Those intervals do not overlap: the hook is ahead on Haiku here.

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest No memory 40% (95% interval 20%–64%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest Curated + hook 100% (95% interval 72%–100%, n 10). Lowest No memory 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row5 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Changelog rule and late-fee rate, pooled · 95% Wilson intervals

Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Split-check intervals are computed from the pass and fail outcomes in the Sonnet receipts and Haiku receipts. With the hook, Haiku passed the changelog check 8 of 8 times (67.6% to 100%) against 0 of 8 with no memory (0% to 32.4%). On Sonnet it was 10 of 12 (55.2% to 95.3%) against 6 of 12 (25.4% to 74.6%), and those overlap. Both Sonnet misses asked and wrote no code; the hook had nothing to check.

The script did not supply the late-fee rate (1.25%). Hooks could supply facts; we did not test that. With the hook alone, 0 of 3 Sonnet sessions used it (0% to 56.2%; direct receipt calculation). Two asked and wrote no code; one guessed. Across five memory-file conditions holding the rate, Sonnet used it 15 of 15 times (79.6% to 100%). Each had three late-fee sessions.

Unanswered headless questions count as failures. A person could answer live.

What the hook cost

A block adds turns. On Sonnet, median input per session was 125,674 tokens with the hook alone and 78,455 with no memory: 1.6x (calculation). On Haiku it was 530,935 against 301,233: 1.8x (receipt calculation). The hook blocked 5 times in 15 Sonnet sessions and 10 times in 10 Haiku sessions.

The 10 unblocked Sonnet hook sessions still had a median 126,604 input tokens (range 53,420 to 148,139; receipt calculation). We did not isolate why.

Input sums uncached input, cache reads and writes across turns (CLI-counter calculation). Most came from the prompt cache, including CLI context.

Session input ranges overlap; they are not confidence intervals. Sonnet: 53,420 to 176,503 with the hook, 48,251 to 173,108 with none (n = 15 each). Haiku: 286,447 to 733,979 with the hook, 148,407 to 529,443 with none (n = 10 each; range calculation from receipts). These ratios describe this run.

On Sonnet (n = 15), median time was 27.3 s with the hook alone (range 14.4 s to 38.3 s). With no memory it was 18.0 s (range 12.4 s to 44.8 s). On Haiku (n = 10) it was 68.5 s (range 42.3 s to 94.4 s) against 54.2 s (range 30.2 s to 66.2 s). Ranges overlap and are not confidence intervals; medians and ranges are linked-receipt calculations. Up to four sessions shared one machine. No-memory sessions often stopped to ask or skipped the changelog.

Time per session

Median wall time in seconds

  • Claude Sonnet 5.5
  • Claude Haiku 4.5 (square)
Sorted by gap, largest first.
No memory
/init CLAUDE.md
Stop hook only
Curated + hook
Curated, 11 lines
Handbook, 210 lines
Dreamed notes
Raw notes, 60 lines

Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: slowest Stop hook only 27.3 s (n 15). Fastest No memory 18 s (n 15). Claude Haiku 4.5: slowest Stop hook only 68.5 s (n 10). Fastest Handbook, 210 lines 49.9 s (n 10).

Notesn 10–15 per row

Up to four sessions ran at a time on one machine. Sessions without memory were often shorter because they stopped to ask or skipped the changelog.

Source: Agent memory study: 8 kinds of project memory on Claude Code

A hook plus a short memory file

A full pass requires passing hidden tests and every applicable check. Table intervals are direct receipt calculations (95% Wilson, z = 1.96).

ConditionSonnet 5.5 (n 15)Haiku 4.5 (n 10)
No memory9/15 (35.7% to 80.2%)2/10 (5.7% to 51.0%)
Stop hook only12/15 (54.8% to 93.0%)8/10 (49.0% to 94.3%)
11-line file plus hook15/15 (79.6% to 100%)9/10 (59.6% to 98.2%)

On Sonnet, all interval pairs overlap: no condition is ahead. The 11-line file alone also passed 15 of 15 (95% interval 79.6% to 100%): a ceiling for this task set.

On Haiku, the hook alone and no memory overlap (49.0% to 51.0%), so that gap is unclear. The file plus hook leads no memory: intervals do not overlap. The file alone (7 of 10; 39.7% to 89.2%) overlaps both, so we cannot claim the hook added to the file.

How to use hooks well

These suggestions concern one synthetic repository. The curated file’s author knew the tasks. We did not compare CLAUDE.md, skills and hooks as rule locations.

  1. Put checkable rules in a hook, such as a changelog line.
  2. Put facts in a short memory file, such as the late-fee rate.
  3. Keep the hook message short: the file, the rule and the fix. Our reading; we did not vary message length.
  4. Test on a few real tasks with and without the hook, and compare pass rate, time and input tokens.

Frequently asked questions

What is a Stop hook in Claude Code?

A Stop hook runs when Claude finishes. In the tested version, exit code 2 blocked finishing and sent standard-error feedback to Claude.

Do hooks replace CLAUDE.md?

Our hook checked code; it did not supply the late-fee rate. A hook could supply facts.

Do hooks cost tokens?

Feedback can add turns and tokens. Sonnet’s median input ratio was 1.6x (calculation; n = 15 per condition). The ranges above overlap. We did not isolate the gap’s cause.

When should I use a hook?

Use one for a rule that a script can check, and keep facts in a short file. We tested one Stop hook on one repository.

The data behind this explainer

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.