• Claude Code
  • Agent Memory
  • CLAUDE.md
  • Hooks

Does CLAUDE.md help? We tested 8 kinds of agent memory on Claude Code

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a handbook and a Stop hook. Memory mattered where the repo was silent.

TL;DR

  • We ran 200 Claude Code sessions (120 with Sonnet 5.5, 80 with Haiku 4.5) on 5 tasks in one small repository, with 8 kinds of memory: none, the real /init file, a curated 11-line file, 56 raw notes, the same notes after one "dreaming" pass, a 211-line handbook, a Stop hook, and curated plus hook. Hidden tests and rule checks graded every session. We declared the protocol first and counted every attempt.
  • Sonnet already follows what it can see. With no memory it passed 95% of the rule checks that the code shows and 100% of the rules that the folders hint at.
  • Memory carries what the repository cannot show. Team-knowledge checks: 6 of 15 with no memory, 15 of 15 with an 11-line file.
  • Asked to "charge our standard late fee", Sonnet stopped and asked for the rate 7 of 9 times when no memory held it. With the rate in any memory file, it was right 15 of 15 times, even in messy notes that also held an old rate. Haiku 4.5 never asked: it invented a rate and reported the task done in 6 of 6 sessions.
  • /init wrote a good file, but it copied a broken test command from the README. 13 of 15 Sonnet sessions then ran that command.
  • Messy memory misled the smaller model. Haiku 4.5 with the raw notes ran a stale test command in 10 of 10 sessions, though the corrected command was in the same file. After one "dreaming" pass over the same notes: 1 of 10. It is the clearest difference in the study.
  • A Stop hook enforced every code rule, but it cannot carry a fact, and it used 1.6× the input tokens of no memory.
  • Rule of thumb from the data: write down what the repo cannot show, enforce what a script can check, and delete the rest.

Full data, method and every memory file: /benchmarks/agent-memory.

Live story · 62 sDoes memory help Claude Code? 8 kinds of agent memory, tested

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.

Transcript
  1. Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
  2. Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  3. Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  4. With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
  5. Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
  6. The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  7. The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
  8. Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  9. Write down what the repo cannot show. Enforce what a script can check.

Why we tested this now

Memory is the feature every coding agent shipped this year. Claude Code has CLAUDE.md files and auto memory. Codex has AGENTS.md and Memories. GitHub Copilot has Copilot Memory. On October 5, Devin launched "Memory and Dreaming": a nightly pass that merges notes, drops stale ones and rebuilds the index.

The evidence behind all of this is thin. In our review of 63 papers and vendor pages:

  • Only GitHub published a controlled test of a memory feature: pull requests merged 90% vs 83%, with no sample size given.
  • No one published a controlled test of a shipped "dreaming" feature. Devin, Codex, Letta and Anthropic's managed agents describe consolidation; none publish a result with a baseline.
  • We found no measured rate of harm from stale memory in coding agents.
  • Independent studies disagree. ETH Zurich found that LLM-written context files did not raise success and added 20% to 23% cost. A study by AWS and HSBC found that random rule files helped exactly as much as curated ones (63.8% each, against 50.0% with none).

So we ran a small, careful experiment of our own.

Disclosure: we build Agent, a product that carries work and corrections across AI tools. We have a stake in memory. That is why the protocol, the raw data and every memory file are public.

The experiment

We built tally, a small invoicing ledger in Node.js with no dependencies. It has the kind of conventions real teams have:

RuleWhere an agent could learn itClass
Money is integer cents; no toFixed or parseFloatThe existing code uses parseAmount() and formatCents()Code shows it
Throw TallyError with a code, never new ErrorEvery existing errorCode shows it
Time comes from nowIso(), never new Date()The existing codeCode shows it
Log with log(), never console.logThe existing codeCode shows it
Migrations are append-only; a schema change needs a new fileA migrations/ folder with numbered filesFolders hint at it
The report is generated from a templateA scripts/gen-report.mjs and a gen:report scriptFolders hint at it
Every user-visible change gets a changelog lineNowhereTeam knowledge
The late fee is 1.25% (finance decision, 2026-09-15)NowhereTeam knowledge

One more trap is real and common: the README's test command fails on Node 25. node --test test/ treats the folder as a file. The right command is node --test.

Five tasks, each written the way a developer would ask: add refunds; fix a report that prints $5.1; add an optional phone number; "charge our standard late fee"; and a control task (fix slugify()) where memory should not matter.

Eight conditions:

ConditionWhat the agent got
No memoryNothing
/initThe CLAUDE.md that Claude Code's own /init wrote for this repository (29 lines)
Curated9 facts, one line each (11 lines)
Raw notes56 dated notes from "past sessions": every current fact, plus repeats, one-off events and 5 notes that a later note contradicts (the late fee, money format, the clock helper, the test command, the dev-server port)
DreamedThe raw notes after one consolidation pass by Claude Sonnet with a generic prompt (30 notes under 10 headings)
HandbookThe 9 curated facts inside a 211-line engineering handbook with about 170 generic rules
Stop hookNo memory file. A Claude Code Stop hook runs a rule checker and blocks the agent from finishing while a rule is broken
Curated + hookBoth

Grading. After each session, hidden tests ran in a sandbox. Then deterministic checks read only the lines the agent added. A full pass means the tests pass and every rule that applies passes.

Isolation. Claude Code 2.1.286 ran headless with file and shell tools in its OS sandbox. A probe showed that --setting-sources project alone still loads your personal ~/.claude/CLAUDE.md, so we excluded it with claudeMdExcludes and turned auto memory off. Each session got a fresh repository. Sonnet ran 3 repetitions per cell, Haiku 2.

Result 1: the model already reads the code

Rules the code already shows

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Rules the folders hint at

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Team knowledge only

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 3 series: Rules the code already shows, Rules the folders hint at, Team knowledge only. Rules the code already shows: highest Curated, 11 lines 100% (95% interval 94%–100%, n 60). Lowest /init CLAUDE.md 95% (95% interval 86%–98%, n 60). All intervals overlap. Rules the folders hint at: all at 100%.

NotesWhiskers: 95% Wilson intervaln 15–60 per row6 of 8 (Rules the code already shows) at 100%: this task set cannot separate them.

Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

Source: Agent memory study: 8 kinds of project memory on Claude Code

With no memory at all, Sonnet passed 95% of the code-visible checks and 100% of the folder-visible ones. It found the report generator by itself. It added new migrations instead of editing old ones. It used TallyError because every other error in the code does.

The only code-rule misses were on the report task: 3 sessions formatted money with toFixed(2). The /init file did not prevent that (3 of 3 again). Every file that said "never toFixed for money" did.

This matches what Anthropic says about CLAUDE.md: do not describe what Claude can learn by reading the code. It also matches the ETH finding that repository overviews do not help an agent find files.

Result 2: memory carries what the repository cannot

  • Used the current rate (1.25%)
  • Asked for the rate, wrote no code
  • Guessed another rate
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One square per count; counts at the right are exact and in legend order.

8 rows, 3 series: Used the current rate (1.25%), Asked for the rate, wrote no code, Guessed another rate. Used the current rate (1.25%): highest Curated, 11 lines 3 (n 3). Lowest Stop hook only 0 (n 3). Asked for the rate, wrote no code: highest No memory 3 (n 3). Lowest Curated + hook 0 (n 3).

Notesn = 3 per row

The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%

Late-fee task, 3 sessions per condition. "Asked" means the agent searched the repository, found no rate and stopped with a question instead of code.

Source: Agent memory study: 8 kinds of project memory on Claude Code

The late-fee task is the clearest case. The rate is in no file. Without it in memory, Sonnet searched the source, the README, the changelog and the git history, then stopped with a question in 7 of 9 sessions. In one session it wrote:

I haven't written anything yet. The repo doesn't define what "our standard late fee" is, and I'd rather not guess a billing number.

When it did guess (2 sessions), it said so and picked 1.5% and 5%. That is honest behavior. But each question is a round trip to a person: the re-briefing that memory is supposed to remove.

With the rate in any memory file, Sonnet used 1.25% in 15 of 15 sessions. That includes the raw notes, where a July note says 2% and a September note says 1.25%. Sonnet took the newer one every time.

The smaller model behaved differently, and worse. Haiku 4.5 never asked. Without the rate it invented one in 6 of 6 sessions (5%, 2.5%, 2% and 1.5%) and reported the task done. One summary called its number "a standard late fee rate". None of its final messages said the number was a guess; all 9 of Sonnet's did, or asked.

  • Used the current rate (1.25%)
  • Guessed another rate
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One square per count; counts at the right are exact and in legend order.

8 rows, 2 series: Used the current rate (1.25%), Guessed another rate. Used the current rate (1.25%): highest Curated, 11 lines 2 (n 2). Lowest Stop hook only 0 (n 2). Guessed another rate: highest No memory 2 (n 2). Lowest Curated + hook 0 (n 2).

Notesn = 2 per row

The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%

Late-fee task, 2 sessions per condition. "Asked" means the agent searched the repository, found no rate and stopped with a question instead of code.

Source: Agent memory study: 8 kinds of project memory on Claude Code

A missing fact costs a strong model a question. It costs a small model a silent, confident error. If cheaper models do your subagent work, the facts they cannot see belong in memory.

The changelog rule split the same way. With no memory, Sonnet added a changelog line for both new features in every session, but never for the bug fix. Every file that stated the rule brought it to 12 of 12, except where a session stopped to ask about the late fee and wrote nothing.

Result 3: /init summarises what was already there

The real /init output is a good file. It found the generated report, the append-only migrations, integer cents and the changelog rule. It did not know the late fee, because no file holds it.

It also copied the README's test command. 13 of 15 Sonnet sessions with the /init file ran the broken command, against 12 of 15 with no memory. For Sonnet, every memory file that named the right command brought this to 0.

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest /init CLAUDE.md 87% (95% interval 62%–96%, n 15). Lowest Curated + hook 0% (95% interval 0%–20%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest No memory 100% (95% interval 72%–100%, n 10). Lowest Curated + hook 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row

Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals

The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Lesson: /init writes down your repository's documentation, errors included. Fix stale docs at the source, and add what the docs never said.

Result 4: messy memory, and what "dreaming" fixed

The raw notes hold every current fact, and the old ones too. In June a note says "Run tests with npm test". In September another note says that command fails on Node 25 and gives the right one.

Sonnet handled the mess. With the raw notes it never ran the stale command (0 of 15), used the newer late fee every time and passed 14 of 15 sessions in full. The one miss was a refund that did not re-open the balance; it did not come from the notes.

Haiku 4.5 did not. With the same raw notes it ran the stale npm test in 10 of 10 sessions. After one dreaming pass: 1 of 10. The 95% intervals do not overlap (72–100% against 2–40%), so by our declared rule this is a clear difference.

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest No memory 40% (95% interval 20%–64%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest Curated + hook 100% (95% interval 72%–100%, n 10). Lowest No memory 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row5 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Changelog rule and late-fee rate, pooled · 95% Wilson intervals

Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Team knowledge for Haiku followed the same order: 0% with no memory, 60% with the raw notes, 80% with the dreamed notes or the curated file, and 100% with curated plus hook.

We also scored the dreaming pass itself. We ran it 3 times on the 56 raw notes.

Current facts kept (of 10)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

Stale notes left (of 5)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

Transient notes left (of 9)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

One panel per series, all on the same axis.

6 rows, 3 series: Current facts kept (of 10), Stale notes left (of 5), Transient notes left (of 9). Current facts kept (of 10): highest Raw notes 10. Lowest /init 5. Stale notes left (of 5): highest Raw notes 5. Lowest Curated 0.

Notes

Fixed-pattern checks on each memory file: 10 current facts, 5 stale notes, 9 transient notes

Each dream is one Claude Sonnet call with no tools over the 56 raw notes. A stale note that the new file marks as replaced does not count as left. The one transient note every dream kept is the current dev-server port.

Source: Agent memory study: 8 kinds of project memory on Claude Code

All 3 dreams kept 10 of 10 current facts and left 0 of 5 stale notes. Each kept 1 of 9 transient notes: the current dev-server port, which is arguably still useful. Each dream took about 7 seconds and one call. This is one small file, not months of notes. The ACE paper (ICLR 2026) shows the risk at scale: one rewrite collapsed 18,282 tokens of context to 122 tokens and cut accuracy below the no-memory baseline. Read the diff of a dream before you accept it.

Result 5: long files cost tokens, and a small model pays more

For Sonnet, the 211-line handbook did as well as the 11-line file: 15 of 15 full passes. It used 31% more input tokens (median 85,223 against 65,045, our arithmetic).

For Haiku, the same 9 facts inside the handbook were followed less often. The changelog line worked in 1 of 8 sessions inside the handbook and in 6 of 8 in the 11-line file. Full passes: 3 of 10 against 7 of 10. With 10 sessions per condition the intervals overlap, but the direction matches the instruction-density research: the more rules around a rule, the more often one gets dropped.

If you run smaller or cheaper models as subagents, short memory matters more.

Result 6: hooks enforce rules, they do not carry facts

The Stop hook ran a 52-line checker for the code rules. Sonnet with the hook alone passed 100% of the code and folder rules. But the hook has nothing to say about a late-fee rate. Hook-only sessions asked or guessed on that task, like no memory.

Enforcement has a cost. Each block sends the agent back for another round of work.

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows. Highest Stop hook only 125,674 (n 15). Lowest Curated, 11 lines 65,045 (n 15).

Notesn = 15 per row

Median input tokens per session, cache reads included (Claude Sonnet 5.5)

Input = uncached input + cache reads + cache writes over every turn, as the CLI reports it. Most of it is read from the prompt cache. The hook adds turns: each block sends the agent back to work.

Source: Agent memory study: 8 kinds of project memory on Claude Code

The hook alone used 1.6× the median input tokens of no memory. The curated file used the fewest input tokens of any condition: the agent explored less because it already knew where things were. Curated plus hook cost about the same as curated alone, because the agent rarely broke a rule that the file already stated.

Result 7: cost per correct result

A failed session still costs money. So we divided each condition's list-price estimate by its full passes (a calculation; the sessions ran on a subscription).

Calculation
  • Claude Sonnet 5.5
  • Claude Haiku 4.5 (square)
Sorted by gap, largest first.
/init CLAUDE.md
No memory
Handbook, 210 lines
Curated, 11 lines
Dreamed notes
Raw notes, 60 lines
Curated + hook
Stop hook only

Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).

List-price calculation, not a run. 8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest No memory $0.14 (n 9). Lowest Curated, 11 lines $0.082 (n 15). Claude Haiku 4.5: highest /init CLAUDE.md $0.43 (n 2). Lowest Curated + hook $0.096 (n 9).

Notesn 2–15 per row

Sum of the CLI's cost estimates for a condition, divided by its full passes

Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.

Source: Agent memory study: 8 kinds of project memory on Claude Code

For Sonnet, a fully correct result cost $0.082 with the 11-line file against $0.139 with no memory: 41% less (our calculation). For Haiku the gap was wider: $0.096 with curated plus hook against $0.379 with no memory, because only 2 of 10 no-memory sessions passed in full. A short memory file paid for itself.

The overall score, read honestly

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest /init CLAUDE.md 60% (95% interval 36%–80%, n 15). All intervals overlap. Claude Haiku 4.5: highest Curated + hook 90% (95% interval 60%–98%, n 10). Lowest /init CLAUDE.md 20% (95% interval 5.7%–51%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row4 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Hidden tests pass and every convention check passes · 95% Wilson intervals

Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.

Source: Agent memory study: 8 kinds of project memory on Claude Code

For Sonnet, every condition with the team facts in memory beat no memory on full passes; /init tied it. But with 15 sessions per condition, most 95% intervals overlap, including curated against none by a hair. For Haiku, one pair is clear: curated plus hook passed 9 of 10 in full, no memory 2 of 10. Read the rest as direction, not a ranking. The per-class results above are clearer than the total, because a total mixes rules the model never needed help with and facts it could not know.

What other research says, and where this fits

FindingSourceWhat it means next to our data
LLM-written context files: success −0.5 and −2 points (not significant), cost +20% to +23%ETH Zurich, arXiv 2602.11988Same shape as our /init result: a summary of the repo adds tokens, not knowledge
Random rule files = curated rule files (63.8% each, 50.0% with none); only "do not" rules helped one by oneAWS and HSBC, arXiv 2604.11088Much of a rule file's effect is priming. The part that is not: facts and constraints the model cannot infer
Unfiltered memory summaries scored below no memory (22.2% vs 26.3%); perfect summaries 34.3%SWE-ContextBench, arXiv 2602.08316Bad memory can be worse than none, which our Haiku raw-notes result echoes
GPT-4o followed all rules 94% of the time with 1 rule, 21% with 10Harada et al., EMNLP 2025 Findings"Follow every rule" is harder than "follow each rule"; long files raise the odds of a miss
A rewrite collapsed 18,282 tokens of context to 122; accuracy fell below baselineACE, arXiv 2510.04618Consolidation can destroy memory; review it
Payloads planted in CLAUDE.md or AGENTS.md persisted across sessions in 96.7% of runs (Opus 4.7, n = 10 per cell)"Bad Memory", arXiv 2607.14611Memory files are an attack surface. Review them like code
Anthropic removed over 80% of Claude Code's system prompt "with no measurable loss"Anthropic, via @trq212Newer models need fewer instructions, not more

The full review, with 63 sources and every number checked against its source, is in our research notes.

What goes where

Put it in…WhenExamples
Nowhere (delete it)The code already shows itFolder layout, stack, style that a linter enforces, architecture overviews
A hook, test or lintA script can check itNo floats for money, append-only migrations, generated files in sync, a changelog entry
A short memory fileOnly people know itDecisions with a date ("late fee 1.25%, 2026-09-15"), gotchas ("run node --test"), "do not" rules with the reason
A skillA procedure you need sometimesRelease steps, a migration recipe, an incident runbook

A memory checklist

  1. Write only what the repository cannot show: decisions, numbers, gotchas and the reason behind a rule.
  2. Turn every rule a script can check into a hook, test or lint. Keep one line in memory for the reason.
  3. Date every decision. When a fact changes, replace the old line. Do not append a second rate below the first.
  4. Fix stale docs at the source. /init copies them.
  5. Keep the file short. 11 lines did as well as anything we tried; Anthropic's target is under 200 lines.
  6. Consolidate on a schedule: merge repeats, drop one-off events, resolve contradictions. Then read the diff.
  7. Treat memory as code. Review changes and look for planted instructions.
  8. Test with and without the file. Run 3 to 5 real tasks both ways and keep the lines that change the outcome.

Copy these

  • A curated file in the shape that worked. Ours is 9 lines: facts, a date on each decision, the reason when there is one. The exact file is in the published memory files.
  • The Stop hook. A 52-line Node script that checks the changed lines and exits with code 2 to send problems back to the agent. It stops blocking after 3 tries, so a session cannot loop. Same file.
  • The dreaming prompt and all three dream outputs: dreams.json.

Limits

  • One small synthetic repository, and one author of its conventions. The curated file is an upper bound: it was written with knowledge of the tasks.
  • The hook checks the same code rules as the grader. It shows what rules written as code can do.
  • n = 15 per condition for Sonnet and 10 for Haiku. Most full-pass intervals overlap.
  • Headless sessions cannot get an answer to a question, so asking counts as a failure here. In a live session a person would answer; the cost is the round trip.
  • Costs are the CLI's list-price estimates for subscription sessions, not invoices.

FAQ

Does CLAUDE.md help Claude Code?

In our test, yes for knowledge the repository cannot show, and barely for anything else. Sonnet 5.5 followed visible code patterns without it. Team-only knowledge went from 6 of 15 checks to 15 of 15 with an 11-line file.

How long should CLAUDE.md be?

Short enough that every line carries something the code cannot. Our 11-line file matched every longer file for Sonnet and used the fewest tokens. For Haiku 4.5, the same facts inside a 211-line handbook were followed less often (the changelog rule in 1 of 8 sessions, against 6 of 8 in the short file).

Should I run /init?

It is a fine start, but it describes what is already in your repository, including stale docs. Add the decisions and gotchas it cannot know, and fix wrong docs at the source.

What is agent "dreaming"?

A scheduled pass that rewrites an agent's memory: it merges repeats, drops stale and one-off notes, and resolves contradictions. Devin shipped it on 2026-10-05. In our test one pass cleaned 56 notes correctly 3 times out of 3, and it helped the smaller model most.

CLAUDE.md or hooks?

Both, for different jobs. A hook enforces a rule that a script can check, every time. A memory file carries facts and the reasons behind rules. A hook alone cost 1.6× the input tokens and could not carry a single fact.

Method and data

The data behind this post

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.