Does CLAUDE.md help? We tested 8 kinds of agent memory on Claude Code
200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a handbook and a Stop hook. Memory mattered where the repo was silent.
TL;DR
- We ran 200 Claude Code sessions (120 with Sonnet 5.5, 80 with Haiku 4.5) on 5 tasks in one small repository, with 8 kinds of memory: none, the real
/initfile, a curated 11-line file, 56 raw notes, the same notes after one "dreaming" pass, a 211-line handbook, a Stop hook, and curated plus hook. Hidden tests and rule checks graded every session. We declared the protocol first and counted every attempt. - Sonnet already follows what it can see. With no memory it passed 95% of the rule checks that the code shows and 100% of the rules that the folders hint at.
- Memory carries what the repository cannot show. Team-knowledge checks: 6 of 15 with no memory, 15 of 15 with an 11-line file.
- Asked to "charge our standard late fee", Sonnet stopped and asked for the rate 7 of 9 times when no memory held it. With the rate in any memory file, it was right 15 of 15 times, even in messy notes that also held an old rate. Haiku 4.5 never asked: it invented a rate and reported the task done in 6 of 6 sessions.
/initwrote a good file, but it copied a broken test command from the README. 13 of 15 Sonnet sessions then ran that command.- Messy memory misled the smaller model. Haiku 4.5 with the raw notes ran a stale test command in 10 of 10 sessions, though the corrected command was in the same file. After one "dreaming" pass over the same notes: 1 of 10. It is the clearest difference in the study.
- A Stop hook enforced every code rule, but it cannot carry a fact, and it used 1.6× the input tokens of no memory.
- Rule of thumb from the data: write down what the repo cannot show, enforce what a script can check, and delete the rest.
Full data, method and every memory file: /benchmarks/agent-memory.
Does memory help Claude Code? 8 kinds of agent memory, tested
200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.
Transcript
- Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
- Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
- Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
- Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
- The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
- Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- Write down what the repo cannot show. Enforce what a script can check.
Why we tested this now
Memory is the feature every coding agent shipped this year. Claude Code has CLAUDE.md files and auto memory. Codex has AGENTS.md and Memories. GitHub Copilot has Copilot Memory. On October 5, Devin launched "Memory and Dreaming": a nightly pass that merges notes, drops stale ones and rebuilds the index.
The evidence behind all of this is thin. In our review of 63 papers and vendor pages:
- Only GitHub published a controlled test of a memory feature: pull requests merged 90% vs 83%, with no sample size given.
- No one published a controlled test of a shipped "dreaming" feature. Devin, Codex, Letta and Anthropic's managed agents describe consolidation; none publish a result with a baseline.
- We found no measured rate of harm from stale memory in coding agents.
- Independent studies disagree. ETH Zurich found that LLM-written context files did not raise success and added 20% to 23% cost. A study by AWS and HSBC found that random rule files helped exactly as much as curated ones (63.8% each, against 50.0% with none).
So we ran a small, careful experiment of our own.
Disclosure: we build Agent, a product that carries work and corrections across AI tools. We have a stake in memory. That is why the protocol, the raw data and every memory file are public.
The experiment
We built tally, a small invoicing ledger in Node.js with no dependencies. It has the kind of conventions real teams have:
| Rule | Where an agent could learn it | Class |
|---|---|---|
Money is integer cents; no toFixed or parseFloat | The existing code uses parseAmount() and formatCents() | Code shows it |
Throw TallyError with a code, never new Error | Every existing error | Code shows it |
Time comes from nowIso(), never new Date() | The existing code | Code shows it |
Log with log(), never console.log | The existing code | Code shows it |
| Migrations are append-only; a schema change needs a new file | A migrations/ folder with numbered files | Folders hint at it |
| The report is generated from a template | A scripts/gen-report.mjs and a gen:report script | Folders hint at it |
| Every user-visible change gets a changelog line | Nowhere | Team knowledge |
| The late fee is 1.25% (finance decision, 2026-09-15) | Nowhere | Team knowledge |
One more trap is real and common: the README's test command fails on Node 25. node --test test/ treats the folder as a file. The right command is node --test.
Five tasks, each written the way a developer would ask: add refunds; fix a report that prints $5.1; add an optional phone number; "charge our standard late fee"; and a control task (fix slugify()) where memory should not matter.
Eight conditions:
| Condition | What the agent got |
|---|---|
| No memory | Nothing |
/init | The CLAUDE.md that Claude Code's own /init wrote for this repository (29 lines) |
| Curated | 9 facts, one line each (11 lines) |
| Raw notes | 56 dated notes from "past sessions": every current fact, plus repeats, one-off events and 5 notes that a later note contradicts (the late fee, money format, the clock helper, the test command, the dev-server port) |
| Dreamed | The raw notes after one consolidation pass by Claude Sonnet with a generic prompt (30 notes under 10 headings) |
| Handbook | The 9 curated facts inside a 211-line engineering handbook with about 170 generic rules |
| Stop hook | No memory file. A Claude Code Stop hook runs a rule checker and blocks the agent from finishing while a rule is broken |
| Curated + hook | Both |
Grading. After each session, hidden tests ran in a sandbox. Then deterministic checks read only the lines the agent added. A full pass means the tests pass and every rule that applies passes.
Isolation. Claude Code 2.1.286 ran headless with file and shell tools in its OS sandbox. A probe showed that --setting-sources project alone still loads your personal ~/.claude/CLAUDE.md, so we excluded it with claudeMdExcludes and turned auto memory off. Each session got a fresh repository. Sonnet ran 3 repetitions per cell, Haiku 2.
Result 1: the model already reads the code
With no memory at all, Sonnet passed 95% of the code-visible checks and 100% of the folder-visible ones. It found the report generator by itself. It added new migrations instead of editing old ones. It used TallyError because every other error in the code does.
The only code-rule misses were on the report task: 3 sessions formatted money with toFixed(2). The /init file did not prevent that (3 of 3 again). Every file that said "never toFixed for money" did.
This matches what Anthropic says about CLAUDE.md: do not describe what Claude can learn by reading the code. It also matches the ETH finding that repository overviews do not help an agent find files.
Result 2: memory carries what the repository cannot
The late-fee task is the clearest case. The rate is in no file. Without it in memory, Sonnet searched the source, the README, the changelog and the git history, then stopped with a question in 7 of 9 sessions. In one session it wrote:
I haven't written anything yet. The repo doesn't define what "our standard late fee" is, and I'd rather not guess a billing number.
When it did guess (2 sessions), it said so and picked 1.5% and 5%. That is honest behavior. But each question is a round trip to a person: the re-briefing that memory is supposed to remove.
With the rate in any memory file, Sonnet used 1.25% in 15 of 15 sessions. That includes the raw notes, where a July note says 2% and a September note says 1.25%. Sonnet took the newer one every time.
The smaller model behaved differently, and worse. Haiku 4.5 never asked. Without the rate it invented one in 6 of 6 sessions (5%, 2.5%, 2% and 1.5%) and reported the task done. One summary called its number "a standard late fee rate". None of its final messages said the number was a guess; all 9 of Sonnet's did, or asked.
A missing fact costs a strong model a question. It costs a small model a silent, confident error. If cheaper models do your subagent work, the facts they cannot see belong in memory.
The changelog rule split the same way. With no memory, Sonnet added a changelog line for both new features in every session, but never for the bug fix. Every file that stated the rule brought it to 12 of 12, except where a session stopped to ask about the late fee and wrote nothing.
Result 3: /init summarises what was already there
The real /init output is a good file. It found the generated report, the append-only migrations, integer cents and the changelog rule. It did not know the late fee, because no file holds it.
It also copied the README's test command. 13 of 15 Sonnet sessions with the /init file ran the broken command, against 12 of 15 with no memory. For Sonnet, every memory file that named the right command brought this to 0.
Lesson: /init writes down your repository's documentation, errors included. Fix stale docs at the source, and add what the docs never said.
Result 4: messy memory, and what "dreaming" fixed
The raw notes hold every current fact, and the old ones too. In June a note says "Run tests with npm test". In September another note says that command fails on Node 25 and gives the right one.
Sonnet handled the mess. With the raw notes it never ran the stale command (0 of 15), used the newer late fee every time and passed 14 of 15 sessions in full. The one miss was a refund that did not re-open the balance; it did not come from the notes.
Haiku 4.5 did not. With the same raw notes it ran the stale npm test in 10 of 10 sessions. After one dreaming pass: 1 of 10. The 95% intervals do not overlap (72–100% against 2–40%), so by our declared rule this is a clear difference.
Team knowledge for Haiku followed the same order: 0% with no memory, 60% with the raw notes, 80% with the dreamed notes or the curated file, and 100% with curated plus hook.
We also scored the dreaming pass itself. We ran it 3 times on the 56 raw notes.
All 3 dreams kept 10 of 10 current facts and left 0 of 5 stale notes. Each kept 1 of 9 transient notes: the current dev-server port, which is arguably still useful. Each dream took about 7 seconds and one call. This is one small file, not months of notes. The ACE paper (ICLR 2026) shows the risk at scale: one rewrite collapsed 18,282 tokens of context to 122 tokens and cut accuracy below the no-memory baseline. Read the diff of a dream before you accept it.
Result 5: long files cost tokens, and a small model pays more
For Sonnet, the 211-line handbook did as well as the 11-line file: 15 of 15 full passes. It used 31% more input tokens (median 85,223 against 65,045, our arithmetic).
For Haiku, the same 9 facts inside the handbook were followed less often. The changelog line worked in 1 of 8 sessions inside the handbook and in 6 of 8 in the 11-line file. Full passes: 3 of 10 against 7 of 10. With 10 sessions per condition the intervals overlap, but the direction matches the instruction-density research: the more rules around a rule, the more often one gets dropped.
If you run smaller or cheaper models as subagents, short memory matters more.
Result 6: hooks enforce rules, they do not carry facts
The Stop hook ran a 52-line checker for the code rules. Sonnet with the hook alone passed 100% of the code and folder rules. But the hook has nothing to say about a late-fee rate. Hook-only sessions asked or guessed on that task, like no memory.
Enforcement has a cost. Each block sends the agent back for another round of work.
The hook alone used 1.6× the median input tokens of no memory. The curated file used the fewest input tokens of any condition: the agent explored less because it already knew where things were. Curated plus hook cost about the same as curated alone, because the agent rarely broke a rule that the file already stated.
Result 7: cost per correct result
A failed session still costs money. So we divided each condition's list-price estimate by its full passes (a calculation; the sessions ran on a subscription).
For Sonnet, a fully correct result cost $0.082 with the 11-line file against $0.139 with no memory: 41% less (our calculation). For Haiku the gap was wider: $0.096 with curated plus hook against $0.379 with no memory, because only 2 of 10 no-memory sessions passed in full. A short memory file paid for itself.
The overall score, read honestly
For Sonnet, every condition with the team facts in memory beat no memory on full passes; /init tied it. But with 15 sessions per condition, most 95% intervals overlap, including curated against none by a hair. For Haiku, one pair is clear: curated plus hook passed 9 of 10 in full, no memory 2 of 10. Read the rest as direction, not a ranking. The per-class results above are clearer than the total, because a total mixes rules the model never needed help with and facts it could not know.
What other research says, and where this fits
| Finding | Source | What it means next to our data |
|---|---|---|
| LLM-written context files: success −0.5 and −2 points (not significant), cost +20% to +23% | ETH Zurich, arXiv 2602.11988 | Same shape as our /init result: a summary of the repo adds tokens, not knowledge |
| Random rule files = curated rule files (63.8% each, 50.0% with none); only "do not" rules helped one by one | AWS and HSBC, arXiv 2604.11088 | Much of a rule file's effect is priming. The part that is not: facts and constraints the model cannot infer |
| Unfiltered memory summaries scored below no memory (22.2% vs 26.3%); perfect summaries 34.3% | SWE-ContextBench, arXiv 2602.08316 | Bad memory can be worse than none, which our Haiku raw-notes result echoes |
| GPT-4o followed all rules 94% of the time with 1 rule, 21% with 10 | Harada et al., EMNLP 2025 Findings | "Follow every rule" is harder than "follow each rule"; long files raise the odds of a miss |
| A rewrite collapsed 18,282 tokens of context to 122; accuracy fell below baseline | ACE, arXiv 2510.04618 | Consolidation can destroy memory; review it |
Payloads planted in CLAUDE.md or AGENTS.md persisted across sessions in 96.7% of runs (Opus 4.7, n = 10 per cell) | "Bad Memory", arXiv 2607.14611 | Memory files are an attack surface. Review them like code |
| Anthropic removed over 80% of Claude Code's system prompt "with no measurable loss" | Anthropic, via @trq212 | Newer models need fewer instructions, not more |
The full review, with 63 sources and every number checked against its source, is in our research notes.
What goes where
| Put it in… | When | Examples |
|---|---|---|
| Nowhere (delete it) | The code already shows it | Folder layout, stack, style that a linter enforces, architecture overviews |
| A hook, test or lint | A script can check it | No floats for money, append-only migrations, generated files in sync, a changelog entry |
| A short memory file | Only people know it | Decisions with a date ("late fee 1.25%, 2026-09-15"), gotchas ("run node --test"), "do not" rules with the reason |
| A skill | A procedure you need sometimes | Release steps, a migration recipe, an incident runbook |
A memory checklist
- Write only what the repository cannot show: decisions, numbers, gotchas and the reason behind a rule.
- Turn every rule a script can check into a hook, test or lint. Keep one line in memory for the reason.
- Date every decision. When a fact changes, replace the old line. Do not append a second rate below the first.
- Fix stale docs at the source.
/initcopies them. - Keep the file short. 11 lines did as well as anything we tried; Anthropic's target is under 200 lines.
- Consolidate on a schedule: merge repeats, drop one-off events, resolve contradictions. Then read the diff.
- Treat memory as code. Review changes and look for planted instructions.
- Test with and without the file. Run 3 to 5 real tasks both ways and keep the lines that change the outcome.
Copy these
- A curated file in the shape that worked. Ours is 9 lines: facts, a date on each decision, the reason when there is one. The exact file is in the published memory files.
- The Stop hook. A 52-line Node script that checks the changed lines and exits with code 2 to send problems back to the agent. It stops blocking after 3 tries, so a session cannot loop. Same file.
- The dreaming prompt and all three dream outputs: dreams.json.
Limits
- One small synthetic repository, and one author of its conventions. The curated file is an upper bound: it was written with knowledge of the tasks.
- The hook checks the same code rules as the grader. It shows what rules written as code can do.
- n = 15 per condition for Sonnet and 10 for Haiku. Most full-pass intervals overlap.
- Headless sessions cannot get an answer to a question, so asking counts as a failure here. In a live session a person would answer; the cost is the round trip.
- Costs are the CLI's list-price estimates for subscription sessions, not invoices.
FAQ
Does CLAUDE.md help Claude Code?
In our test, yes for knowledge the repository cannot show, and barely for anything else. Sonnet 5.5 followed visible code patterns without it. Team-only knowledge went from 6 of 15 checks to 15 of 15 with an 11-line file.
How long should CLAUDE.md be?
Short enough that every line carries something the code cannot. Our 11-line file matched every longer file for Sonnet and used the fewest tokens. For Haiku 4.5, the same facts inside a 211-line handbook were followed less often (the changelog rule in 1 of 8 sessions, against 6 of 8 in the short file).
Should I run /init?
It is a fine start, but it describes what is already in your repository, including stale docs. Add the decisions and gotchas it cannot know, and fix wrong docs at the source.
What is agent "dreaming"?
A scheduled pass that rewrites an agent's memory: it merges repeats, drops stale and one-off notes, and resolves contradictions. Devin shipped it on 2026-10-05. In our test one pass cleaned 56 notes correctly 3 times out of 3, and it helped the smaller model most.
CLAUDE.md or hooks?
Both, for different jobs. A hook enforces a rule that a script can check, every time. A memory file carries facts and the reasons behind rules. A hook alone cost 1.6× the input tokens and could not carry a single fact.
Method and data
- Study page with every chart, table and caveat: /benchmarks/agent-memory
- Raw data: Sonnet attempts, Haiku attempts, memory files, dreams; the declared protocol is summarised in the study's Method
- Related: prompt caching, measured and the hidden context tax of coding CLIs