{"i":8,"study":{"slug":"agent-memory","title":"Does memory help Claude Code? 8 kinds of agent memory, tested","seoTitle":"Does CLAUDE.md help? Agent memory tested on Claude Code","description":"200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.","question":"Does project memory make Claude Code do better work, and which kind of memory?","answer":"Claude Sonnet 5.5 followed almost every rule it could see without any memory (95% of code rules, 100% of folder rules), but only 6/15 of the team-knowledge checks; with an 11-line curated file it passed 15/15. When a memory file held the late-fee rate, Sonnet used it 15/15 times, even from raw notes that also held a stale rate; without it, Sonnet stopped and asked 7 of 9 times, while Haiku invented a rate and flagged it in 0/6 sessions. The smaller Haiku 4.5 was misled by messy memory: with the raw notes it ran the stale test command in 10/10 sessions, and after one dreaming pass in 1/10. A Stop hook enforced every code rule but could not carry a fact, and it used 1.6× the input tokens of no memory. Full-pass intervals overlap for most pairs (n = 15 per condition).","date":"2026-10-06","updated":"2026-10-06","tags":["claude-code","agent-memory","claude-md","hooks","context-engineering"],"caveats":["One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.","The hook checks the same code rules as the grader. It shows what rules written as code can do; it cannot carry a fact such as the late-fee rate.","n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.","Costs are the CLI's list-price estimates for subscription sessions, not invoices.","In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip."],"sourceIds":["agent-memory-study"],"stats":{"$k":["id","label","value","unit","display","n","note","ci"],"$r":[["memory-sessions","Claude Code sessions, every one graded (120 Sonnet 5.5, 80 Haiku 4.5)",200,"count","200",200,"0 planned sessions not run.","\u0001"],["memory-team-knowledge-none","Team-knowledge checks passed with no memory (Sonnet 5.5)",0.4,"rate","40% (6/15)",15,"\u0001",[0.1982,0.6425]],["memory-team-knowledge-curated","Team-knowledge checks passed with an 11-line curated file (Sonnet 5.5)",1,"rate","100% (15/15)",15,"\u0001",[0.7961,1]],["memory-late-fee-known","Late fee right when any memory file held the rate (Sonnet 5.5)",1,"rate","15/15",15,"\u0001",[0.7961,1]],["memory-late-fee-asked","Without the rate, sessions that asked for it and wrote no code (Sonnet 5.5)",0.7778,"rate","7/9",9,"The other 2 guessed a rate and said it was a placeholder.","\u0001"],["memory-late-fee-flagged-sonnet","Without the rate: sessions whose final message said the rate was unknown or a guess (Sonnet 5.5)",1,"rate","9/9",9,"\u0001","\u0001"],["memory-late-fee-flagged-haiku","Without the rate: sessions whose final message said the rate was unknown or a guess (Haiku 4.5)",0,"rate","0/6",6,"Every other Haiku session invented a rate and reported the task done.","\u0001"],["memory-init-broken-command","Sessions with the /init file that ran the broken test command it copied from the README (Sonnet 5.5)",0.8667,"rate","13/15",15,"\u0001","\u0001"],["memory-hook-input-ratio","Median input tokens per session, Stop hook only vs no memory (calculation, Sonnet 5.5)",1.6,"ratio","1.6×",15,"\u0001","\u0001"],["memory-haiku-raw-broken-command","Haiku 4.5 with the raw notes: sessions that ran the stale test command",1,"rate","10/10",10,"\u0001","\u0001"],["memory-haiku-dreamed-broken-command","Haiku 4.5 with the dreamed notes: sessions that ran the stale test command",0.1,"rate","1/10",10,"\u0001","\u0001"],["memory-total-cost","List-price estimate of every session (calculation, not an invoice)",17.49,"usd","$17.49",200,"\u0001","\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","whisker","series","note","sourceIds","polarity"],"$r":[["memory-full-pass","Full pass rate by kind of memory","Hidden tests pass and every convention check passes · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["No memory",0.6,0.3575,0.8018,15,"\u0001"],["/init CLAUDE.md",0.6,0.3575,0.8018,15,"\u0001"],["Curated, 11 lines",1,0.7961,1,15,true],["Raw notes, 60 lines",0.9333,0.7018,0.9881,15,"\u0001"],["Dreamed notes",1,0.7961,1,15,"\u0001"],["Handbook, 210 lines",1,0.7961,1,15,"\u0001"],["Stop hook only",0.8,0.5481,0.9295,15,"\u0001"],["Curated + hook",1,0.7961,1,15,"\u0001"]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.2,0.0567,0.5098,10],["/init CLAUDE.md",0.2,0.0567,0.5098,10],["Curated, 11 lines",0.7,0.3968,0.8922,10],["Raw notes, 60 lines",0.6,0.3127,0.8318,10],["Dreamed notes",0.7,0.3968,0.8922,10],["Handbook, 210 lines",0.3,0.1078,0.6032,10],["Stop hook only",0.8,0.4902,0.9433,10],["Curated + hook",0.9,0.5958,0.9821,10]]}}],"Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.",["agent-memory-study"],"\u0001"],["memory-knowledge-class","Where memory helps: what the repo shows vs what only the team knows","Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)","grouped-bar","rate","\u0001",{"$k":["name","points"],"$r":[["Rules the code already shows",{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.95,0.863,0.9829,60],["/init CLAUDE.md",0.95,0.863,0.9829,60],["Curated, 11 lines",1,0.9398,1,60],["Raw notes, 60 lines",1,0.9398,1,60],["Dreamed notes",1,0.9398,1,60],["Handbook, 210 lines",1,0.9398,1,60],["Stop hook only",1,0.9398,1,60],["Curated + hook",1,0.9398,1,60]]}],["Rules the folders hint at",{"$k":["label","value","lo","hi","n"],"$r":[["No memory",1,0.8454,1,21],["/init CLAUDE.md",1,0.8454,1,21],["Curated, 11 lines",1,0.8454,1,21],["Raw notes, 60 lines",1,0.8454,1,21],["Dreamed notes",1,0.8454,1,21],["Handbook, 210 lines",1,0.8454,1,21],["Stop hook only",1,0.8454,1,21],["Curated + hook",1,0.8454,1,21]]}],["Team knowledge only",{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.4,0.1982,0.6425,15],["/init CLAUDE.md",0.6667,0.4171,0.8482,15],["Curated, 11 lines",1,0.7961,1,15],["Raw notes, 60 lines",1,0.7961,1,15],["Dreamed notes",1,0.7961,1,15],["Handbook, 210 lines",1,0.7961,1,15],["Stop hook only",0.6667,0.4171,0.8482,15],["Curated + hook",1,0.7961,1,15]]}]]},"Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.",["agent-memory-study"],"\u0001"],["memory-team-knowledge-by-model","Team knowledge followed, Sonnet vs Haiku","Changelog rule and late-fee rate, pooled · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.4,0.1982,0.6425,15],["/init CLAUDE.md",0.6667,0.4171,0.8482,15],["Curated, 11 lines",1,0.7961,1,15],["Raw notes, 60 lines",1,0.7961,1,15],["Dreamed notes",1,0.7961,1,15],["Handbook, 210 lines",1,0.7961,1,15],["Stop hook only",0.6667,0.4171,0.8482,15],["Curated + hook",1,0.7961,1,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0,0,0.2775,10],["/init CLAUDE.md",0.1,0.0179,0.4042,10],["Curated, 11 lines",0.8,0.4902,0.9433,10],["Raw notes, 60 lines",0.6,0.3127,0.8318,10],["Dreamed notes",0.8,0.4902,0.9433,10],["Handbook, 210 lines",0.3,0.1078,0.6032,10],["Stop hook only",0.8,0.4902,0.9433,10],["Curated + hook",1,0.7225,1,10]]}}],"Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.",["agent-memory-study"],"\u0001"],["memory-late-fee","\"Charge our standard late fee\": what Sonnet 5.5 did","The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%","stacked-bar","count","\u0001",{"$k":["name","points"],"$r":[["Used the current rate (1.25%)",{"$k":["label","value","n","highlight"],"$r":[["No memory",0,3,"\u0001"],["/init CLAUDE.md",0,3,"\u0001"],["Curated, 11 lines",3,3,true],["Raw notes, 60 lines",3,3,true],["Dreamed notes",3,3,true],["Handbook, 210 lines",3,3,true],["Stop hook only",0,3,"\u0001"],["Curated + hook",3,3,true]]}],["Asked for the rate, wrote no code",{"$k":["label","value","n"],"$r":[["No memory",3,3],["/init CLAUDE.md",2,3],["Curated, 11 lines",0,3],["Raw notes, 60 lines",0,3],["Dreamed notes",0,3],["Handbook, 210 lines",0,3],["Stop hook only",2,3],["Curated + hook",0,3]]}],["Guessed another rate",{"$k":["label","value","n"],"$r":[["No memory",0,3],["/init CLAUDE.md",1,3],["Curated, 11 lines",0,3],["Raw notes, 60 lines",0,3],["Dreamed notes",0,3],["Handbook, 210 lines",0,3],["Stop hook only",1,3],["Curated + hook",0,3]]}]]},"Late-fee task, 3 sessions per condition. \"Asked\" means the agent searched the repository, found no rate and stopped with a question instead of code.",["agent-memory-study"],"\u0001"],["memory-late-fee-haiku","\"Charge our standard late fee\": what Haiku 4.5 did","The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%","stacked-bar","count","\u0001",[{"name":"Used the current rate (1.25%)","points":{"$k":["label","value","n","highlight"],"$r":[["No memory",0,2,"\u0001"],["/init CLAUDE.md",0,2,"\u0001"],["Curated, 11 lines",2,2,true],["Raw notes, 60 lines",2,2,true],["Dreamed notes",2,2,true],["Handbook, 210 lines",2,2,true],["Stop hook only",0,2,"\u0001"],["Curated + hook",2,2,true]]}},{"name":"Guessed another rate","points":{"$k":["label","value","n"],"$r":[["No memory",2,2],["/init CLAUDE.md",2,2],["Curated, 11 lines",0,2],["Raw notes, 60 lines",0,2],["Dreamed notes",0,2],["Handbook, 210 lines",0,2],["Stop hook only",2,2],["Curated + hook",0,2]]}}],"Late-fee task, 2 sessions per condition. \"Asked\" means the agent searched the repository, found no rate and stopped with a question instead of code.",["agent-memory-study"],"\u0001"],["memory-broken-test-command","A stale README command: who still ran it?","Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals","dot-range","rate","ci95",[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",0.8,0.5481,0.9295,15],["/init CLAUDE.md",0.8667,0.6212,0.9626,15],["Curated, 11 lines",0,0,0.2039,15],["Raw notes, 60 lines",0,0,0.2039,15],["Dreamed notes",0,0,0.2039,15],["Handbook, 210 lines",0,0,0.2039,15],["Stop hook only",0.6,0.3575,0.8018,15],["Curated + hook",0,0,0.2039,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","lo","hi","n"],"$r":[["No memory",1,0.7225,1,10],["/init CLAUDE.md",1,0.7225,1,10],["Curated, 11 lines",0,0,0.2775,10],["Raw notes, 60 lines",1,0.7225,1,10],["Dreamed notes",0.1,0.0179,0.4042,10],["Handbook, 210 lines",0,0,0.2775,10],["Stop hook only",1,0.7225,1,10],["Curated + hook",0,0,0.2775,10]]}}],"The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old \"run npm test\" note and a later correction. The /init file repeats the README.",["agent-memory-study"],"lower"],["memory-input-tokens","What memory costs in context","Median input tokens per session, cache reads included (Claude Sonnet 5.5)","bar","tokens","\u0001",[{"name":"Median input tokens per session","points":{"$k":["label","value","n","highlight"],"$r":[["No memory",78455,15,"\u0001"],["/init CLAUDE.md",82262,15,"\u0001"],["Curated, 11 lines",65045,15,true],["Raw notes, 60 lines",76963,15,"\u0001"],["Dreamed notes",68218,15,"\u0001"],["Handbook, 210 lines",85223,15,"\u0001"],["Stop hook only",125674,15,"\u0001"],["Curated + hook",69224,15,"\u0001"]]}}],"Input = uncached input + cache reads + cache writes over every turn, as the CLI reports it. Most of it is read from the prompt cache. The hook adds turns: each block sends the agent back to work.",["agent-memory-study"],"\u0001"],["memory-cost-per-full-pass","List-price cost per fully correct result (calculation)","Sum of the CLI's cost estimates for a condition, divided by its full passes","grouped-bar","usd","\u0001",[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","n"],"$r":[["No memory",0.1386,9],["/init CLAUDE.md",0.1359,9],["Curated, 11 lines",0.0818,15],["Raw notes, 60 lines",0.1009,14],["Dreamed notes",0.0896,15],["Handbook, 210 lines",0.1006,15],["Stop hook only",0.1278,12],["Curated + hook",0.0843,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","n"],"$r":[["No memory",0.3786,2],["/init CLAUDE.md",0.4317,2],["Curated, 11 lines",0.1095,7],["Raw notes, 60 lines",0.1255,6],["Dreamed notes",0.1153,7],["Handbook, 210 lines",0.2615,3],["Stop hook only",0.1419,8],["Curated + hook",0.096,9]]}}],"Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.",["agent-memory-study"],"\u0001"],["memory-wall-time","Time per session","Median wall time in seconds","grouped-bar","seconds","\u0001",[{"name":"Claude Sonnet 5.5","points":{"$k":["label","value","n"],"$r":[["No memory",18,15],["/init CLAUDE.md",19,15],["Curated, 11 lines",21.9,15],["Raw notes, 60 lines",26.8,15],["Dreamed notes",27.2,15],["Handbook, 210 lines",23.6,15],["Stop hook only",27.3,15],["Curated + hook",22,15]]}},{"name":"Claude Haiku 4.5","points":{"$k":["label","value","n"],"$r":[["No memory",54.2,10],["/init CLAUDE.md",52.9,10],["Curated, 11 lines",51.7,10],["Raw notes, 60 lines",51,10],["Dreamed notes",51.8,10],["Handbook, 210 lines",49.9,10],["Stop hook only",68.5,10],["Curated + hook",52.8,10]]}}],"Up to four sessions ran at a time on one machine. Sessions without memory were often shorter because they stopped to ask or skipped the changelog.",["agent-memory-study"],"\u0001"],["memory-dreaming-scorecard","Dreaming: what one consolidation pass kept and dropped","Fixed-pattern checks on each memory file: 10 current facts, 5 stale notes, 9 transient notes","grouped-bar","count","\u0001",{"$k":["name","points"],"$r":[["Current facts kept (of 10)",{"$k":["label","value"],"$r":[["Raw notes",10],["Dream 1",10],["Dream 2",10],["Dream 3",10],["Curated",10],["/init",5]]}],["Stale notes left (of 5)",{"$k":["label","value"],"$r":[["Raw notes",5],["Dream 1",0],["Dream 2",0],["Dream 3",0],["Curated",0],["/init",1]]}],["Transient notes left (of 9)",{"$k":["label","value"],"$r":[["Raw notes",9],["Dream 1",1],["Dream 2",1],["Dream 3",1],["Curated",0],["/init",0]]}]]},"Each dream is one Claude Sonnet call with no tools over the 56 raw notes. A stale note that the new file marks as replaced does not count as left. The one transient note every dream kept is the current dev-server port.",["agent-memory-study"],"\u0001"]]},"related":["caching-consistency","cli-model-latency-tokens"],"hero":{"statIds":["memory-team-knowledge-none","memory-team-knowledge-curated"]}}}