Does memory help Claude Code? 8 kinds of agent memory, tested
Does project memory make Claude Code do better work, and which kind of memory?
Published · 10 charts · Download the data or a carousel
40%
95% CI 20%–64% · n = 15
6/15 · Team-knowledge checks passed with no memory (Sonnet 5.5)
100%
95% CI 80%–100% · n = 15
15/15 · Team-knowledge checks passed with an 11-line curated file (Sonnet 5.5)
The answer
Claude Sonnet 5.5 followed almost every rule it could see without any memory (95% of code rules, 100% of folder rules), but only 6/15 of the team-knowledge checks; with an 11-line curated file it passed 15/15. When a memory file held the late-fee rate, Sonnet used it 15/15 times, even from raw notes that also held a stale rate; without it, Sonnet stopped and asked 7 of 9 times, while Haiku invented a rate and flagged it in 0/6 sessions. The smaller Haiku 4.5 was misled by messy memory: with the raw notes it ran the stale test command in 10/10 sessions, and after one dreaming pass in 1/10. A Stop hook enforced every code rule but could not carry a fact, and it used 1.6× the input tokens of no memory. Full-pass intervals overlap for most pairs (n = 15 per condition).
Live story
Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.
Does memory help Claude Code? 8 kinds of agent memory, tested
200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.
Transcript
- Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
- Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
- Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
- Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
- The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
- Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- Write down what the repo cannot show. Enforce what a script can check.
Key numbers
200
Claude Code sessions, every one graded (120 Sonnet 5.5, 80 Haiku 4.5)
n = 200
15/15
Late fee right when any memory file held the rate (Sonnet 5.5)
95% CI 80%–100% · n = 15
7/9
Without the rate, sessions that asked for it and wrote no code (Sonnet 5.5)
n = 9
9/9
Without the rate: sessions whose final message said the rate was unknown or a guess (Sonnet 5.5)
n = 9
0/6
Without the rate: sessions whose final message said the rate was unknown or a guess (Haiku 4.5)
n = 6
13/15
Sessions with the /init file that ran the broken test command it copied from the README (Sonnet 5.5)
n = 15
1.6×
Median input tokens per session, Stop hook only vs no memory (calculation, Sonnet 5.5)
n = 15
10/10
Haiku 4.5 with the raw notes: sessions that ran the stale test command
n = 10
1/10
Haiku 4.5 with the dreamed notes: sessions that ran the stale test command
n = 10
$17.49
List-price estimate of every session (calculation, not an invoice)
n = 200
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
- Claude Sonnet 5.5
- Claude Haiku 4.5
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | 95% interval | n |
|---|---|---|---|---|
| No memory | 60% | 20% | Claude Sonnet 5.5: 36%–80%; Claude Haiku 4.5: 5.7%–51% | 15 |
| /init CLAUDE.md | 60% | 20% | Claude Sonnet 5.5: 36%–80%; Claude Haiku 4.5: 5.7%–51% | 15 |
| Curated, 11 lines | 100% | 70% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 40%–89% | 15 |
| Raw notes, 60 lines | 93% | 60% | Claude Sonnet 5.5: 70%–99%; Claude Haiku 4.5: 31%–83% | 15 |
| Dreamed notes | 100% | 70% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 40%–89% | 15 |
| Handbook, 210 lines | 100% | 30% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 11%–60% | 15 |
| Stop hook only | 80% | 80% | Claude Sonnet 5.5: 55%–93%; Claude Haiku 4.5: 49%–94% | 15 |
| Curated + hook | 100% | 90% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 60%–98% | 15 |
8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest /init CLAUDE.md 60% (95% interval 36%–80%, n 15). All intervals overlap. Claude Haiku 4.5: highest Curated + hook 90% (95% interval 60%–98%, n 10). Lowest /init CLAUDE.md 20% (95% interval 5.7%–51%, n 10). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 10–15 per row4 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.
Hidden tests pass and every convention check passes · 95% Wilson intervals
Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.
Source: Agent memory study: 8 kinds of project memory on Claude Code
Rules the code already shows
Rules the folders hint at
Team knowledge only
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Rules the code already shows | Rules the folders hint at | Team knowledge only | 95% interval | n |
|---|---|---|---|---|---|
| No memory | 95% | 100% | 40% | Rules the code already shows: 86%–98%; Rules the folders hint at: 85%–100%; Team knowledge only: 20%–64% | 60 |
| /init CLAUDE.md | 95% | 100% | 67% | Rules the code already shows: 86%–98%; Rules the folders hint at: 85%–100%; Team knowledge only: 42%–85% | 60 |
| Curated, 11 lines | 100% | 100% | 100% | Rules the code already shows: 94%–100%; Rules the folders hint at: 85%–100%; Team knowledge only: 80%–100% | 60 |
| Raw notes, 60 lines | 100% | 100% | 100% | Rules the code already shows: 94%–100%; Rules the folders hint at: 85%–100%; Team knowledge only: 80%–100% | 60 |
| Dreamed notes | 100% | 100% | 100% | Rules the code already shows: 94%–100%; Rules the folders hint at: 85%–100%; Team knowledge only: 80%–100% | 60 |
| Handbook, 210 lines | 100% | 100% | 100% | Rules the code already shows: 94%–100%; Rules the folders hint at: 85%–100%; Team knowledge only: 80%–100% | 60 |
| Stop hook only | 100% | 100% | 67% | Rules the code already shows: 94%–100%; Rules the folders hint at: 85%–100%; Team knowledge only: 42%–85% | 60 |
| Curated + hook | 100% | 100% | 100% | Rules the code already shows: 94%–100%; Rules the folders hint at: 85%–100%; Team knowledge only: 80%–100% | 60 |
8 rows, 3 series: Rules the code already shows, Rules the folders hint at, Team knowledge only. Rules the code already shows: highest Curated, 11 lines 100% (95% interval 94%–100%, n 60). Lowest /init CLAUDE.md 95% (95% interval 86%–98%, n 60). All intervals overlap. Rules the folders hint at: all at 100%.
NotesWhiskers: 95% Wilson intervaln 15–60 per row6 of 8 (Rules the code already shows) at 100%: this task set cannot separate them.
Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)
Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.
Source: Agent memory study: 8 kinds of project memory on Claude Code
- Claude Sonnet 5.5
- Claude Haiku 4.5
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | 95% interval | n |
|---|---|---|---|---|
| No memory | 40% | 0% | Claude Sonnet 5.5: 20%–64%; Claude Haiku 4.5: 0%–28% | 15 |
| /init CLAUDE.md | 67% | 10% | Claude Sonnet 5.5: 42%–85%; Claude Haiku 4.5: 1.8%–40% | 15 |
| Curated, 11 lines | 100% | 80% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 49%–94% | 15 |
| Raw notes, 60 lines | 100% | 60% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 31%–83% | 15 |
| Dreamed notes | 100% | 80% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 49%–94% | 15 |
| Handbook, 210 lines | 100% | 30% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 11%–60% | 15 |
| Stop hook only | 67% | 80% | Claude Sonnet 5.5: 42%–85%; Claude Haiku 4.5: 49%–94% | 15 |
| Curated + hook | 100% | 100% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 72%–100% | 15 |
8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest No memory 40% (95% interval 20%–64%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest Curated + hook 100% (95% interval 72%–100%, n 10). Lowest No memory 0% (95% interval 0%–28%, n 10). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 10–15 per row5 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.
Changelog rule and late-fee rate, pooled · 95% Wilson intervals
Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.
Source: Agent memory study: 8 kinds of project memory on Claude Code
- Used the current rate (1.25%)
- Asked for the rate, wrote no code
- Guessed another rate
One square per count; counts at the right are exact and in legend order.
| Item | Used the current rate (1.25%) | Asked for the rate, wrote no code | Guessed another rate | n |
|---|---|---|---|---|
| No memory | 0 | 3 | 0 | 3 |
| /init CLAUDE.md | 0 | 2 | 1 | 3 |
| Curated, 11 lines | 3 | 0 | 0 | 3 |
| Raw notes, 60 lines | 3 | 0 | 0 | 3 |
| Dreamed notes | 3 | 0 | 0 | 3 |
| Handbook, 210 lines | 3 | 0 | 0 | 3 |
| Stop hook only | 0 | 2 | 1 | 3 |
| Curated + hook | 3 | 0 | 0 | 3 |
8 rows, 3 series: Used the current rate (1.25%), Asked for the rate, wrote no code, Guessed another rate. Used the current rate (1.25%): highest Curated, 11 lines 3 (n 3). Lowest Stop hook only 0 (n 3). Asked for the rate, wrote no code: highest No memory 3 (n 3). Lowest Curated + hook 0 (n 3).
Notesn = 3 per row
The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%
Late-fee task, 3 sessions per condition. "Asked" means the agent searched the repository, found no rate and stopped with a question instead of code.
Source: Agent memory study: 8 kinds of project memory on Claude Code
- Used the current rate (1.25%)
- Guessed another rate
One square per count; counts at the right are exact and in legend order.
| Item | Used the current rate (1.25%) | Guessed another rate | n |
|---|---|---|---|
| No memory | 0 | 2 | 2 |
| /init CLAUDE.md | 0 | 2 | 2 |
| Curated, 11 lines | 2 | 0 | 2 |
| Raw notes, 60 lines | 2 | 0 | 2 |
| Dreamed notes | 2 | 0 | 2 |
| Handbook, 210 lines | 2 | 0 | 2 |
| Stop hook only | 0 | 2 | 2 |
| Curated + hook | 2 | 0 | 2 |
8 rows, 2 series: Used the current rate (1.25%), Guessed another rate. Used the current rate (1.25%): highest Curated, 11 lines 2 (n 2). Lowest Stop hook only 0 (n 2). Guessed another rate: highest No memory 2 (n 2). Lowest Curated + hook 0 (n 2).
Notesn = 2 per row
The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%
Late-fee task, 2 sessions per condition. "Asked" means the agent searched the repository, found no rate and stopped with a question instead of code.
Source: Agent memory study: 8 kinds of project memory on Claude Code
- Claude Sonnet 5.5
- Claude Haiku 4.5
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | 95% interval | n |
|---|---|---|---|---|
| No memory | 80% | 100% | Claude Sonnet 5.5: 55%–93%; Claude Haiku 4.5: 72%–100% | 15 |
| /init CLAUDE.md | 87% | 100% | Claude Sonnet 5.5: 62%–96%; Claude Haiku 4.5: 72%–100% | 15 |
| Curated, 11 lines | 0% | 0% | Claude Sonnet 5.5: 0%–20%; Claude Haiku 4.5: 0%–28% | 15 |
| Raw notes, 60 lines | 0% | 100% | Claude Sonnet 5.5: 0%–20%; Claude Haiku 4.5: 72%–100% | 15 |
| Dreamed notes | 0% | 10% | Claude Sonnet 5.5: 0%–20%; Claude Haiku 4.5: 1.8%–40% | 15 |
| Handbook, 210 lines | 0% | 0% | Claude Sonnet 5.5: 0%–20%; Claude Haiku 4.5: 0%–28% | 15 |
| Stop hook only | 60% | 100% | Claude Sonnet 5.5: 36%–80%; Claude Haiku 4.5: 72%–100% | 15 |
| Curated + hook | 0% | 0% | Claude Sonnet 5.5: 0%–20%; Claude Haiku 4.5: 0%–28% | 15 |
8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest /init CLAUDE.md 87% (95% interval 62%–96%, n 15). Lowest Curated + hook 0% (95% interval 0%–20%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest No memory 100% (95% interval 72%–100%, n 10). Lowest Curated + hook 0% (95% interval 0%–28%, n 10). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 10–15 per row
Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals
The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.
Source: Agent memory study: 8 kinds of project memory on Claude Code
| Item | Median input tokens per session | n |
|---|---|---|
| No memory | 78,455 | 15 |
| /init CLAUDE.md | 82,262 | 15 |
| Curated, 11 lines | 65,045 | 15 |
| Raw notes, 60 lines | 76,963 | 15 |
| Dreamed notes | 68,218 | 15 |
| Handbook, 210 lines | 85,223 | 15 |
| Stop hook only | 125,674 | 15 |
| Curated + hook | 69,224 | 15 |
8 rows. Highest Stop hook only 125,674 (n 15). Lowest Curated, 11 lines 65,045 (n 15).
Notesn = 15 per row
Median input tokens per session, cache reads included (Claude Sonnet 5.5)
Input = uncached input + cache reads + cache writes over every turn, as the CLI reports it. Most of it is read from the prompt cache. The hook adds turns: each block sends the agent back to work.
Source: Agent memory study: 8 kinds of project memory on Claude Code
- Claude Sonnet 5.5
- Claude Haiku 4.5 (square)
Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | n |
|---|---|---|---|
| No memory | $0.14 | $0.38 | 9 |
| /init CLAUDE.md | $0.14 | $0.43 | 9 |
| Curated, 11 lines | $0.082 | $0.11 | 15 |
| Raw notes, 60 lines | $0.1 | $0.13 | 14 |
| Dreamed notes | $0.09 | $0.12 | 15 |
| Handbook, 210 lines | $0.1 | $0.26 | 15 |
| Stop hook only | $0.13 | $0.14 | 12 |
| Curated + hook | $0.084 | $0.096 | 15 |
List-price calculation, not a run. 8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest No memory $0.14 (n 9). Lowest Curated, 11 lines $0.082 (n 15). Claude Haiku 4.5: highest /init CLAUDE.md $0.43 (n 2). Lowest Curated + hook $0.096 (n 9).
Notesn 2–15 per row
Sum of the CLI's cost estimates for a condition, divided by its full passes
Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.
Source: Agent memory study: 8 kinds of project memory on Claude Code
Time per session
Median wall time in seconds
- Claude Sonnet 5.5
- Claude Haiku 4.5 (square)
Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | n |
|---|---|---|---|
| No memory | 18 s | 54.2 s | 15 |
| /init CLAUDE.md | 19 s | 52.9 s | 15 |
| Curated, 11 lines | 21.9 s | 51.7 s | 15 |
| Raw notes, 60 lines | 26.8 s | 51 s | 15 |
| Dreamed notes | 27.2 s | 51.8 s | 15 |
| Handbook, 210 lines | 23.6 s | 49.9 s | 15 |
| Stop hook only | 27.3 s | 68.5 s | 15 |
| Curated + hook | 22 s | 52.8 s | 15 |
8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: slowest Stop hook only 27.3 s (n 15). Fastest No memory 18 s (n 15). Claude Haiku 4.5: slowest Stop hook only 68.5 s (n 10). Fastest Handbook, 210 lines 49.9 s (n 10).
Notesn 10–15 per row
Up to four sessions ran at a time on one machine. Sessions without memory were often shorter because they stopped to ask or skipped the changelog.
Source: Agent memory study: 8 kinds of project memory on Claude Code
Current facts kept (of 10)
Stale notes left (of 5)
Transient notes left (of 9)
One panel per series, all on the same axis.
| Item | Current facts kept (of 10) | Stale notes left (of 5) | Transient notes left (of 9) |
|---|---|---|---|
| Raw notes | 10 | 5 | 9 |
| Dream 1 | 10 | 0 | 1 |
| Dream 2 | 10 | 0 | 1 |
| Dream 3 | 10 | 0 | 1 |
| Curated | 10 | 0 | 0 |
| /init | 5 | 1 | 0 |
6 rows, 3 series: Current facts kept (of 10), Stale notes left (of 5), Transient notes left (of 9). Current facts kept (of 10): highest Raw notes 10. Lowest /init 5. Stale notes left (of 5): highest Raw notes 5. Lowest Curated 0.
Notes
Fixed-pattern checks on each memory file: 10 current facts, 5 stale notes, 9 transient notes
Each dream is one Claude Sonnet call with no tools over the 56 raw notes. A stale note that the new file marks as replaced does not count as left. The one transient note every dream kept is the current dev-server port.
Source: Agent memory study: 8 kinds of project memory on Claude Code
Tables
Every condition, Claude Sonnet 5.5
Memory by column: each cell is k of n from the table
- Full pass
- Hidden tests pass
- Code-visible rules
- Folder-visible rules
- Team knowledge
- Ran the stale test command
Shade is the share k of n in each cell. Totals and other columns are in the Table view.
| Memory | Full pass | Hidden tests pass | Code-visible rules | Folder-visible rules | Team knowledge | Ran the stale test command | Median input tokens | Median time (s) | Hook blocks |
|---|---|---|---|---|---|---|---|---|---|
| No memory | 9/15 | 12/15 | 57/60 | 21/21 | 6/15 | 12/15 | 78,455 | 18 s | 0 |
| /init CLAUDE.md | 9/15 | 12/15 | 57/60 | 21/21 | 10/15 | 13/15 | 82,262 | 19 s | 0 |
| Curated, 11 lines | 15/15 | 15/15 | 60/60 | 21/21 | 15/15 | 0/15 | 65,045 | 21.9 s | 0 |
| Raw notes, 60 lines | 14/15 | 14/15 | 60/60 | 21/21 | 15/15 | 0/15 | 76,963 | 26.8 s | 0 |
| Dreamed notes | 15/15 | 15/15 | 60/60 | 21/21 | 15/15 | 0/15 | 68,218 | 27.2 s | 0 |
| Handbook, 210 lines | 15/15 | 15/15 | 60/60 | 21/21 | 15/15 | 0/15 | 85,223 | 23.6 s | 0 |
| Stop hook only | 12/15 | 12/15 | 60/60 | 21/21 | 10/15 | 9/15 | 125,674 | 27.3 s | 5 |
| Curated + hook | 15/15 | 15/15 | 60/60 | 21/21 | 15/15 | 0/15 | 69,224 | 22 s | 0 |
Counts from the table, not an intervaln = 15–60 per cell
8 rows by 6 columns. 27 of 48 k/n cells are full (n of n) and 5 are zero.
Every condition, Claude Haiku 4.5
Memory by column: each cell is k of n from the table
- Full pass
- Hidden tests pass
- Code-visible rules
- Folder-visible rules
- Team knowledge
- Ran the stale test command
Shade is the share k of n in each cell. Totals and other columns are in the Table view.
| Memory | Full pass | Hidden tests pass | Code-visible rules | Folder-visible rules | Team knowledge | Ran the stale test command | Median input tokens | Median time (s) | Hook blocks |
|---|---|---|---|---|---|---|---|---|---|
| No memory | 2/10 | 6/10 | 40/40 | 12/14 | 0/10 | 10/10 | 301,233 | 54.2 s | 0 |
| /init CLAUDE.md | 2/10 | 7/10 | 38/40 | 14/14 | 1/10 | 10/10 | 323,519 | 52.9 s | 0 |
| Curated, 11 lines | 7/10 | 9/10 | 40/40 | 14/14 | 8/10 | 0/10 | 284,212 | 51.7 s | 0 |
| Raw notes, 60 lines | 6/10 | 10/10 | 40/40 | 14/14 | 6/10 | 10/10 | 285,263 | 51 s | 0 |
| Dreamed notes | 7/10 | 9/10 | 40/40 | 14/14 | 8/10 | 1/10 | 267,405 | 51.8 s | 0 |
| Handbook, 210 lines | 3/10 | 10/10 | 40/40 | 14/14 | 3/10 | 0/10 | 313,796 | 49.9 s | 0 |
| Stop hook only | 8/10 | 8/10 | 40/40 | 14/14 | 8/10 | 10/10 | 530,935 | 68.5 s | 10 |
| Curated + hook | 9/10 | 9/10 | 40/40 | 14/14 | 10/10 | 0/10 | 340,491 | 52.8 s | 3 |
Counts from the table, not an intervaln = 10–40 per cell
8 rows by 6 columns. 21 of 48 k/n cells are full (n of n) and 4 are zero.
The 56 raw notes, labelled by what a consolidation pass should do with each (paths in the test repository shown under tally/)
| Saved | Note | Should be |
|---|---|---|
| 2026-06-12 | Project is `tally`, a small invoicing ledger for the billing service. Plain ESM, no dependencies. | kept |
| 2026-06-12 | Shah prefers small PRs with one change each. | kept |
| 2026-06-14 | Dev server for the billing service ran on port 3001 this afternoon. | removed: transient |
| 2026-06-18 | Run tests with `npm test`. | removed: stale |
| 2026-06-20 | Amounts are dollars as floats. Format them with `toFixed(2)`. | removed: stale |
| 2026-06-20 | Use 2-space indentation and single quotes. | kept |
| 2026-06-22 | Branch `feat/customers-v1` pushed; waiting for review. | removed: transient |
| 2026-06-25 | The time helper is `now()` in `tally/src/clock.js`. | removed: stale |
| 2026-06-27 | Shah is out on Friday; do not merge anything then. | removed: transient |
| 2026-07-01 | Customers live in `tally/src/customers.js`. | kept |
| 2026-07-03 | Throw `TallyError` with a code from `CODES` in `tally/src/errors.js`; the billing service maps codes to HTTP statuses. | kept |
| 2026-07-03 | Error codes are public API. Do not rename an existing code. | kept |
| 2026-07-05 | Prefer named exports. | kept |
| 2026-07-08 | Used jq to inspect the payment fixtures; worked fine. | removed: transient |
| 2026-07-10 | Do not use `console.log` in `tally/src/`; use `log()` from `tally/src/log.js`. | kept |
| 2026-07-12 | Staging database was reset today. | removed: transient |
| 2026-07-14 | Late fee is 2% of the unpaid balance. | removed: stale |
| 2026-07-15 | PR #212 (partial payments) merged. | removed: transient |
| 2026-07-18 | The report module is new; Shah wants it in the next release. | kept |
| 2026-07-22 | Shah wants a line under `## Unreleased` in CHANGELOG.md for every user-visible change. | kept |
| 2026-07-22 | Keep CHANGELOG entries to one line. | kept |
| 2026-07-25 | `slugify()` is used for export file names. | kept |
| 2026-07-29 | Remember: no `console.log` in src, use the log helper. | merged: repeat |
| 2026-08-01 | Released 1.3.0 (monthly report, partial payments). | removed: transient |
| 2026-08-03 | CI was slow today because of a runner outage. | removed: transient |
| 2026-08-05 | Incident: someone edited migration 0002 and production schemas drifted. Migrations are append-only now: add a new `tally/migrations/NNNN_<name>.json`, never edit an old one. | kept |
| 2026-08-07 | Store loads the schema from `tally/migrations/*.json` in file-name order. | kept |
| 2026-08-09 | Shah prefers early returns over nested ifs. | kept |
| 2026-08-12 | Migrated all money to integer cents (`amountCents`). Parse input with `parseAmount()`, print with `formatCents()`. No floats for money. | kept |
| 2026-08-12 | `toFixed` caused rounding bugs in the old float code. | kept |
| 2026-08-14 | Reminder: throw TallyError, not Error. | merged: repeat |
| 2026-08-16 | Coffee machine on floor 3 is broken (Shah mentioned it). | removed: transient |
| 2026-08-19 | CI flaky again; reran twice. | removed: transient |
| 2026-08-22 | Payment amounts must not exceed the open balance (`E_OVERPAYMENT`). | kept |
| 2026-08-24 | Branch `fix/report-totals` pushed. | removed: transient |
| 2026-08-27 | Integer cents everywhere. Do not divide by 100 except when formatting. | merged: repeat |
| 2026-08-30 | Renamed the time helper: use `nowIso()` from `tally/src/time.js`. Tests inject a clock with `setClock()`. Do not call `new Date()` or `Date.now()` in src. | kept |
| 2026-09-02 | Throw TallyError with a CODES entry; add a new code when none fits. | kept |
| 2026-09-04 | Shah likes tests next to each feature change. | kept |
| 2026-09-06 | Looked at Stripe's refund API for ideas; not needed yet. | removed: transient |
| 2026-09-08 | Port 3001 is taken by another service now; billing dev server moved to 3002. | removed: transient |
| 2026-09-10 | Lost a fix because I edited `tally/src/report.gen.js` directly; it is generated. Edit `tally/templates/report.tmpl.js` and run `node tally/scripts/gen-report.mjs`. | kept |
| 2026-09-11 | The report template uses a `__CURRENCY__` placeholder. | kept |
| 2026-09-13 | Shah asked for clearer error messages in validation. | kept |
| 2026-09-15 | Finance changed the late fee to 1.25% of the unpaid balance, effective now. Round half up to the cent. | kept |
| 2026-09-16 | Email validation added to customers. | kept |
| 2026-09-18 | Reminder: changelog line for every user-visible change. | merged: repeat |
| 2026-09-20 | Released 1.4.0 (integer cents, email validation). | removed: transient |
| 2026-09-21 | `node --test test/` fails on Node 25 (it treats the folder as a file). Run `node --test` with no path. `npm test` uses the broken form. | kept |
| 2026-09-23 | Prefer `Object.freeze` for constant maps. | kept |
| 2026-09-24 | Do not edit migration files that already exist. | merged: repeat |
| 2026-09-26 | Shah reviews PRs in the morning. | kept |
| 2026-09-28 | Use the log helper for audit events (payments, refunds). | kept |
| 2026-09-30 | The billing service reads `CODES` at start-up; new codes need no other change. | kept |
| 2026-10-01 | Staging is on the new cluster. | removed: transient |
| 2026-10-02 | Generated file again: never hand-edit `tally/src/report.gen.js`. | merged: repeat |
Full passes per task (Claude Sonnet 5.5, of 3)
| Task | No memory | /init CLAUDE.md | Curated, 11 lines | Raw notes, 60 lines | Dreamed notes | Handbook, 210 lines | Stop hook only | Curated + hook |
|---|---|---|---|---|---|---|---|---|
| Refunds | 3 | 3 | 3 | 2 | 3 | 3 | 3 | 3 |
| Report decimals | 0 | 0 | 3 | 3 | 3 | 3 | 3 | 3 |
| Customer phone | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 |
| Late fee | 0 | 0 | 3 | 3 | 3 | 3 | 0 | 3 |
| Slugify (control) | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 |
Method
- One small Node.js repository with real conventions: money in integer cents, coded errors, a clock helper, a log helper, append-only migrations, a generated report, a changelog rule and a late-fee rate that exists only in memory.
- Five tasks: refunds, a report formatting bug, an optional phone field, a late fee "at our standard rate", and a control task where memory does not matter.
- Eight conditions: no memory, the real /init output, a curated 11-line file, 56 raw dated notes (repeats, one-off events and notes that a later note contradicts), the same notes after one "dreaming" pass, the curated facts inside a 210-line handbook, a Stop hook that blocks finishing while a rule is broken, and curated plus hook.
- Claude Code 2.1.286 headless; Sonnet 3 repetitions per cell, Haiku 2; tools Bash, Read, Edit, Write, Glob, Grep; OS sandbox; the tester's own instruction files excluded and auto memory off (both checked with a probe).
- Grading after the session: hidden tests run in a sandbox, plus deterministic checks on the lines the agent added. Full pass = tests and every applicable check. Every attempt counts.
- The protocol was declared before the first counted session. One amendment, made before any Haiku session, added the Haiku lane. One erratum: the raw notes file holds 56 notes, not 58 as first written. The frozen handbook has 210 lines, not 211 as the protocol first stated. Measurements did not change.
Caveats
- One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
- The hook checks the same code rules as the grader. It shows what rules written as code can do; it cannot carry a fact such as the late-fee rate.
- n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- Costs are the CLI's list-price estimates for subscription sessions, not invoices.
- In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
Sources
Agent memory study: 8 kinds of project memory on Claude Code
One small Node.js repository, 5 tasks, 8 memory conditions (none, /init, curated, raw notes, dreamed notes, long handbook, Stop hook, curated + hook). Claude Code 2.1.286 headless: Sonnet lane 3 repetitions, Haiku lane 2. Hidden tests and deterministic convention checks; protocol declared before the first session; every attempt kept. The exact memory files and three dreaming passes are published.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Does memory help Claude Code? 8 kinds of agent memory, tested”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/agent-memory.
Explainers that cite this study
Read the methods and terms in the context of these recorded results.
- Memory consolidation (dreaming) for AI agents, measured
- Benchmark saturation: when every model scores 100%
- Claude Code hooks, explained with a measured Stop hook
- What is CLAUDE.md? Agent memory files, measured
- Cost per correct answer: the LLM price that counts failures
- How many runs does an LLM eval need? Sample size, with real intervals
- What is context engineering? What to put in the context, measured
Models and comparisons in this study
Write-ups on this study
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
AI coding cost per developer: a formula built on recorded work
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
Claude Code cost per task: a price ladder from one decision to one agent run
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.
Claude Code Stop hook, tested: it enforced the rules, used 1.6x the median input and lacked a fact
Claude Code Stop hook: 60/60 Sonnet code-rule checks (95% interval 94–100%). Median input was 1.6x no memory (calculation). Limits and raw data.
Claude Haiku 4.5 vs Sonnet 5.5: all 80 comparison rows, and where the small model loses
Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.
How long does an AI coding agent take per task? Minutes, calls and where the time goes
Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.
How many runs do you need to compare two AI models? A sample-size table
Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.
Is Claude Haiku cheaper than Sonnet? Cost per correct answer, with retries
Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.
Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
The cheapest way to run an AI coding agent: 7 levers from measured runs
7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.
Which Claude model should you use? A task-by-task guide from our measurements
Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.
Why AI coding agents fail on real pull requests: every failure from 12 attempts
32 AI coding agent misses from our own runs, sorted into 8 classes: wrong answers, gates, caps, lost context and more. Counts, not rates.
Your AI agent says it is done. Is it? Claimed vs verified in our runs
6 of 6 Haiku 4.5 sessions invented a late-fee rate and reported done. Sonnet 5.5 flagged the gap in 9 of 9. Claimed vs verified, with small n.
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
Does CLAUDE.md help? We tested 8 kinds of agent memory on Claude Code
200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a handbook and a Stop hook. Memory mattered where the repo was silent.
More studies
All benchmarksDoes a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.
Does a new Claude Code session reuse the prompt cache of an earlier one?
30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.