• Claude Code
  • Agent Memory
  • CLAUDE.md
  • Hooks
  • Context Engineering

Does memory help Claude Code? 8 kinds of agent memory, tested

Does project memory make Claude Code do better work, and which kind of memory?

Published · 10 charts · Download the data or a carousel

40%

95% CI 20%–64% · n = 15

6/15 · Team-knowledge checks passed with no memory (Sonnet 5.5)

100%

95% CI 80%–100% · n = 15

15/15 · Team-knowledge checks passed with an 11-line curated file (Sonnet 5.5)

The answer

Claude Sonnet 5.5 followed almost every rule it could see without any memory (95% of code rules, 100% of folder rules), but only 6/15 of the team-knowledge checks; with an 11-line curated file it passed 15/15. When a memory file held the late-fee rate, Sonnet used it 15/15 times, even from raw notes that also held a stale rate; without it, Sonnet stopped and asked 7 of 9 times, while Haiku invented a rate and flagged it in 0/6 sessions. The smaller Haiku 4.5 was misled by messy memory: with the raw notes it ran the stale test command in 10/10 sessions, and after one dreaming pass in 1/10. A Stop hook enforced every code rule but could not carry a fact, and it used 1.6× the input tokens of no memory. Full-pass intervals overlap for most pairs (n = 15 per condition).

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 62 sDoes memory help Claude Code? 8 kinds of agent memory, tested

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.

Transcript
  1. Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
  2. Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  3. Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  4. With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
  5. Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
  6. The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  7. The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
  8. Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  9. Write down what the repo cannot show. Enforce what a script can check.

Key numbers

200

Claude Code sessions, every one graded (120 Sonnet 5.5, 80 Haiku 4.5)

n = 200

15/15

Late fee right when any memory file held the rate (Sonnet 5.5)

95% CI 80%–100% · n = 15

7/9

Without the rate, sessions that asked for it and wrote no code (Sonnet 5.5)

n = 9

9/9

Without the rate: sessions whose final message said the rate was unknown or a guess (Sonnet 5.5)

n = 9

0/6

Without the rate: sessions whose final message said the rate was unknown or a guess (Haiku 4.5)

n = 6

13/15

Sessions with the /init file that ran the broken test command it copied from the README (Sonnet 5.5)

n = 15

1.6×

Median input tokens per session, Stop hook only vs no memory (calculation, Sonnet 5.5)

n = 15

10/10

Haiku 4.5 with the raw notes: sessions that ran the stale test command

n = 10

1/10

Haiku 4.5 with the dreamed notes: sessions that ran the stale test command

n = 10

$17.49

List-price estimate of every session (calculation, not an invoice)

n = 200

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest /init CLAUDE.md 60% (95% interval 36%–80%, n 15). All intervals overlap. Claude Haiku 4.5: highest Curated + hook 90% (95% interval 60%–98%, n 10). Lowest /init CLAUDE.md 20% (95% interval 5.7%–51%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row4 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Hidden tests pass and every convention check passes · 95% Wilson intervals

Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)

Rules the code already shows

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Rules the folders hint at

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Team knowledge only

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 3 series: Rules the code already shows, Rules the folders hint at, Team knowledge only. Rules the code already shows: highest Curated, 11 lines 100% (95% interval 94%–100%, n 60). Lowest /init CLAUDE.md 95% (95% interval 86%–98%, n 60). All intervals overlap. Rules the folders hint at: all at 100%.

NotesWhiskers: 95% Wilson intervaln 15–60 per row6 of 8 (Rules the code already shows) at 100%: this task set cannot separate them.

Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)
  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest No memory 40% (95% interval 20%–64%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest Curated + hook 100% (95% interval 72%–100%, n 10). Lowest No memory 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row5 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Changelog rule and late-fee rate, pooled · 95% Wilson intervals

Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)
  • Used the current rate (1.25%)
  • Asked for the rate, wrote no code
  • Guessed another rate
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One square per count; counts at the right are exact and in legend order.

8 rows, 3 series: Used the current rate (1.25%), Asked for the rate, wrote no code, Guessed another rate. Used the current rate (1.25%): highest Curated, 11 lines 3 (n 3). Lowest Stop hook only 0 (n 3). Asked for the rate, wrote no code: highest No memory 3 (n 3). Lowest Curated + hook 0 (n 3).

Notesn = 3 per row

The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%

Late-fee task, 3 sessions per condition. "Asked" means the agent searched the repository, found no rate and stopped with a question instead of code.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)
  • Used the current rate (1.25%)
  • Guessed another rate
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One square per count; counts at the right are exact and in legend order.

8 rows, 2 series: Used the current rate (1.25%), Guessed another rate. Used the current rate (1.25%): highest Curated, 11 lines 2 (n 2). Lowest Stop hook only 0 (n 2). Guessed another rate: highest No memory 2 (n 2). Lowest Curated + hook 0 (n 2).

Notesn = 2 per row

The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%

Late-fee task, 2 sessions per condition. "Asked" means the agent searched the repository, found no rate and stopped with a question instead of code.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)
  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest /init CLAUDE.md 87% (95% interval 62%–96%, n 15). Lowest Curated + hook 0% (95% interval 0%–20%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest No memory 100% (95% interval 72%–100%, n 10). Lowest Curated + hook 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row

Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals

The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows. Highest Stop hook only 125,674 (n 15). Lowest Curated, 11 lines 65,045 (n 15).

Notesn = 15 per row

Median input tokens per session, cache reads included (Claude Sonnet 5.5)

Input = uncached input + cache reads + cache writes over every turn, as the CLI reports it. Most of it is read from the prompt cache. The hook adds turns: each block sends the agent back to work.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)
Calculation
  • Claude Sonnet 5.5
  • Claude Haiku 4.5 (square)
Sorted by gap, largest first.
/init CLAUDE.md
No memory
Handbook, 210 lines
Curated, 11 lines
Dreamed notes
Raw notes, 60 lines
Curated + hook
Stop hook only

Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).

List-price calculation, not a run. 8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest No memory $0.14 (n 9). Lowest Curated, 11 lines $0.082 (n 15). Claude Haiku 4.5: highest /init CLAUDE.md $0.43 (n 2). Lowest Curated + hook $0.096 (n 9).

Notesn 2–15 per row

Sum of the CLI's cost estimates for a condition, divided by its full passes

Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)

Time per session

Median wall time in seconds

  • Claude Sonnet 5.5
  • Claude Haiku 4.5 (square)
Sorted by gap, largest first.
No memory
/init CLAUDE.md
Stop hook only
Curated + hook
Curated, 11 lines
Handbook, 210 lines
Dreamed notes
Raw notes, 60 lines

Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: slowest Stop hook only 27.3 s (n 15). Fastest No memory 18 s (n 15). Claude Haiku 4.5: slowest Stop hook only 68.5 s (n 10). Fastest Handbook, 210 lines 49.9 s (n 10).

Notesn 10–15 per row

Up to four sessions ran at a time on one machine. Sessions without memory were often shorter because they stopped to ask or skipped the changelog.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)

Current facts kept (of 10)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

Stale notes left (of 5)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

Transient notes left (of 9)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

One panel per series, all on the same axis.

6 rows, 3 series: Current facts kept (of 10), Stale notes left (of 5), Transient notes left (of 9). Current facts kept (of 10): highest Raw notes 10. Lowest /init 5. Stale notes left (of 5): highest Raw notes 5. Lowest Curated 0.

Notes

Fixed-pattern checks on each memory file: 10 current facts, 5 stale notes, 9 transient notes

Each dream is one Claude Sonnet call with no tools over the 56 raw notes. A stale note that the new file marks as replaced does not count as left. The one transient note every dream kept is the current dev-server port.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Share card (PNG)

Tables

Every condition, Claude Sonnet 5.5

Memory by column: each cell is k of n from the table

MemoryFull pass1Hidden tests pass2Code-visible rules3Folder-visible rules4Team knowledge5Ran the stale test command6
No memory9/1512/1557/6021/216/1512/15
/init CLAUDE.md9/1512/1557/6021/2110/1513/15
Curated, 11 lines15/1515/1560/6021/2115/150/15
Raw notes, 60 lines14/1514/1560/6021/2115/150/15
Dreamed notes15/1515/1560/6021/2115/150/15
Handbook, 210 lines15/1515/1560/6021/2115/150/15
Stop hook only12/1512/1560/6021/2110/159/15
Curated + hook15/1515/1560/6021/2115/150/15
  1. Full pass
  2. Hidden tests pass
  3. Code-visible rules
  4. Folder-visible rules
  5. Team knowledge
  6. Ran the stale test command

Shade is the share k of n in each cell. Totals and other columns are in the Table view.

Counts from the table, not an intervaln = 15–60 per cell

8 rows by 6 columns. 27 of 48 k/n cells are full (n of n) and 5 are zero.

Every condition, Claude Haiku 4.5

Memory by column: each cell is k of n from the table

MemoryFull pass1Hidden tests pass2Code-visible rules3Folder-visible rules4Team knowledge5Ran the stale test command6
No memory2/106/1040/4012/140/1010/10
/init CLAUDE.md2/107/1038/4014/141/1010/10
Curated, 11 lines7/109/1040/4014/148/100/10
Raw notes, 60 lines6/1010/1040/4014/146/1010/10
Dreamed notes7/109/1040/4014/148/101/10
Handbook, 210 lines3/1010/1040/4014/143/100/10
Stop hook only8/108/1040/4014/148/1010/10
Curated + hook9/109/1040/4014/1410/100/10
  1. Full pass
  2. Hidden tests pass
  3. Code-visible rules
  4. Folder-visible rules
  5. Team knowledge
  6. Ran the stale test command

Shade is the share k of n in each cell. Totals and other columns are in the Table view.

Counts from the table, not an intervaln = 10–40 per cell

8 rows by 6 columns. 21 of 48 k/n cells are full (n of n) and 4 are zero.

The 56 raw notes, labelled by what a consolidation pass should do with each (paths in the test repository shown under tally/)

SavedNoteShould be
2026-06-12Project is `tally`, a small invoicing ledger for the billing service. Plain ESM, no dependencies.kept
2026-06-12Shah prefers small PRs with one change each.kept
2026-06-14Dev server for the billing service ran on port 3001 this afternoon.removed: transient
2026-06-18Run tests with `npm test`.removed: stale
2026-06-20Amounts are dollars as floats. Format them with `toFixed(2)`.removed: stale
2026-06-20Use 2-space indentation and single quotes.kept
2026-06-22Branch `feat/customers-v1` pushed; waiting for review.removed: transient
2026-06-25The time helper is `now()` in `tally/src/clock.js`.removed: stale
2026-06-27Shah is out on Friday; do not merge anything then.removed: transient
2026-07-01Customers live in `tally/src/customers.js`.kept
2026-07-03Throw `TallyError` with a code from `CODES` in `tally/src/errors.js`; the billing service maps codes to HTTP statuses.kept
2026-07-03Error codes are public API. Do not rename an existing code.kept
2026-07-05Prefer named exports.kept
2026-07-08Used jq to inspect the payment fixtures; worked fine.removed: transient
2026-07-10Do not use `console.log` in `tally/src/`; use `log()` from `tally/src/log.js`.kept
2026-07-12Staging database was reset today.removed: transient
2026-07-14Late fee is 2% of the unpaid balance.removed: stale
2026-07-15PR #212 (partial payments) merged.removed: transient
2026-07-18The report module is new; Shah wants it in the next release.kept
2026-07-22Shah wants a line under `## Unreleased` in CHANGELOG.md for every user-visible change.kept
2026-07-22Keep CHANGELOG entries to one line.kept
2026-07-25`slugify()` is used for export file names.kept
2026-07-29Remember: no `console.log` in src, use the log helper.merged: repeat
2026-08-01Released 1.3.0 (monthly report, partial payments).removed: transient
2026-08-03CI was slow today because of a runner outage.removed: transient
2026-08-05Incident: someone edited migration 0002 and production schemas drifted. Migrations are append-only now: add a new `tally/migrations/NNNN_<name>.json`, never edit an old one.kept
2026-08-07Store loads the schema from `tally/migrations/*.json` in file-name order.kept
2026-08-09Shah prefers early returns over nested ifs.kept
2026-08-12Migrated all money to integer cents (`amountCents`). Parse input with `parseAmount()`, print with `formatCents()`. No floats for money.kept
2026-08-12`toFixed` caused rounding bugs in the old float code.kept
2026-08-14Reminder: throw TallyError, not Error.merged: repeat
2026-08-16Coffee machine on floor 3 is broken (Shah mentioned it).removed: transient
2026-08-19CI flaky again; reran twice.removed: transient
2026-08-22Payment amounts must not exceed the open balance (`E_OVERPAYMENT`).kept
2026-08-24Branch `fix/report-totals` pushed.removed: transient
2026-08-27Integer cents everywhere. Do not divide by 100 except when formatting.merged: repeat
2026-08-30Renamed the time helper: use `nowIso()` from `tally/src/time.js`. Tests inject a clock with `setClock()`. Do not call `new Date()` or `Date.now()` in src.kept
2026-09-02Throw TallyError with a CODES entry; add a new code when none fits.kept
2026-09-04Shah likes tests next to each feature change.kept
2026-09-06Looked at Stripe's refund API for ideas; not needed yet.removed: transient
2026-09-08Port 3001 is taken by another service now; billing dev server moved to 3002.removed: transient
2026-09-10Lost a fix because I edited `tally/src/report.gen.js` directly; it is generated. Edit `tally/templates/report.tmpl.js` and run `node tally/scripts/gen-report.mjs`.kept
2026-09-11The report template uses a `__CURRENCY__` placeholder.kept
2026-09-13Shah asked for clearer error messages in validation.kept
2026-09-15Finance changed the late fee to 1.25% of the unpaid balance, effective now. Round half up to the cent.kept
2026-09-16Email validation added to customers.kept
2026-09-18Reminder: changelog line for every user-visible change.merged: repeat
2026-09-20Released 1.4.0 (integer cents, email validation).removed: transient
2026-09-21`node --test test/` fails on Node 25 (it treats the folder as a file). Run `node --test` with no path. `npm test` uses the broken form.kept
2026-09-23Prefer `Object.freeze` for constant maps.kept
2026-09-24Do not edit migration files that already exist.merged: repeat
2026-09-26Shah reviews PRs in the morning.kept
2026-09-28Use the log helper for audit events (payments, refunds).kept
2026-09-30The billing service reads `CODES` at start-up; new codes need no other change.kept
2026-10-01Staging is on the new cluster.removed: transient
2026-10-02Generated file again: never hand-edit `tally/src/report.gen.js`.merged: repeat

Full passes per task (Claude Sonnet 5.5, of 3)

TaskNo memory/init CLAUDE.mdCurated, 11 linesRaw notes, 60 linesDreamed notesHandbook, 210 linesStop hook onlyCurated + hook
Refunds33323333
Report decimals00333333
Customer phone33333333
Late fee00333303
Slugify (control)33333333

Method

  1. One small Node.js repository with real conventions: money in integer cents, coded errors, a clock helper, a log helper, append-only migrations, a generated report, a changelog rule and a late-fee rate that exists only in memory.
  2. Five tasks: refunds, a report formatting bug, an optional phone field, a late fee "at our standard rate", and a control task where memory does not matter.
  3. Eight conditions: no memory, the real /init output, a curated 11-line file, 56 raw dated notes (repeats, one-off events and notes that a later note contradicts), the same notes after one "dreaming" pass, the curated facts inside a 210-line handbook, a Stop hook that blocks finishing while a rule is broken, and curated plus hook.
  4. Claude Code 2.1.286 headless; Sonnet 3 repetitions per cell, Haiku 2; tools Bash, Read, Edit, Write, Glob, Grep; OS sandbox; the tester's own instruction files excluded and auto memory off (both checked with a probe).
  5. Grading after the session: hidden tests run in a sandbox, plus deterministic checks on the lines the agent added. Full pass = tests and every applicable check. Every attempt counts.
  6. The protocol was declared before the first counted session. One amendment, made before any Haiku session, added the Haiku lane. One erratum: the raw notes file holds 56 notes, not 58 as first written. The frozen handbook has 210 lines, not 211 as the protocol first stated. Measurements did not change.

Caveats

  • One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  • The hook checks the same code rules as the grader. It shows what rules written as code can do; it cannot carry a fact such as the late-fee rate.
  • n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  • Costs are the CLI's list-price estimates for subscription sessions, not invoices.
  • In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.

Sources

  • Agent memory study: 8 kinds of project memory on Claude Code

    Our recorded runs ·

    One small Node.js repository, 5 tasks, 8 memory conditions (none, /init, curated, raw notes, dreamed notes, long handbook, Stop hook, curated + hook). Claude Code 2.1.286 headless: Sonnet lane 3 repetitions, Haiku lane 2. Hidden tests and deterministic convention checks; protocol declared before the first session; every attempt kept. The exact memory files and three dreaming passes are published.

    Raw data: agent-memory/attempts-sonnet.json, agent-memory/attempts-haiku.json, agent-memory/dreams.json, agent-memory/memory-files.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Does memory help Claude Code? 8 kinds of agent memory, tested”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/agent-memory.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

Models and comparisons in this study

More studies

All benchmarks
  • Prompt Caching
  • Cache Reuse

Does a new Claude Code session reuse the prompt cache of an earlier one?

30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.

0of 2 (95% interval 0% to 66%) · Later sessions with at least 50% of turn-1 input cached, A: new folder each time · n = 2

4 chartsUpdated October 7, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.