Explainer · Memory consolidation

Memory consolidation (dreaming) for AI agents, measured

Definition

Memory consolidation, also called dreaming, is a pass in which a model rewrites saved agent notes into one shorter file. It aims to keep current facts and drop stale notes and one-off events.

Agent team · · 5 min read · Every number is from the public studies

Current facts kept (of 10)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

Stale notes left (of 5)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

Transient notes left (of 9)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

One panel per series, all on the same axis.

6 rows, 3 series: Current facts kept (of 10), Stale notes left (of 5), Transient notes left (of 9). Current facts kept (of 10): highest Raw notes 10. Lowest /init 5. Stale notes left (of 5): highest Raw notes 5. Lowest Curated 0.

Notes

Fixed-pattern checks on each memory file: 10 current facts, 5 stale notes, 9 transient notes

Each dream is one Claude Sonnet call with no tools over the 56 raw notes. A stale note that the new file marks as replaced does not count as left. The one transient note every dream kept is the current dev-server port.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Each of our 3 Claude Sonnet 5.5 passes kept 10 of 10 checked current facts and left 0 of 5 stale notes. Haiku 4.5 ran a stale command in 10 of 10 raw-note sessions and 1 of 10 dreamed-note sessions. The 95% Wilson intervals were 72% to 100% and 2% to 40% (n = 10 each).

When raw notes hold conflicting commands

A later decision can make an old note wrong. We counted stale-command use, not why a model chose it.

Our small synthetic repository's README test command fails on Node 25. The raw file holds 56 dated notes, including an old instruction to run npm test and a later correction. Running the stale command is a bad outcome, so lower is better.

Sonnet 5.5 ran it in 0 of 15 raw-note sessions (95% Wilson interval 0% to 20%). Haiku with no memory also ran it in 10 of 10 sessions (72% to 100%).

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest /init CLAUDE.md 87% (95% interval 62%–96%, n 15). Lowest Curated + hook 0% (95% interval 0%–20%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest No memory 100% (95% interval 72%–100%, n 10). Lowest Curated + hook 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row

Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals

The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.

Source: Agent memory study: 8 kinds of project memory on Claude Code

What the passes kept and dropped

Each pass was one Claude Sonnet 5.5 call with no tools over the same 56 notes. The generic prompt asked for merged duplicates, fewer one-off details and decision dates. For conflicting notes, it asked for the newest unless an older note remained true.

Sessions used the first pass. Its memory file was in place before the first counted session. We measured a plain model call, not a product's built-in dreaming feature.

Current facts kept (of 10)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

Stale notes left (of 5)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

Transient notes left (of 9)

Raw notes
Dream 1
Dream 2
Dream 3
Curated
/init

One panel per series, all on the same axis.

6 rows, 3 series: Current facts kept (of 10), Stale notes left (of 5), Transient notes left (of 9). Current facts kept (of 10): highest Raw notes 10. Lowest /init 5. Stale notes left (of 5): highest Raw notes 5. Lowest Curated 0.

Notes

Fixed-pattern checks on each memory file: 10 current facts, 5 stale notes, 9 transient notes

Each dream is one Claude Sonnet call with no tools over the 56 raw notes. A stale note that the new file marks as replaced does not count as left. The one transient note every dream kept is the current dev-server port.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Fixed text patterns checked 10 current facts, 5 stale notes and 9 transient (one-off) notes. A stale note marked as replaced does not count as left.

  • Raw notes: 10 current facts, 5 stale notes and 9 transient notes.
  • Each of the 3 passes: 10, 0 and 1.
  • Hand-written curated file: 10, 0 and 0.
  • The /init file: 5, 1 and 0.

Each pass dropped 8 of 9 transient notes (calculation: 9 minus 1), keeping the current dev-server port. The first shortened 4,332 characters to 2,403: 45% fewer (calculation: 1 minus 2,403 / 4,332).

What changed in later sessions, and what it cost

Haiku's stale-command intervals do not overlap, so the lower rate is clear in this sample. We did not test why it followed the old note. Sonnet had no room to improve: 0 of 15 with either file (95% Wilson interval 0% to 20% each).

A full pass means hidden tests and every rule check pass. All session-rate intervals here are 95% Wilson intervals.

  • Sonnet 5.5: raw notes 14 of 15 (93%, interval 70% to 99%); dreamed notes 15 of 15 (100%, 80% to 100%). The intervals overlap. Dreamed notes hit this task set's ceiling.
  • Haiku 4.5: raw notes 6 of 10 (60%, 31% to 83%); dreamed notes 7 of 10 (70%, 40% to 89%). The intervals overlap, so no full-pass change is clear.

Sonnet median input was 76,963 tokens with raw notes and 68,218 with dreamed notes (n = 15 each). That is 8,745 fewer, about 11% (calculation: (76,963 minus 68,218) / 76,963). Sonnet attempts ranged from 49,257 to 134,351 with raw notes and 47,137 to 146,205 with dreamed notes.

Haiku medians were 285,263 and 267,405 tokens (n = 10 each). Haiku attempts ranged from 123,674 to 467,653 with raw notes and 137,193 to 486,214 with dreamed notes.

These ranges are not intervals. They overlap for each model, so no reduction is clear. Input sums uncached tokens, cache reads and cache writes across all turns. Most came from the prompt cache; see also the CLI context tax.

The pass receipts show a median 6.8 s, range 6.7 to 6.8 s (n = 3). Output ranged from 975 to 1,052 tokens. CLI list-price estimates ranged from $0.0107 to $0.0201 (calculation, not an invoice). These are ranges, not intervals. Calls used a subscription; we did not price input-token savings.

How to run a pass safely

This is advice, not tested steps.

  1. Keep the original notes in version control.
  2. Ask for decision dates and the newest of conflicting notes.
  3. Read the diff for lost facts and stale notes in new words.
  4. Test real sessions for a known error.

Limits of this test

  • One small synthetic repository, one pass model, one generic prompt and one set of 56 notes.
  • Each condition has 5 tasks: 3 repetitions per task for Sonnet and 2 for Haiku.
  • n = 15 Sonnet sessions and 10 Haiku sessions per condition. Repeated tasks do not represent independent repositories. Most full-pass intervals overlap (how to read them).
  • Experimenters wrote the facts. Pattern counts are not a general fact-retention rate.
  • Sessions ran headless. Asking a question counts as failure because no person can answer.
  • We did not test larger files, months of notes, other pass models or schedules.

See the memory study, Does CLAUDE.md help?, Claude Haiku 4.5 and Haiku vs Sonnet.

Frequently asked questions

What is dreaming in AI agent memory?

Dreaming is a consolidation pass that rewrites saved notes as a shorter, current file. We tested a plain call, not a built-in product feature.

Does consolidating notes improve results?

Haiku's stale-command use fell from 10 of 10 to 1 of 10 sessions. The 95% Wilson intervals do not overlap: 72% to 100% and 2% to 40% (n = 10 each).

Full-pass intervals overlap for both models; the results above show no clear full-pass change. Sonnet with dreamed notes hit this task set's ceiling.

Can consolidation delete a fact I need?

It can. All 3 passes kept the 10 checked current facts, but pattern checks cannot rule out other losses. Read the diff before accepting a rewrite.

How often should I consolidate memory?

We did not measure a schedule. Sessions used the first pass on one set of notes.

Watch the data

Live story · 62 sDoes memory help Claude Code? 8 kinds of agent memory, tested

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.

Transcript
  1. Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
  2. Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  3. Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  4. With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
  5. Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
  6. The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  7. The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
  8. Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  9. Write down what the repo cannot show. Enforce what a script can check.

The data behind this explainer

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.