Explainer · CLAUDE.md

What is CLAUDE.md? Agent memory files, measured

Definition

CLAUDE.md is a Markdown file in your project that Claude Code reads into its context at the start of every session. It holds facts the code cannot show, such as a team decision or the current test command, and it helped most there in our tests. Codex CLI reads AGENTS.md for the same job, but we measured Claude Code only.

Agent team · · 5 min read · Every number is from the public studies

Rules the code already shows

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Rules the folders hint at

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Team knowledge only

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 3 series: Rules the code already shows, Rules the folders hint at, Team knowledge only. Rules the code already shows: highest Curated, 11 lines 100% (95% interval 94%–100%, n 60). Lowest /init CLAUDE.md 95% (95% interval 86%–98%, n 60). All intervals overlap. Rules the folders hint at: all at 100%.

NotesWhiskers: 95% Wilson intervaln 15–60 per row6 of 8 (Rules the code already shows) at 100%: this task set cannot separate them.

Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

Source: Agent memory study: 8 kinds of project memory on Claude Code

What CLAUDE.md is for

We tested it in one small synthetic repository with five tasks and eight kinds of memory. We graded 200 sessions: 120 with Claude Sonnet 5.5 and 80 with Claude Haiku 4.5 (blog post).

Rules the code already shows

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Rules the folders hint at

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Team knowledge only

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 3 series: Rules the code already shows, Rules the folders hint at, Team knowledge only. Rules the code already shows: highest Curated, 11 lines 100% (95% interval 94%–100%, n 60). Lowest /init CLAUDE.md 95% (95% interval 86%–98%, n 60). All intervals overlap. Rules the folders hint at: all at 100%.

NotesWhiskers: 95% Wilson intervaln 15–60 per row6 of 8 (Rules the code already shows) at 100%: this task set cannot separate them.

Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

Source: Agent memory study: 8 kinds of project memory on Claude Code

With no memory, Sonnet passed 57 of 60 checks on rules that the code shows (95%; 95% Wilson interval 86.3% to 98.3%). It passed 21 of 21 checks on rules that the folders hint at (84.5% to 100%). Both sit near a ceiling, so a file has little room to help.

Two facts only the team knew: a changelog rule and a late-fee rate. With no memory, Sonnet passed 6 of 15 of those checks (19.8% to 64.3%). With an 11-line file it passed 15 of 15 (79.6% to 100%). The intervals do not overlap, so the file is ahead on team facts.

Without the rate in memory, Sonnet's final message called it unknown or a guess in 9 of 9 sessions (70.1% to 100%; calculation). Haiku 4.5's final message did so in 0 of 6 sessions (0% to 39.0%; calculation). All 6 guessed another rate. Sonnet asked and wrote no code in 7 of the 9.

What /init writes

/init is the Claude Code command that writes a first CLAUDE.md from your repository. In our test it copied a mistake.

Our README gave a test command that fails on Node 25, and the /init file repeated it. In 13 of 15 Sonnet sessions with that file, the agent ran the broken command (86.7%; 62.1% to 96.3%). With no memory, 12 of 15 did (80.0%; 54.8% to 93.0%). The intervals overlap, so that is a tie. With the 11-line file, which named the right command, 0 of 15 sessions ran the broken one (0% to 20.4%). That is lower than /init: the intervals do not overlap.

In full passes, the /init file did no better than no file: 9 of 15 each (35.8% to 80.2%). That is a tie.

Treat /init output as a draft: fix stale docs at the source and add what only your team knows. We did not test an edited /init file.

A short file or a long handbook

For Sonnet, length made no measurable difference. The 11-line file and a 211-line handbook with the same facts both passed 15 of 15 sessions in full (79.6% to 100%). That is a tie. The set has a ceiling: even 15 of 15 reaches down to 79.6%.

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest /init CLAUDE.md 60% (95% interval 36%–80%, n 15). All intervals overlap. Claude Haiku 4.5: highest Curated + hook 90% (95% interval 60%–98%, n 10). Lowest /init CLAUDE.md 20% (95% interval 5.7%–51%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row4 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Hidden tests pass and every convention check passes · 95% Wilson intervals

Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.

Source: Agent memory study: 8 kinds of project memory on Claude Code

No memory passed 9 of 15 in full (35.8% to 80.2%). That overlaps the file's by a hair, so that gap is unclear.

The smaller model, Haiku 4.5, ran 10 sessions per condition. It passed 2 of 10 in full with no memory (5.7% to 51.0%) and 7 of 10 with the 11-line file (39.7% to 89.2%). Those intervals overlap, so that gap is unclear. On team-fact checks alone, the file is ahead: 0 of 10 against 8 of 10 (0% to 27.8% against 49.0% to 94.3%). The handbook passed 3 of 10 on both measures (10.8% to 60.3%). That overlaps the other two, so the effect of a long file is unclear.

Stale lines did not help the smaller model. The raw notes held an old test command and its later correction. Haiku ran the old command in 10 of 10 sessions (72.3% to 100%), the same as with no memory. After one dreaming pass, in which a model merges and prunes notes, it ran that command in 1 of 10 (1.8% to 40.4%).

What it costs in tokens

The file loads into every session, so it adds to the context.

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows. Highest Stop hook only 125,674 (n 15). Lowest Curated, 11 lines 65,045 (n 15).

Notesn = 15 per row

Median input tokens per session, cache reads included (Claude Sonnet 5.5)

Input = uncached input + cache reads + cache writes over every turn, as the CLI reports it. Most of it is read from the prompt cache. The hook adds turns: each block sends the agent back to work.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Median input tokens per Sonnet session (n = 15 per condition): no memory 78,455; /init 82,262; 11-line file 65,045; handbook 85,223. The handbook's median was 31% above the short file's (calculation: 85,223 ÷ 65,045 = 1.31). The chart shows medians with no run range, so the order is this run's result, not a tested ranking. No memory used more than the short file, so the medians do not track file length.

At list price (a calculation, not a bill), one fully correct Sonnet result cost $0.082 with the 11-line file and $0.139 with no memory. With the file and with none, the 15 sessions cost about the same in total ($1.23 and $1.25; calculation). The gap comes from full passes, 15 and 9. Their intervals overlap, so it is unclear.

A CLAUDE.md checklist

  1. Write facts the code cannot show: dated team decisions, current commands and reasons.
  2. Leave out what the code shows. Sonnet passed 57 of 60 such checks with no file.
  3. Keep it short. Our token medians have no run range, so we claim no saving.
  4. Remove stale lines. Replace an old command. Do not add a correction below it.
  5. Test it. Run a few real tasks with and without the file. Keep the lines that change the result.

Limits. One small repository and one author of its rules. The 11-line file is an upper bound: we wrote it knowing the tasks. Sample sizes were 15 (Sonnet) and 10 (Haiku), so most full-pass intervals overlap. In headless mode a question gets no answer, so asking counts as a failure. The checklist is our advice, not a tested result.

Frequently asked questions

Does CLAUDE.md actually help?

Yes, for facts the code cannot show. Sonnet 5.5 passed 6 of 15 team-fact checks with no memory (19.8% to 64.3%) and 15 of 15 with an 11-line file (79.6% to 100%). The test used one small repository.

Should I run /init?

Run it as a first draft. Then edit it. It copied our README's stale test command. Sonnet ran that command in 13 of 15 sessions with the /init file (62.1% to 96.3%). With no file it ran it in 12 of 15 (54.8% to 93.0%). That is a tie.

How long should a CLAUDE.md be?

Short enough that each line adds something the code cannot. Sonnet passed 15 of 15 sessions in full with 11 lines and with 211 (79.6% to 100% each). We tested the same facts at only those two lengths.

Is AGENTS.md the same thing?

It plays the same role for Codex CLI: a Markdown file of project instructions. We measured Claude Code only, so we cannot say these numbers hold for AGENTS.md.

Is CLAUDE.md the same as Claude Code auto memory?

They are two features. We turned auto memory off in every session, so we have no data on it.

Watch the data

Live story · 62 sDoes memory help Claude Code? 8 kinds of agent memory, tested

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.

Transcript
  1. Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
  2. Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  3. Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  4. With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
  5. Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
  6. The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  7. The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
  8. Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  9. Write down what the repo cannot show. Enforce what a script can check.

The data behind this explainer

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.