Explainer · Context engineering
What is context engineering? What to put in the context, measured
Definition
Context engineering chooses what enters a model's context window on each call: instructions, memory, tools, retrieved files and history. It covers more than wording. In our tests, curated memory improved team-knowledge checks. Haiku used a stale command less often with consolidated notes than raw notes.
Agent team · · 5 min read · Every number is from the public studies
| Item | Median input tokens per session | n |
|---|---|---|
| No memory | 78,455 | 15 |
| /init CLAUDE.md | 82,262 | 15 |
| Curated, 11 lines | 65,045 | 15 |
| Raw notes, 60 lines | 76,963 | 15 |
| Dreamed notes | 68,218 | 15 |
| Handbook, 210 lines | 85,223 | 15 |
| Stop hook only | 125,674 | 15 |
| Curated + hook | 69,224 | 15 |
8 rows. Highest Stop hook only 125,674 (n 15). Lowest Curated, 11 lines 65,045 (n 15).
Notesn = 15 per row
Median input tokens per session, cache reads included (Claude Sonnet 5.5)
Input = uncached input + cache reads + cache writes over every turn, as the CLI reports it. Most of it is read from the prompt cache. The hook adds turns: each block sends the agent back to work.
Source: Agent memory study: 8 kinds of project memory on Claude Code
What goes into the context
You choose memory and retrieved files, such as CLAUDE.md. The CLI adds its own context.
For a prompt requesting a one-word answer, Claude Code with Haiku 4.5 sent a median 6,761 input tokens (range 6,756 to 6,765), tools and MCP servers off. Codex CLI (default model) sent 17,051 (17,049 to 20,408) in a read-only sandbox. Each ran 5 times. Claude counts include cache reads and writes; Codex input already includes cache reads. These are ranges, not confidence intervals. Different models mean these figures do not isolate CLI overhead.
Rate intervals below are 95% Wilson intervals unless stated otherwise.
Add what the code cannot show
The memory study ran 200 Claude Code sessions on tally, a small synthetic Node.js repo: 120 Sonnet 5.5, 80 Haiku 4.5, 8 memory conditions, hidden tests. Headless questions count as failures; a person could answer in live sessions.
Without memory, Sonnet followed 95% of code-visible rules (57 of 60; 86% to 98%) and 100% of folder-visible rules (21 of 21; 85% to 100%). Curated memory scored 100% on both: 60 of 60 (94% to 100%) and 21 of 21 (85% to 100%). Intervals overlap; curated results hit this task set's ceiling. Checks come from 15 sessions per condition, not independent tasks.
For team knowledge, Sonnet scored 40% without memory (6 of 15; 20% to 64%) versus 100% with an 11-line curated file (15 of 15; 80% to 100%). We wrote it knowing the tasks. Non-overlapping intervals put the file ahead.
Without the late-fee rate, Sonnet called it unknown or a guess in 9 of 9 sessions (70% to 100%, calculation from counts). Haiku did so in 0 of 6 (0% to 39%, calculation from counts). It wrote a rate and reported completion.
Remove what is stale or noisy
The README's test command fails on Node 25. /init copied it; raw notes kept it alongside a correction.
/initrepeated the error: 13 of 15 Sonnet sessions used the broken command (87%; 62% to 96%); without memory, 12 of 15 did (80%; 55% to 93%).- Raw notes did not reduce Haiku's stale-command use: with 56 notes, 10 of 10 sessions used it (72% to 100%), matching no memory (10 of 10; 72% to 100%). Sonnet used it in 0 of 15 (0% to 20%).
- Consolidated notes reduced use: Haiku used it in 1 of 10 (2% to 40%). The intervals do not overlap.
Each consolidation pass was one Claude Sonnet call, no tools. All 3 kept 10 of 10 current facts, 0 of 5 stale notes and 1 of 9 one-off notes. These are file scores, not success rates. Notes marked replaced did not count as kept.
Size and price
Sonnet passed every check in 15 of 15 sessions with the 11-line file and 15 of 15 with a 210-line handbook (80% to 100% each). Both hit the ceiling; no clear accuracy gain. Median input was 65,045 versus 85,223 tokens per session; ranges 42,802 to 136,178 versus 55,290 to 150,160 (n = 15 each). Counts include cache reads and writes. Overlapping ranges show no clear token-use difference despite the higher median.
Caching makes stable context cheap, not free:
Cache shares are calculations from token counters. Claude Code turn 1 read about 19% from cache: Sonnet 5.5 and Opus 5.5 range 18.680% to 18.689% (n = 6). The study attributes this to the CLI prefix, without testing that cause separately. Turns 2 to 5 averaged 97% (88% to 99%, n = 24 across 6 sessions), 3 sessions per model.
Codex CLI read 55% on turn 1 (48% to 59%, n = 3), 99% on turns 2 to 5 (97% to 99%, n = 12). Its longer ledger prevents a matched comparison. These are ranges, not confidence intervals.
By calculation from tokens and list prices, caching cut Sonnet cost 50% ($0.1350 vs $0.2698): 15 turns, 3 sessions. Subscription calls mean these are not invoices. Every write used the one-hour rate, twice normal input price.
No clear speed effect: Sonnet's median was 1.64 s on turn 1 (1.58 to 1.79 s, n = 3), 1.61 s later (1.35 to 5.63 s, n = 12). Ranges overlap and are not intervals. Different questions prevent isolating a cache speed effect.
A short checklist
- Count every input token, CLI prefix included.
- Add facts the repo cannot show, briefly.
- Delete stale and one-off notes; consolidate once.
- Check commands before adding them to memory.
- Script rules you can check. The Stop hook passed 60 of 60 Sonnet code-rule checks (94% to 100%; 15 sessions). It used the grader's rules; it could not supply missing facts.
- Keep the context prefix stable for caching.
- Test your tasks; ours used one small synthetic repo.
We build Agent to carry work and corrections across AI tools. We have a stake in memory. The full write-up publishes protocol, data and memory files. Try Agent.
Frequently asked questions
How is context engineering different from prompt engineering?
Prompt engineering words one instruction; context engineering selects all accompanying tokens. We did not test wording or which matters more.
Does a bigger context make an agent better?
Sonnet's results above hit the ceiling; token ranges overlap. Haiku passed 7 of 10 with the short file, 3 of 10 with the handbook (40% to 89%; 11% to 60%). Overlapping intervals leave the gap unclear.
What should never go into an agent's context?
Our stale-note results above favour consolidation. We did not measure secrets; this is advice: keep keys and passwords out of context.
Does prompt caching make long context free?
No: reads cost less, writes cost more. Sonnet's 50% saving above is a calculation. Later Claude Code sessions showed no earlier session's ledger reuse (0 of 4; 0% to 49%, calculation from counts). Their first turns still reported 1,463 cache-read tokens. We did not test why. See prompt caching explained.
Watch the data
Does memory help Claude Code? 8 kinds of agent memory, tested
200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.
Transcript
- Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
- Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
- Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
- Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
- The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
- Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- Write down what the repo cannot show. Enforce what a script can check.