Explainer · Context engineering

What is context engineering? What to put in the context, measured

Definition

Context engineering chooses what enters a model's context window on each call: instructions, memory, tools, retrieved files and history. It covers more than wording. In our tests, curated memory improved team-knowledge checks. Haiku used a stale command less often with consolidated notes than raw notes.

Agent team · · 5 min read · Every number is from the public studies

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows. Highest Stop hook only 125,674 (n 15). Lowest Curated, 11 lines 65,045 (n 15).

Notesn = 15 per row

Median input tokens per session, cache reads included (Claude Sonnet 5.5)

Input = uncached input + cache reads + cache writes over every turn, as the CLI reports it. Most of it is read from the prompt cache. The hook adds turns: each block sends the agent back to work.

Source: Agent memory study: 8 kinds of project memory on Claude Code

What goes into the context

You choose memory and retrieved files, such as CLAUDE.md. The CLI adds its own context.

For a prompt requesting a one-word answer, Claude Code with Haiku 4.5 sent a median 6,761 input tokens (range 6,756 to 6,765), tools and MCP servers off. Codex CLI (default model) sent 17,051 (17,049 to 20,408) in a read-only sandbox. Each ran 5 times. Claude counts include cache reads and writes; Codex input already includes cache reads. These are ranges, not confidence intervals. Different models mean these figures do not isolate CLI overhead.

Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows. Highest Codex CLI (default model) 17,051 (n 5). Lowest Claude Code · Claude Haiku 4.5 6,761 (n 5).

Notesn = 5 per row

Per call, mostly the CLI’s own system prompt and tool definitions

Claude Code sums its disjoint input, cache-read and cache-write fields. Codex CLI reports 17,051 input tokens, 13,184 of them read from the cache. The prompt itself is a few tokens.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

Rate intervals below are 95% Wilson intervals unless stated otherwise.

Add what the code cannot show

The memory study ran 200 Claude Code sessions on tally, a small synthetic Node.js repo: 120 Sonnet 5.5, 80 Haiku 4.5, 8 memory conditions, hidden tests. Headless questions count as failures; a person could answer in live sessions.

Rules the code already shows

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Rules the folders hint at

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Team knowledge only

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 3 series: Rules the code already shows, Rules the folders hint at, Team knowledge only. Rules the code already shows: highest Curated, 11 lines 100% (95% interval 94%–100%, n 60). Lowest /init CLAUDE.md 95% (95% interval 86%–98%, n 60). All intervals overlap. Rules the folders hint at: all at 100%.

NotesWhiskers: 95% Wilson intervaln 15–60 per row6 of 8 (Rules the code already shows) at 100%: this task set cannot separate them.

Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Without memory, Sonnet followed 95% of code-visible rules (57 of 60; 86% to 98%) and 100% of folder-visible rules (21 of 21; 85% to 100%). Curated memory scored 100% on both: 60 of 60 (94% to 100%) and 21 of 21 (85% to 100%). Intervals overlap; curated results hit this task set's ceiling. Checks come from 15 sessions per condition, not independent tasks.

For team knowledge, Sonnet scored 40% without memory (6 of 15; 20% to 64%) versus 100% with an 11-line curated file (15 of 15; 80% to 100%). We wrote it knowing the tasks. Non-overlapping intervals put the file ahead.

Without the late-fee rate, Sonnet called it unknown or a guess in 9 of 9 sessions (70% to 100%, calculation from counts). Haiku did so in 0 of 6 (0% to 39%, calculation from counts). It wrote a rate and reported completion.

Remove what is stale or noisy

The README's test command fails on Node 25. /init copied it; raw notes kept it alongside a correction.

  • /init repeated the error: 13 of 15 Sonnet sessions used the broken command (87%; 62% to 96%); without memory, 12 of 15 did (80%; 55% to 93%).
  • Raw notes did not reduce Haiku's stale-command use: with 56 notes, 10 of 10 sessions used it (72% to 100%), matching no memory (10 of 10; 72% to 100%). Sonnet used it in 0 of 15 (0% to 20%).
  • Consolidated notes reduced use: Haiku used it in 1 of 10 (2% to 40%). The intervals do not overlap.

Each consolidation pass was one Claude Sonnet call, no tools. All 3 kept 10 of 10 current facts, 0 of 5 stale notes and 1 of 9 one-off notes. These are file scores, not success rates. Notes marked replaced did not count as kept.

Size and price

Sonnet passed every check in 15 of 15 sessions with the 11-line file and 15 of 15 with a 210-line handbook (80% to 100% each). Both hit the ceiling; no clear accuracy gain. Median input was 65,045 versus 85,223 tokens per session; ranges 42,802 to 136,178 versus 55,290 to 150,160 (n = 15 each). Counts include cache reads and writes. Overlapping ranges show no clear token-use difference despite the higher median.

Caching makes stable context cheap, not free:

  • Claude Sonnet 5.5 · Claude Code
  • Claude Opus 5.5 · Claude Code
  • GPT-6.1 Sol (medium) · Codex CLI*

* The note below the chart says what this route does not report.

5 turn in the sessions, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (medium) · Codex CLI. Claude Sonnet 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3). Claude Opus 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3).

Notesn = 3 per row

Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each

Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Cache shares are calculations from token counters. Claude Code turn 1 read about 19% from cache: Sonnet 5.5 and Opus 5.5 range 18.680% to 18.689% (n = 6). The study attributes this to the CLI prefix, without testing that cause separately. Turns 2 to 5 averaged 97% (88% to 99%, n = 24 across 6 sessions), 3 sessions per model.

Codex CLI read 55% on turn 1 (48% to 59%, n = 3), 99% on turns 2 to 5 (97% to 99%, n = 12). Its longer ledger prevents a matched comparison. These are ranges, not confidence intervals.

By calculation from tokens and list prices, caching cut Sonnet cost 50% ($0.1350 vs $0.2698): 15 turns, 3 sessions. Subscription calls mean these are not invoices. Every write used the one-hour rate, twice normal input price.

No clear speed effect: Sonnet's median was 1.64 s on turn 1 (1.58 to 1.79 s, n = 3), 1.61 s later (1.35 to 5.63 s, n = 12). Ranges overlap and are not intervals. Different questions prevent isolating a cache speed effect.

A short checklist

  1. Count every input token, CLI prefix included.
  2. Add facts the repo cannot show, briefly.
  3. Delete stale and one-off notes; consolidate once.
  4. Check commands before adding them to memory.
  5. Script rules you can check. The Stop hook passed 60 of 60 Sonnet code-rule checks (94% to 100%; 15 sessions). It used the grader's rules; it could not supply missing facts.
  6. Keep the context prefix stable for caching.
  7. Test your tasks; ours used one small synthetic repo.

We build Agent to carry work and corrections across AI tools. We have a stake in memory. The full write-up publishes protocol, data and memory files. Try Agent.

Frequently asked questions

How is context engineering different from prompt engineering?

Prompt engineering words one instruction; context engineering selects all accompanying tokens. We did not test wording or which matters more.

Does a bigger context make an agent better?

Sonnet's results above hit the ceiling; token ranges overlap. Haiku passed 7 of 10 with the short file, 3 of 10 with the handbook (40% to 89%; 11% to 60%). Overlapping intervals leave the gap unclear.

What should never go into an agent's context?

Our stale-note results above favour consolidation. We did not measure secrets; this is advice: keep keys and passwords out of context.

Does prompt caching make long context free?

No: reads cost less, writes cost more. Sonnet's 50% saving above is a calculation. Later Claude Code sessions showed no earlier session's ledger reuse (0 of 4; 0% to 49%, calculation from counts). Their first turns still reported 1,463 cache-read tokens. We did not test why. See prompt caching explained.

Watch the data

Live story · 62 sDoes memory help Claude Code? 8 kinds of agent memory, tested

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.

Transcript
  1. Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
  2. Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  3. Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  4. With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
  5. Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
  6. The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  7. The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
  8. Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  9. Write down what the repo cannot show. Enforce what a script can check.

The data behind this explainer

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.