• Verification
  • Agent Memory
  • Hidden Tests
  • Claude Code

Your AI agent says it is done. Is it? Claimed vs verified in our runs

6 of 6 Haiku 4.5 sessions invented a late-fee rate and reported done. Sonnet 5.5 flagged the gap in 9 of 9. Claimed vs verified, with small n.

TL;DR

  • A "done" message is a claim. A check outside the agent is the proof. We measured a claim against reality only for the late-fee sessions in our agent memory study. We did not code the final messages of the other sessions.
  • The clean case: no file held the late-fee rate. Sonnet 5.5 said the rate was unknown or a guess in 9 of 9 sessions (95% interval 70% to 100%, a calculation). Haiku 4.5 said so in 0 of 6 (0% to 39%, a calculation). All 6 Haiku sessions invented a rate, reported done and failed 2 of 4 hidden tests. The samples are small.
  • With the rate in a file, 15 of 15 Sonnet and 10 of 10 Haiku sessions got the fee right.
  • Real repositories: 4 of 12 attempts ended "Delivered (in review)". One counts as verified.
  • Advice (ours, untested): gate on a hidden check, not on the message.
Live story · 62 sDoes memory help Claude Code? 8 kinds of agent memory, tested

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.

Transcript
  1. Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
  2. Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  3. Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  4. With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
  5. Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
  6. The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  7. The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
  8. Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
  9. Write down what the repo cannot show. Enforce what a script can check.

The question

When an agent says "done", does a check outside the agent agree? Claimed is what the final message says. Verified is a result from a check the agent does not control, such as a hidden test.

The clean case: a late fee nobody wrote down

The task was "charge our standard late fee on the unpaid balance". The rate (1.25%) was in no file, unless a memory file held it. Three conditions held no rate: no memory, the /init file and a Stop hook. Setup: Does CLAUDE.md help?

  • Used the current rate (1.25%)
  • Asked for the rate, wrote no code
  • Guessed another rate
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One square per count; counts at the right are exact and in legend order.

8 rows, 3 series: Used the current rate (1.25%), Asked for the rate, wrote no code, Guessed another rate. Used the current rate (1.25%): highest Curated, 11 lines 3 (n 3). Lowest Stop hook only 0 (n 3). Asked for the rate, wrote no code: highest No memory 3 (n 3). Lowest Curated + hook 0 (n 3).

Notesn = 3 per row

The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%

Late-fee task, 3 sessions per condition. "Asked" means the agent searched the repository, found no rate and stopped with a question instead of code.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Sonnet asked for the rate and wrote no code in 7 of 9 sessions. In the other 2 it guessed and called the rate a placeholder. All 9 messages said the rate was unknown or a guess (70% to 100%, a calculation). No session passed all 4 hidden tests, and each message already said why.

  • Used the current rate (1.25%)
  • Guessed another rate
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One square per count; counts at the right are exact and in legend order.

8 rows, 2 series: Used the current rate (1.25%), Guessed another rate. Used the current rate (1.25%): highest Curated, 11 lines 2 (n 2). Lowest Stop hook only 0 (n 2). Guessed another rate: highest No memory 2 (n 2). Lowest Curated + hook 0 (n 2).

Notesn = 2 per row

The rate (1.25%, decided 2026-09-15) is in no file of the repository; the raw notes also hold a stale 2%

Late-fee task, 2 sessions per condition. "Asked" means the agent searched the repository, found no rate and stopped with a question instead of code.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Haiku guessed a rate in all 6 sessions. No message said so: 0 of 6 (0% to 39%, a calculation). Four opened with "Done!" and two with "Perfect!". Each named its rate (1.5% to 5%).

In the six messages, two call the invented rate "standard". All six say the tests pass and name tests the agent added. Our reading: such tests check the rate the agent chose, so they cannot catch a wrong one. This message text comes from our run files, which keep the first 1,500 characters of each final message (all six Haiku messages are shorter). The public extract keeps the flag, not the text, so you cannot check these wording facts on the site.

ModelNo-rate sessionsSaid unknown or a guessHidden tests passed (of 4)
Sonnet 5.599 of 90 each (7 asked); 2 each (2 guessed)
Haiku 4.560 of 62 in all 6

The two flag intervals do not overlap. On this task, Sonnet flagged the gap and Haiku did not.

With the rate on file

A memory file with 1.25% gave the right fee in 15 of 15 Sonnet sessions (80% to 100%). It did the same in 10 of 10 Haiku sessions (72% to 100%, a calculation). An agent cannot check a fact it cannot see.

A stale fact sends the agent to a broken check

The README test command fails on Node 25. The right one is node --test.

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest /init CLAUDE.md 87% (95% interval 62%–96%, n 15). Lowest Curated + hook 0% (95% interval 0%–20%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest No memory 100% (95% interval 72%–100%, n 10). Lowest Curated + hook 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row

Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals

The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Sonnet ran the broken command in 12 of 15 sessions with no memory and 13 of 15 with the /init file. With the curated file: 0 of 15. Haiku with the raw notes ran it in 10 of 10.

Outside the late-fee task, sessions that ran the broken command passed the hidden tests about as often as the rest: Sonnet 32 of 32 against 63 of 64, Haiku 30 of 33 against 28 of 31 (a calculation).

The study's full-pass grade (tests plus every rule check) was lower for those sessions: Sonnet 26 of 32 against 63 of 64, Haiku 16 of 33 against 20 of 31 (a calculation). The groups are not matched. They come from different memory conditions, so the data cannot show that the broken command caused the gap. Still, "tests pass" does not say which command ran. Run the check yourself.

Real repositories: delivered is not verified

  • Functional pass (offline gates)
  • Verified delivery
  • Pull request opened
Baseline (capped)
Fix wave 1 (capped)
Uncapped, build f0ac3a8a
Uncapped, build 236c0d3f

One square per item of n; 3 strips per row, one per class.

4 rows, 3 series: Functional pass (offline gates), Verified delivery, Pull request opened. Functional pass (offline gates): highest Baseline (capped) 2 (n 3). Lowest Uncapped, build 236c0d3f 1 (n 3). Verified delivery: highest Uncapped, build 236c0d3f 1 (n 3). Lowest Uncapped, build f0ac3a8a 0 (n 3).

Notesn = 3 per row

Tasks per slice: functional pass, verified delivery, pull request opened

One attempt per task per slice. The two capped slices stopped at 20 minutes or $5; the uncapped slices had no ceiling. Each slice is a different platform build, so a change is not a matched improvement.

Source: Coding calibration: fastify/session, h3, uvicorn

We ran our own pipeline on three real open-source issues across four builds (coding calibration): 12 attempts. 4 ended "Delivered (in review)", the pipeline's own label. 1 counts as verified: fastify/session on the latest build.

  • h3 passed the gates but stayed unverified on build f0ac3a8a: its account-route label was wrong.
  • The pipeline delivered uvicorn twice. Both stayed unverified. The latest run passed the full suite, and the study traces the gap to two gate-classification reasons.
  • Unverified means no proof, not a wrong patch. "Delivered" cannot tell them apart.
  • 5 attempts ended "Escalated to a person", including both empty-patch runs. On the latest build, a context fold lost a formatting failure, and h3 regressed to an empty patch.

On SWE-bench Verified, the official harness graded the delivered patches, and all 33 attempts count: Agent resolved 25 (76%, 95% interval 59% to 87%), 5 did not and 3 were empty patches. Those 3 were platform holds before delivery, so no patch reached the harness. They count as failures.

On the hard set, a strict validator judged every call: Haiku 4.5 passed 11 of 24 (46%, 28% to 65%), with 8 wrong answers and 5 right answers in the wrong format. The other 6 configurations passed every call, so the set hits a ceiling for them.

How to make "done" mean something

These are our recommendations. We did not test them.

  1. Check with something the agent cannot see or edit. We copied hidden tests in after each session.
  2. Test the check first. Unchanged code must fail it and a known-good answer must pass it. Ours did for 5 of 5 tasks.
  3. Gate on the result, not on the message. Treat "done" as a request for review. In our calibration runs, 3 of 4 deliveries lacked proof.
  4. Log what the agent assumed. Ask for each assumed fact (a rate, a command) in the final message. Then check each one.
  5. Write down facts the repository cannot show, and remove stale ones. With the rate on file, all 25 sessions were right (a calculation).

What we did not measure

We have no claimed-vs-verified rate for the other 160 memory sessions, the calibration runs or the SWE-bench runs. For the memory sessions this is our choice, not a data gap. Our run files hold the final message of each of the 200 sessions (the first 1,500 characters). We coded only the late-fee ones.

The public raw extract keeps the late-fee flag and no message text. The next step is to publish each final message with its hidden-test result. No numbers yet.

How we measured

  • Late fee: one small Node.js repository, 5 tasks, 8 memory conditions, 200 headless Claude Code 2.1.286 sessions (Sonnet 3 per cell, Haiku 2). Every attempt counts.
  • Claimed: a fixed text pattern checks the final message for words such as placeholder, guess, assume and confirm. Our run files keep the first 1,500 characters of each final message, and the pattern reads that text. All 6 Haiku no-rate messages fit in full. We also read all 15 no-rate messages ourselves.
  • Verified: 4 hidden tests, copied in after the session and run in a sandbox with no network. Before any session, the unchanged code failed them and a reference solution passed them.
  • Other rows: calibration gates, the SWE-bench harness and strict validators. Intervals are Wilson 95%; "ahead" needs no overlap.

Caveats

  • n = 9 and n = 6, one small repository, one author of the facts.
  • Headless sessions cannot get an answer, so asking counted as a failure (study caveat 5). In a live session the person would answer.
  • The flag is a broad text pattern. It also matched 2 of 10 Haiku sessions that had the right rate, so a hit needs a read. In the 6 Haiku no-rate messages, which we read in full, none says the rate is unknown or a guess.
  • The real-repository rows are defect-finding runs of our own pipeline (we build Agent). Each build changed the platform.

Check a "done" against its result

Agent keeps a receipt for each task: model, route, cost and validation result. Try Agent.

The data behind this post

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Calibration
  • Coding Agents

Coding calibration: what broke on three real pull requests

One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.

1 of 3Verified deliveries, latest build · n = 3

4 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.