Your AI agent says it is done. Is it? Claimed vs verified in our runs
6 of 6 Haiku 4.5 sessions invented a late-fee rate and reported done. Sonnet 5.5 flagged the gap in 9 of 9. Claimed vs verified, with small n.
TL;DR
- A "done" message is a claim. A check outside the agent is the proof. We measured a claim against reality only for the late-fee sessions in our agent memory study. We did not code the final messages of the other sessions.
- The clean case: no file held the late-fee rate. Sonnet 5.5 said the rate was unknown or a guess in 9 of 9 sessions (95% interval 70% to 100%, a calculation). Haiku 4.5 said so in 0 of 6 (0% to 39%, a calculation). All 6 Haiku sessions invented a rate, reported done and failed 2 of 4 hidden tests. The samples are small.
- With the rate in a file, 15 of 15 Sonnet and 10 of 10 Haiku sessions got the fee right.
- Real repositories: 4 of 12 attempts ended "Delivered (in review)". One counts as verified.
- Advice (ours, untested): gate on a hidden check, not on the message.
Does memory help Claude Code? 8 kinds of agent memory, tested
200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.
Transcript
- Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.
- Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
- Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: "Charge our standard late fee" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
- Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.
- The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.
- Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
- Write down what the repo cannot show. Enforce what a script can check.
The question
When an agent says "done", does a check outside the agent agree? Claimed is what the final message says. Verified is a result from a check the agent does not control, such as a hidden test.
The clean case: a late fee nobody wrote down
The task was "charge our standard late fee on the unpaid balance". The rate (1.25%) was in no file, unless a memory file held it. Three conditions held no rate: no memory, the /init file and a Stop hook. Setup: Does CLAUDE.md help?
Sonnet asked for the rate and wrote no code in 7 of 9 sessions. In the other 2 it guessed and called the rate a placeholder. All 9 messages said the rate was unknown or a guess (70% to 100%, a calculation). No session passed all 4 hidden tests, and each message already said why.
Haiku guessed a rate in all 6 sessions. No message said so: 0 of 6 (0% to 39%, a calculation). Four opened with "Done!" and two with "Perfect!". Each named its rate (1.5% to 5%).
In the six messages, two call the invented rate "standard". All six say the tests pass and name tests the agent added. Our reading: such tests check the rate the agent chose, so they cannot catch a wrong one. This message text comes from our run files, which keep the first 1,500 characters of each final message (all six Haiku messages are shorter). The public extract keeps the flag, not the text, so you cannot check these wording facts on the site.
| Model | No-rate sessions | Said unknown or a guess | Hidden tests passed (of 4) |
|---|---|---|---|
| Sonnet 5.5 | 9 | 9 of 9 | 0 each (7 asked); 2 each (2 guessed) |
| Haiku 4.5 | 6 | 0 of 6 | 2 in all 6 |
The two flag intervals do not overlap. On this task, Sonnet flagged the gap and Haiku did not.
With the rate on file
A memory file with 1.25% gave the right fee in 15 of 15 Sonnet sessions (80% to 100%). It did the same in 10 of 10 Haiku sessions (72% to 100%, a calculation). An agent cannot check a fact it cannot see.
A stale fact sends the agent to a broken check
The README test command fails on Node 25. The right one is node --test.
Sonnet ran the broken command in 12 of 15 sessions with no memory and 13 of 15 with the /init file. With the curated file: 0 of 15. Haiku with the raw notes ran it in 10 of 10.
Outside the late-fee task, sessions that ran the broken command passed the hidden tests about as often as the rest: Sonnet 32 of 32 against 63 of 64, Haiku 30 of 33 against 28 of 31 (a calculation).
The study's full-pass grade (tests plus every rule check) was lower for those sessions: Sonnet 26 of 32 against 63 of 64, Haiku 16 of 33 against 20 of 31 (a calculation). The groups are not matched. They come from different memory conditions, so the data cannot show that the broken command caused the gap. Still, "tests pass" does not say which command ran. Run the check yourself.
Real repositories: delivered is not verified
We ran our own pipeline on three real open-source issues across four builds (coding calibration): 12 attempts. 4 ended "Delivered (in review)", the pipeline's own label. 1 counts as verified: fastify/session on the latest build.
- h3 passed the gates but stayed unverified on build f0ac3a8a: its account-route label was wrong.
- The pipeline delivered uvicorn twice. Both stayed unverified. The latest run passed the full suite, and the study traces the gap to two gate-classification reasons.
- Unverified means no proof, not a wrong patch. "Delivered" cannot tell them apart.
- 5 attempts ended "Escalated to a person", including both empty-patch runs. On the latest build, a context fold lost a formatting failure, and h3 regressed to an empty patch.
On SWE-bench Verified, the official harness graded the delivered patches, and all 33 attempts count: Agent resolved 25 (76%, 95% interval 59% to 87%), 5 did not and 3 were empty patches. Those 3 were platform holds before delivery, so no patch reached the harness. They count as failures.
On the hard set, a strict validator judged every call: Haiku 4.5 passed 11 of 24 (46%, 28% to 65%), with 8 wrong answers and 5 right answers in the wrong format. The other 6 configurations passed every call, so the set hits a ceiling for them.
How to make "done" mean something
These are our recommendations. We did not test them.
- Check with something the agent cannot see or edit. We copied hidden tests in after each session.
- Test the check first. Unchanged code must fail it and a known-good answer must pass it. Ours did for 5 of 5 tasks.
- Gate on the result, not on the message. Treat "done" as a request for review. In our calibration runs, 3 of 4 deliveries lacked proof.
- Log what the agent assumed. Ask for each assumed fact (a rate, a command) in the final message. Then check each one.
- Write down facts the repository cannot show, and remove stale ones. With the rate on file, all 25 sessions were right (a calculation).
What we did not measure
We have no claimed-vs-verified rate for the other 160 memory sessions, the calibration runs or the SWE-bench runs. For the memory sessions this is our choice, not a data gap. Our run files hold the final message of each of the 200 sessions (the first 1,500 characters). We coded only the late-fee ones.
The public raw extract keeps the late-fee flag and no message text. The next step is to publish each final message with its hidden-test result. No numbers yet.
How we measured
- Late fee: one small Node.js repository, 5 tasks, 8 memory conditions, 200 headless Claude Code 2.1.286 sessions (Sonnet 3 per cell, Haiku 2). Every attempt counts.
- Claimed: a fixed text pattern checks the final message for words such as placeholder, guess, assume and confirm. Our run files keep the first 1,500 characters of each final message, and the pattern reads that text. All 6 Haiku no-rate messages fit in full. We also read all 15 no-rate messages ourselves.
- Verified: 4 hidden tests, copied in after the session and run in a sandbox with no network. Before any session, the unchanged code failed them and a reference solution passed them.
- Other rows: calibration gates, the SWE-bench harness and strict validators. Intervals are Wilson 95%; "ahead" needs no overlap.
Caveats
- n = 9 and n = 6, one small repository, one author of the facts.
- Headless sessions cannot get an answer, so asking counted as a failure (study caveat 5). In a live session the person would answer.
- The flag is a broad text pattern. It also matched 2 of 10 Haiku sessions that had the right rate, so a hit needs a read. In the 6 Haiku no-rate messages, which we read in full, none says the rate is unknown or a guess.
- The real-repository rows are defect-finding runs of our own pipeline (we build Agent). Each build changed the platform.
What to read next
- Does CLAUDE.md help? We tested 8 kinds of agent memory
- Why we count every failed attempt
- How to read AI benchmarks honestly and Claude Haiku 4.5 vs Sonnet 5.5
Check a "done" against its result
Agent keeps a receipt for each task: model, route, cost and validation result. Try Agent.