- Home
- Best practices
Evidence-led recommendations · updated
Run coding agents.
Check the evidence.
AI coding agent best practices, with the measurements and limits beside each rule.
12 recommendations from 7 recorded studies. These are our recommendations. The studies did not test every practice as an intervention.
Recommendation
Validate every answer.
Check the delivered answer against a known reference or a task-specific test.
What we measured
- Measured0%
Same prompt, 10 times: strict pass rate
Claude Haiku 4.5 · Claude Code · Exact number
n = 10 · 95% interval: 0% to 28%.
One series per prompt; whiskers are 95% Wilson intervals
Read the source study ·Source and measurement limits
Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.
- Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
- Measured100%
Same prompt, 10 times: strict pass rate
Claude Sonnet 5.5 · Claude Code · Exact number
n = 10 · 95% interval: 72% to 100%.
One series per prompt; whiskers are 95% Wilson intervals
Read the source study ·Source and measurement limits
Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.
- Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
How to read this evidence ↓What this does not prove
Repeated output can be consistently wrong. This is a small repeated-prompt test, not a general model ranking.
Recommendation
Grade format apart from content.
Keep strict passes, format misses and wrong answers separate.
What we measured
- Measured91% (139/152)
Calls that passed strictly (hard set)
n = 152 · 95% interval: 86% to 95%.
Read the source study ·Source and measurement limits
The study does not publish a note for this value.
- Provider head-to-head, hard set: eight hard tasks with strict validators ·
- Repricing calculation ·
- Anthropic list prices (Claude models) ·
- OpenAI list prices ·
- Measured95% (144/152)
Calls with a correct answer, format misses included (lenient reading)
n = 152 · 95% interval: 90% to 97%.
Read the source study ·Source and measurement limits
The study does not publish a note for this value.
- Provider head-to-head, hard set: eight hard tasks with strict validators ·
- Repricing calculation ·
- Anthropic list prices (Claude models) ·
- OpenAI list prices ·
How to read this evidence ↓What this does not prove
A lenient extractor can find an answer that passes the same validator. That does not turn the original reply into a strict pass or prove general correctness.
Recommendation
Use the cheapest model that passes.
Compare cost per accepted result alongside the same quality checks.
What we measured
- Calculation$0.014
List-price cost per strict pass on hard tasks (calculation)
Claude Sonnet 5.5 · Claude Code · Cost per strict pass
n = 24 · Interval or range not published for this value.
All calls in a configuration, failures and format misses included, divided by its strict passes
Read the source study ·Source and measurement limits
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
- Provider head-to-head, hard set: eight hard tasks with strict validators ·
- Repricing calculation ·
- Anthropic list prices (Claude models) ·
- OpenAI list prices ·
- Calculation$0.028
List-price cost per strict pass on hard tasks (calculation)
Claude Opus 5.5 · Claude Code · Cost per strict pass
n = 24 · Interval or range not published for this value.
All calls in a configuration, failures and format misses included, divided by its strict passes
Read the source study ·Source and measurement limits
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
- Provider head-to-head, hard set: eight hard tasks with strict validators ·
- Repricing calculation ·
- Anthropic list prices (Claude models) ·
- OpenAI list prices ·
- Measured100%
Pass rate on eight hard tasks
Claude Sonnet 5.5 · Claude Code · Strict pass
n = 24 · 95% interval: 86% to 100%.
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Read the source study ·Source and measurement limits
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
- Provider head-to-head, hard set: eight hard tasks with strict validators ·
- Measured100%
Pass rate on eight hard tasks
Claude Opus 5.5 · Claude Code · Strict pass
n = 24 · 95% interval: 86% to 100%.
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Read the source study ·Source and measurement limits
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
- Provider head-to-head, hard set: eight hard tasks with strict validators ·
How to read this evidence ↓What this does not prove
The cost rows reprice recorded subscription tokens at historical list prices. They are calculations, not invoices or current vendor quotes. The task set has a ceiling.
Recommendation
Start at low effort.
Raise effort when your own checks show that the lower setting is insufficient.
What we measured
- Measured100%
Strict pass rate by effort on eight hard tasks
Claude Sonnet 5.5 (low) · Claude Code · Strict pass
n = 16 · 95% interval: 81% to 100%.
Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals
Read the source study ·Source and measurement limits
Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.
- Effort ladder: the hard task set at each effort level ·
- Measured100%
Strict pass rate by effort on eight hard tasks
Claude Sonnet 5.5 (high) · Claude Code · Strict pass
n = 16 · 95% interval: 81% to 100%.
Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals
Read the source study ·Source and measurement limits
Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.
- Effort ladder: the hard task set at each effort level ·
How to read this evidence ↓What this does not prove
Every configuration passed this task set. Reference cells reuse earlier calls; they are not independent replications. These results do not prove low effort is enough on harder work.
Recommendation
Keep the prompt prefix stable.
Inspect the actual cache counters and put changing content after shared context.
What we measured
- Measured99%
Share of input read from the cache, by turn in a session
Turn 2 · Claude Sonnet 5.5 · Claude Code
n = 3 · Interval or range not published for this value.
Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each
Read the source study ·Source and measurement limits
Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.
- Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
- Measured91%
Share of input read from the cache, by turn in a session
Turn 5 · Claude Sonnet 5.5 · Claude Code
n = 3 · Interval or range not published for this value.
Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each
Read the source study ·Source and measurement limits
Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.
- Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
- Measured1.6 s
Time per turn: first turn vs later turns in a cached session
Claude Sonnet 5.5 · Claude Code · Turn 1 (writes the ledger to the cache)
n = 3 · Observed range, fastest to slowest: 1.6 s to 1.8 s.
Median; whiskers = fastest and slowest turn
Read the source study ·Source and measurement limits
Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.
- Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
- Measured1.6 s
Time per turn: first turn vs later turns in a cached session
Claude Sonnet 5.5 · Claude Code · Turns 2-5 (read the ledger from the cache)
n = 12 · Observed range, fastest to slowest: 1.4 s to 5.6 s.
Median; whiskers = fastest and slowest turn
Read the source study ·Source and measurement limits
Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.
- Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
How to read this evidence ↓What this does not prove
Prefix changes were not tested. The cache shares describe these recorded sessions. Different questions, output sizes and cache lifetimes prevent a causal speed or savings claim.
Recommendation
Keep memory short and curated.
Write down current team facts that the repository does not show.
What we measured
- Measured40% (6/15)
Team-knowledge checks passed with no memory (Sonnet 5.5)
n = 15 · 95% interval: 20% to 64%.
Read the source study ·Source and measurement limits
The study does not publish a note for this value.
- Agent memory study: 8 kinds of project memory on Claude Code ·
- Measured100% (15/15)
Team-knowledge checks passed with an 11-line curated file (Sonnet 5.5)
n = 15 · 95% interval: 80% to 100%.
Read the source study ·Source and measurement limits
The study does not publish a note for this value.
- Agent memory study: 8 kinds of project memory on Claude Code ·
How to read this evidence ↓What this does not prove
One author knew the tasks in one small synthetic repository. This is an upper bound for curated memory. It does not prove that shorter files always work better; checks within a session are not independent.
Recommendation
Review what /init writes.
Check generated project instructions against the current tools and commands.
What we measured
- Measured87%
A stale README command: who still ran it?
/init CLAUDE.md · Claude Sonnet 5.5
n = 15 · 95% interval: 62% to 96%.
Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals
Read the source study ·Source and measurement limits
The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.
- Agent memory study: 8 kinds of project memory on Claude Code ·
- Measured80%
A stale README command: who still ran it?
No memory · Claude Sonnet 5.5
n = 15 · 95% interval: 55% to 93%.
Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals
Read the source study ·Source and measurement limits
The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.
- Agent memory study: 8 kinds of project memory on Claude Code ·
- Measured0%
A stale README command: who still ran it?
Curated, 11 lines · Claude Sonnet 5.5
n = 15 · 95% interval: 0% to 20%.
Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals
Read the source study ·Source and measurement limits
The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.
- Agent memory study: 8 kinds of project memory on Claude Code ·
How to read this evidence ↓What this does not prove
The generated file copied a stale README command. The no-file and generated-file intervals overlap; the comparison does not prove that /init makes performance worse.
Recommendation
Use hooks for rules and files for facts.
Enforce machine-checkable rules in code. Store team facts in reviewed memory.
What we measured
- Measured100%
Where memory helps: what the repo shows vs what only the team knows
Stop hook only · Rules the code already shows
n = 60 · 95% interval: 94% to 100%.
Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)
Read the source study ·Source and measurement limits
Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.
- Agent memory study: 8 kinds of project memory on Claude Code ·
- Measured67%
Where memory helps: what the repo shows vs what only the team knows
Stop hook only · Team knowledge only
n = 15 · 95% interval: 42% to 85%.
Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)
Read the source study ·Source and measurement limits
Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.
- Agent memory study: 8 kinds of project memory on Claude Code ·
How to read this evidence ↓What this does not prove
The hook checked the same code rules as the grader. It did not supply the missing team facts. Pooled checks within each session are not independent.
Recommendation
Route with rules or a small decision model.
Measure the full routing path and check its decision quality separately.
What we measured
- Measured1.42 µs
Time to make one routing decision
Deterministic routing policy (Agent, in process) · Decision time
n = 20,000 · Median to p95: 1.42 µs to 2.33 µs; not a confidence interval.
Median; whiskers = median to 95th percentile
Read the source study ·Source and measurement limits
The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.
- Routing overhead runs: policy microbenchmark and CLI start-up ·
- Routing runs: Jev router vs LLM routing ·
- Jev live run: 246 timed calls on the 82 routing decisions ·
- Measured2.6 s
Time to make one routing decision
Claude Sonnet 5.5 (effort low, via Claude Code) · Decision time
n = 82 · Median to p95: 2.6 s to 4.3 s; not a confidence interval.
Median; whiskers = median to 95th percentile
Read the source study ·Source and measurement limits
The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.
- Routing overhead runs: policy microbenchmark and CLI start-up ·
- Routing runs: Jev router vs LLM routing ·
- Jev live run: 246 timed calls on the 82 routing decisions ·
How to read this evidence ↓What this does not prove
An in-process policy and a model through a CLI are different routes. These timings do not isolate model compute or prove the rules are accurate. The span is median to p95, not a confidence interval.
Recommendation
Use the API for tiny calls.
Compare matched model and effort settings before choosing a route for short replies.
What we measured
- Measured1 s
CLI vs API: time for a one-line answer
OpenAI API · GPT-6.1 Sol · low · Total time
n = 5 · Observed range, fastest to slowest: 1 s to 1.9 s.
Matched cohort, fixed exact reply, 5 runs per configuration
Read the source study ·Source and measurement limits
Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.
- Provider explorer receipts: CLI vs API ·
- Measured4.2 s
CLI vs API: time for a one-line answer
Codex CLI · GPT-6.1 Sol · low · Total time
n = 5 · Observed range, fastest to slowest: 3.9 s to 4.5 s.
Matched cohort, fixed exact reply, 5 runs per configuration
Read the source study ·Source and measurement limits
Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.
- Provider explorer receipts: CLI vs API ·
How to read this evidence ↓What this does not prove
One host and network, a one-line task and small samples. These are full-route timings, not an estimate of long coding sessions or deployment latency.
Recommendation
Count every attempt and show n.
Retain failures and empty outputs in the declared denominator.
What we measured
- Measured76% (25/33)
Agent resolved, all 33 attempted instances
n = 33 · 95% interval: 59% to 87%.
Read the source study ·Source and measurement limits
Both campaigns, one attempt each, failures and empty patches included.
- Agent on SWE-bench Verified, campaign 1 (25 instances) ·
- Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances) ·
- SWE-bench Verified leaderboard, mini-SWE-agent v2 runs ·
- SWE-bench campaign rules and sample design ·
- Anthropic list prices (Claude models) ·
How to read this evidence ↓What this does not prove
This run includes failed delivery and empty patches. It compares whole systems with different harnesses. Overlapping intervals cannot establish superiority or customer value.
Recommendation
Test on your own tasks.
Build a representative task set with known pass rules before picking a model.
What we measured
- Measured100%
Pass rate on eight hard tasks
Claude Sonnet 5.5 · Claude Code · Strict pass
n = 24 · 95% interval: 86% to 100%.
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Read the source study ·Source and measurement limits
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
- Provider head-to-head, hard set: eight hard tasks with strict validators ·
- Measured100%
Pass rate on eight hard tasks
GPT-6.1 Sol (medium) · Codex CLI · Strict pass
n = 16 · 95% interval: 81% to 100%.
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Read the source study ·Source and measurement limits
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
- Provider head-to-head, hard set: eight hard tasks with strict validators ·
- Measured46%
Pass rate on eight hard tasks
Claude Haiku 4.5 · Claude Code · Strict pass
n = 24 · 95% interval: 28% to 65%.
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Read the source study ·Source and measurement limits
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
- Provider head-to-head, hard set: eight hard tasks with strict validators ·
How to read this evidence ↓What this does not prove
Several configurations reached the ceiling. A perfect observed score still has a wide interval. These single-turn tasks with tools off do not establish success on your repository.
Keep the observation and the recommendation separate.
Recorded results describe the named tasks, model, effort, route and date. An interval is not a guarantee on new work. A range shows the recorded runs; a median-to-p95 span describes that distribution. Missing spans stay unknown.
Calculations use historical list prices and recorded tokens. They do not certify a vendor invoice or current price. Reused reference cells, repeated prompts and checks inside a session are not independent replications.
The studies publish methods. The hard-task, effort and cache protocol files cannot establish that their protocols were written before inference. Read each study’s controls and caveats before applying a rule.