Evidence-led recommendations · updated

Run coding agents.
Check the evidence.

AI coding agent best practices, with the measurements and limits beside each rule.

12 recommendations from 7 recorded studies. These are our recommendations. The studies did not test every practice as an intervention.

  1. Recommendation

    Validate every answer.

    Check the delivered answer against a known reference or a task-specific test.

    What we measured

    • Measured0%

      Same prompt, 10 times: strict pass rate

      Claude Haiku 4.5 · Claude Code · Exact number

      n = 10 · 95% interval: 0% to 28%.

      One series per prompt; whiskers are 95% Wilson intervals

      Source and measurement limits

      Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

      • Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
      Read the source study ·
    • Measured100%

      Same prompt, 10 times: strict pass rate

      Claude Sonnet 5.5 · Claude Code · Exact number

      n = 10 · 95% interval: 72% to 100%.

      One series per prompt; whiskers are 95% Wilson intervals

      Source and measurement limits

      Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

      • Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
      Read the source study ·

    What this does not prove

    Repeated output can be consistently wrong. This is a small repeated-prompt test, not a general model ranking.

    How to read this evidence ↓
  2. Recommendation

    Grade format apart from content.

    Keep strict passes, format misses and wrong answers separate.

    What we measured

    • Measured91% (139/152)

      Calls that passed strictly (hard set)

      n = 152 · 95% interval: 86% to 95%.

      Source and measurement limits

      The study does not publish a note for this value.

      • Provider head-to-head, hard set: eight hard tasks with strict validators ·
      • Repricing calculation ·
      • Anthropic list prices (Claude models) ·
      • OpenAI list prices ·
      Read the source study ·
    • Measured95% (144/152)

      Calls with a correct answer, format misses included (lenient reading)

      n = 152 · 95% interval: 90% to 97%.

      Source and measurement limits

      The study does not publish a note for this value.

      • Provider head-to-head, hard set: eight hard tasks with strict validators ·
      • Repricing calculation ·
      • Anthropic list prices (Claude models) ·
      • OpenAI list prices ·
      Read the source study ·

    What this does not prove

    A lenient extractor can find an answer that passes the same validator. That does not turn the original reply into a strict pass or prove general correctness.

    How to read this evidence ↓
  3. Recommendation

    Use the cheapest model that passes.

    Compare cost per accepted result alongside the same quality checks.

    What we measured

    • Calculation$0.014

      List-price cost per strict pass on hard tasks (calculation)

      Claude Sonnet 5.5 · Claude Code · Cost per strict pass

      n = 24 · Interval or range not published for this value.

      All calls in a configuration, failures and format misses included, divided by its strict passes

      Source and measurement limits

      Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

      • Provider head-to-head, hard set: eight hard tasks with strict validators ·
      • Repricing calculation ·
      • Anthropic list prices (Claude models) ·
      • OpenAI list prices ·
      Read the source study ·
    • Calculation$0.028

      List-price cost per strict pass on hard tasks (calculation)

      Claude Opus 5.5 · Claude Code · Cost per strict pass

      n = 24 · Interval or range not published for this value.

      All calls in a configuration, failures and format misses included, divided by its strict passes

      Source and measurement limits

      Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

      • Provider head-to-head, hard set: eight hard tasks with strict validators ·
      • Repricing calculation ·
      • Anthropic list prices (Claude models) ·
      • OpenAI list prices ·
      Read the source study ·
    • Measured100%

      Pass rate on eight hard tasks

      Claude Sonnet 5.5 · Claude Code · Strict pass

      n = 24 · 95% interval: 86% to 100%.

      Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

      Source and measurement limits

      Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

      • Provider head-to-head, hard set: eight hard tasks with strict validators ·
      Read the source study ·
    • Measured100%

      Pass rate on eight hard tasks

      Claude Opus 5.5 · Claude Code · Strict pass

      n = 24 · 95% interval: 86% to 100%.

      Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

      Source and measurement limits

      Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

      • Provider head-to-head, hard set: eight hard tasks with strict validators ·
      Read the source study ·

    What this does not prove

    The cost rows reprice recorded subscription tokens at historical list prices. They are calculations, not invoices or current vendor quotes. The task set has a ceiling.

    How to read this evidence ↓
  4. Recommendation

    Start at low effort.

    Raise effort when your own checks show that the lower setting is insufficient.

    What we measured

    • Measured100%

      Strict pass rate by effort on eight hard tasks

      Claude Sonnet 5.5 (low) · Claude Code · Strict pass

      n = 16 · 95% interval: 81% to 100%.

      Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

      Source and measurement limits

      Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

      • Effort ladder: the hard task set at each effort level ·
      Read the source study ·
    • Measured100%

      Strict pass rate by effort on eight hard tasks

      Claude Sonnet 5.5 (high) · Claude Code · Strict pass

      n = 16 · 95% interval: 81% to 100%.

      Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

      Source and measurement limits

      Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

      • Effort ladder: the hard task set at each effort level ·
      Read the source study ·

    What this does not prove

    Every configuration passed this task set. Reference cells reuse earlier calls; they are not independent replications. These results do not prove low effort is enough on harder work.

    How to read this evidence ↓
  5. Recommendation

    Keep the prompt prefix stable.

    Inspect the actual cache counters and put changing content after shared context.

    What we measured

    • Measured99%

      Share of input read from the cache, by turn in a session

      Turn 2 · Claude Sonnet 5.5 · Claude Code

      n = 3 · Interval or range not published for this value.

      Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each

      Source and measurement limits

      Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.

      • Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
      Read the source study ·
    • Measured91%

      Share of input read from the cache, by turn in a session

      Turn 5 · Claude Sonnet 5.5 · Claude Code

      n = 3 · Interval or range not published for this value.

      Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each

      Source and measurement limits

      Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.

      • Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
      Read the source study ·
    • Measured1.6 s

      Time per turn: first turn vs later turns in a cached session

      Claude Sonnet 5.5 · Claude Code · Turn 1 (writes the ledger to the cache)

      n = 3 · Observed range, fastest to slowest: 1.6 s to 1.8 s.

      Median; whiskers = fastest and slowest turn

      Source and measurement limits

      Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.

      • Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
      Read the source study ·
    • Measured1.6 s

      Time per turn: first turn vs later turns in a cached session

      Claude Sonnet 5.5 · Claude Code · Turns 2-5 (read the ledger from the cache)

      n = 12 · Observed range, fastest to slowest: 1.4 s to 5.6 s.

      Median; whiskers = fastest and slowest turn

      Source and measurement limits

      Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.

      • Caching sessions and repeated prompts (Claude Code and Codex CLI) ·
      Read the source study ·

    What this does not prove

    Prefix changes were not tested. The cache shares describe these recorded sessions. Different questions, output sizes and cache lifetimes prevent a causal speed or savings claim.

    How to read this evidence ↓
  6. Recommendation

    Keep memory short and curated.

    Write down current team facts that the repository does not show.

    What we measured

    • Measured40% (6/15)

      Team-knowledge checks passed with no memory (Sonnet 5.5)

      n = 15 · 95% interval: 20% to 64%.

      Source and measurement limits

      The study does not publish a note for this value.

      • Agent memory study: 8 kinds of project memory on Claude Code ·
      Read the source study ·
    • Measured100% (15/15)

      Team-knowledge checks passed with an 11-line curated file (Sonnet 5.5)

      n = 15 · 95% interval: 80% to 100%.

      Source and measurement limits

      The study does not publish a note for this value.

      • Agent memory study: 8 kinds of project memory on Claude Code ·
      Read the source study ·

    What this does not prove

    One author knew the tasks in one small synthetic repository. This is an upper bound for curated memory. It does not prove that shorter files always work better; checks within a session are not independent.

    How to read this evidence ↓
  7. Recommendation

    Review what /init writes.

    Check generated project instructions against the current tools and commands.

    What we measured

    • Measured87%

      A stale README command: who still ran it?

      /init CLAUDE.md · Claude Sonnet 5.5

      n = 15 · 95% interval: 62% to 96%.

      Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals

      Source and measurement limits

      The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.

      • Agent memory study: 8 kinds of project memory on Claude Code ·
      Read the source study ·
    • Measured80%

      A stale README command: who still ran it?

      No memory · Claude Sonnet 5.5

      n = 15 · 95% interval: 55% to 93%.

      Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals

      Source and measurement limits

      The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.

      • Agent memory study: 8 kinds of project memory on Claude Code ·
      Read the source study ·
    • Measured0%

      A stale README command: who still ran it?

      Curated, 11 lines · Claude Sonnet 5.5

      n = 15 · 95% interval: 0% to 20%.

      Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals

      Source and measurement limits

      The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.

      • Agent memory study: 8 kinds of project memory on Claude Code ·
      Read the source study ·

    What this does not prove

    The generated file copied a stale README command. The no-file and generated-file intervals overlap; the comparison does not prove that /init makes performance worse.

    How to read this evidence ↓
  8. Recommendation

    Use hooks for rules and files for facts.

    Enforce machine-checkable rules in code. Store team facts in reviewed memory.

    What we measured

    • Measured100%

      Where memory helps: what the repo shows vs what only the team knows

      Stop hook only · Rules the code already shows

      n = 60 · 95% interval: 94% to 100%.

      Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

      Source and measurement limits

      Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

      • Agent memory study: 8 kinds of project memory on Claude Code ·
      Read the source study ·
    • Measured67%

      Where memory helps: what the repo shows vs what only the team knows

      Stop hook only · Team knowledge only

      n = 15 · 95% interval: 42% to 85%.

      Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

      Source and measurement limits

      Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

      • Agent memory study: 8 kinds of project memory on Claude Code ·
      Read the source study ·

    What this does not prove

    The hook checked the same code rules as the grader. It did not supply the missing team facts. Pooled checks within each session are not independent.

    How to read this evidence ↓
  9. Recommendation

    Route with rules or a small decision model.

    Measure the full routing path and check its decision quality separately.

    What we measured

    • Measured1.42 µs

      Time to make one routing decision

      Deterministic routing policy (Agent, in process) · Decision time

      n = 20,000 · Median to p95: 1.42 µs to 2.33 µs; not a confidence interval.

      Median; whiskers = median to 95th percentile

      Source and measurement limits

      The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

      • Routing overhead runs: policy microbenchmark and CLI start-up ·
      • Routing runs: Jev router vs LLM routing ·
      • Jev live run: 246 timed calls on the 82 routing decisions ·
      Read the source study ·
    • Measured2.6 s

      Time to make one routing decision

      Claude Sonnet 5.5 (effort low, via Claude Code) · Decision time

      n = 82 · Median to p95: 2.6 s to 4.3 s; not a confidence interval.

      Median; whiskers = median to 95th percentile

      Source and measurement limits

      The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

      • Routing overhead runs: policy microbenchmark and CLI start-up ·
      • Routing runs: Jev router vs LLM routing ·
      • Jev live run: 246 timed calls on the 82 routing decisions ·
      Read the source study ·

    What this does not prove

    An in-process policy and a model through a CLI are different routes. These timings do not isolate model compute or prove the rules are accurate. The span is median to p95, not a confidence interval.

    How to read this evidence ↓
  10. Recommendation

    Use the API for tiny calls.

    Compare matched model and effort settings before choosing a route for short replies.

    What we measured

    • Measured1 s

      CLI vs API: time for a one-line answer

      OpenAI API · GPT-6.1 Sol · low · Total time

      n = 5 · Observed range, fastest to slowest: 1 s to 1.9 s.

      Matched cohort, fixed exact reply, 5 runs per configuration

      Source and measurement limits

      Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.

      • Provider explorer receipts: CLI vs API ·
      Read the source study ·
    • Measured4.2 s

      CLI vs API: time for a one-line answer

      Codex CLI · GPT-6.1 Sol · low · Total time

      n = 5 · Observed range, fastest to slowest: 3.9 s to 4.5 s.

      Matched cohort, fixed exact reply, 5 runs per configuration

      Source and measurement limits

      Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.

      • Provider explorer receipts: CLI vs API ·
      Read the source study ·

    What this does not prove

    One host and network, a one-line task and small samples. These are full-route timings, not an estimate of long coding sessions or deployment latency.

    How to read this evidence ↓
  11. Recommendation

    Count every attempt and show n.

    Retain failures and empty outputs in the declared denominator.

    What we measured

    • Measured76% (25/33)

      Agent resolved, all 33 attempted instances

      n = 33 · 95% interval: 59% to 87%.

      Source and measurement limits

      Both campaigns, one attempt each, failures and empty patches included.

      • Agent on SWE-bench Verified, campaign 1 (25 instances) ·
      • Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances) ·
      • SWE-bench Verified leaderboard, mini-SWE-agent v2 runs ·
      • SWE-bench campaign rules and sample design ·
      • Anthropic list prices (Claude models) ·
      Read the source study ·

    What this does not prove

    This run includes failed delivery and empty patches. It compares whole systems with different harnesses. Overlapping intervals cannot establish superiority or customer value.

    How to read this evidence ↓
  12. Recommendation

    Test on your own tasks.

    Build a representative task set with known pass rules before picking a model.

    What we measured

    • Measured100%

      Pass rate on eight hard tasks

      Claude Sonnet 5.5 · Claude Code · Strict pass

      n = 24 · 95% interval: 86% to 100%.

      Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

      Source and measurement limits

      Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

      • Provider head-to-head, hard set: eight hard tasks with strict validators ·
      Read the source study ·
    • Measured100%

      Pass rate on eight hard tasks

      GPT-6.1 Sol (medium) · Codex CLI · Strict pass

      n = 16 · 95% interval: 81% to 100%.

      Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

      Source and measurement limits

      Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

      • Provider head-to-head, hard set: eight hard tasks with strict validators ·
      Read the source study ·
    • Measured46%

      Pass rate on eight hard tasks

      Claude Haiku 4.5 · Claude Code · Strict pass

      n = 24 · 95% interval: 28% to 65%.

      Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

      Source and measurement limits

      Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

      • Provider head-to-head, hard set: eight hard tasks with strict validators ·
      Read the source study ·

    What this does not prove

    Several configurations reached the ceiling. A perfect observed score still has a wide interval. These single-turn tasks with tools off do not establish success on your repository.

    How to read this evidence ↓

Keep the observation and the recommendation separate.

Recorded results describe the named tasks, model, effort, route and date. An interval is not a guarantee on new work. A range shows the recorded runs; a median-to-p95 span describes that distribution. Missing spans stay unknown.

Calculations use historical list prices and recorded tokens. They do not certify a vendor invoice or current price. Reused reference cells, repeated prompts and checks inside a session are not independent replications.

The studies publish methods. The hard-task, effort and cache protocol files cannot establish that their protocols were written before inference. Read each study’s controls and caveats before applying a rule.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.