AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
TL;DR
Twelve rules, one measurement each. They are our recommendations, based on public data, not tests of every practice.
The numbered sections show sample sizes and uncertainty. All pass-rate intervals below are 95% Wilson intervals.
- Validate every answer. Haiku 4.5 passed 0 of 10 times (95% interval 0% to 28%).
- Grade format apart from content. 5 of 13 hard-set non-passes were format misses.
- Use the cheapest model that passes. Sonnet $0.0143 per strict pass, Opus $0.0282 (calculations); both 24/24 (95% interval 86% to 100% each).
- Start at low effort. All 11 effort configurations passed 16 of 16 (95% interval 81% to 100% each).
- Keep the prompt prefix stable. Turns 2 to 5 read 97% of input from the cache on average (6 sessions; range 88% to 99%).
- Keep memory short and curated. Team knowledge: 6 of 15 with no memory, 15 of 15 with an 11-line file (95% intervals 20%–64% and 80%–100%).
- Review what /init writes. 13 of 15 sessions with its file ran the broken README command (95% interval 62% to 96%).
- Use hooks for rules, files for facts. The Stop hook used 1.6 times the median input tokens (calculation). It used the right late-fee rate in 0/3 sessions (95% interval 0%–56%).
- Route with rules or a small decision model. Rules: median 1.42 µs, p95 2.33 µs (n = 20,000). Sonnet: median 2.60 s, p95 4.30 s (n = 82).
- Use the API for tiny calls. The Codex CLI took 3.5 times as long (calculation, n = 30) and sent a median 19,551 input tokens against 17.
- Count every attempt and show n. Agent resolved 25 of 33 SWE-bench instances (95% interval 59% to 87%).
- Test on your own tasks. 6 of 7 configurations passed every hard-set call (24/24 or 16/16; 95% intervals 86%–100% or 81%–100%).
1. Validate every answer
Check each answer against a known right answer. Haiku 4.5 gave the same wrong number (289, not 282) in 10 of 10 runs. That is 0 of 10 strict passes (95% interval 0% to 28%); Sonnet 5.5 passed 10 of 10 (95% interval 72% to 100%). See Same prompt, ten answers.
Basis: caching-consistency, 10 runs per model. Our recommendation.
2. Grade format apart from content
Score format and content as two checks. On eight hard tasks, 5 of 13 non-passes were right answers in the wrong format; the other 8 were wrong answers. Here, a format miss means a lenient extractor found text that passed the same validator. It does not prove general correctness. See When tasks get hard.
Basis: hard-model-head-to-head, 152 scored calls. All 13 scored non-passes came from Haiku 4.5. Every prompt said to use no code fence. Another 30 Codex attempts were blocked before inference and were not scored. Our recommendation.
3. Use the cheapest model that passes
Pick the cheapest model that passes your checks, and compare cost per strict pass, not price per token. Sonnet 5.5 and Opus 5.5 each passed all 24 hard-set calls (95% interval 86% to 100%), at $0.0143 and $0.0282 per strict pass (calculations). Haiku 4.5, the low-price model, passed 11 of 24 (95% interval 28% to 65%), at $0.0672 per pass (calculation). See Sonnet vs Opus.
Basis: hard-model-head-to-head, 24 calls per model. The provider index reports a 12.6-fold price spread for DeepSeek V4 Flash 0423 across 15 standard-tier providers. This is a calculation on third-party-reported prices from 2026-10-06, using a 3:1 input:output token mix. Endpoints can differ in precision and context. Our recommendation.
4. Start at low effort
Start at the lowest effort and raise it only when a check fails. All 11 effort configurations passed all 16 hard-set calls (95% interval 81% to 100% each). For Opus, high effort cost more than low effort but passed no more calls here. It cost $0.0212 per strict pass at low effort and $0.0337 at high (a calculation). See Does reasoning effort buy quality?
Basis: effort-ladder, 16 calls per configuration. Five cells reuse earlier calls; they are not new runs. The set hits a ceiling. The 16/16 interval reaches about 19 percentage points below 100% (calculation). This does not prove low effort is enough on harder work. Our recommendation.
5. Keep the prompt prefix stable
Keep the start of the prompt fixed and put changing text last. In Claude Code sessions, turns 2 to 5 read 97% of their input from the cache on average (24 later turns across 6 sessions; observed share range 88% to 99%). See How much does prompt caching save?
Basis: caching-consistency, 6 Claude Code sessions. The session-restart probe did not reuse the earlier session's growing cache. We did not test prefix changes, so this recommendation is not a measured effect of keeping a prefix stable. For the recorded tokens from 33 SWE-bench attempts, removing caching changes $87.23 to $343.33, about 3.9 times as much (calculation). The cost study excludes compaction calls from this repricing. Our recommendation.
6. Keep memory short and curated
Write down what your team knows and the code does not show. With no memory file, Sonnet 5.5 passed only 6 of 15 team-knowledge checks. With an 11-line curated file it passed 15 of 15 (95% intervals 20% to 64% and 80% to 100%, no overlap). See Does CLAUDE.md help?
Basis: agent-memory, 15 checks per condition, one small synthetic repository. A 211-line handbook also passed 15 of 15 checks (95% interval 80% to 100%). Its median input was 85,223 tokens, against 65,045 with curated memory (15 sessions each). These totals include cache reads. Team-knowledge checks within a session are not independent. The curated file is an upper bound: its author knew the tasks. Our recommendation.
7. Review what /init writes
Read the file that /init writes before you trust it. The README held a test command that fails on Node 25. The /init file repeated it. 13 of 15 Sonnet 5.5 sessions with that file ran it (95% interval 62% to 96%). With no file, 12 of 15 did (55% to 93%); with the curated file, 0 of 15 (0% to 20%). The /init and no-file intervals overlap. See the memory study.
Basis: agent-memory, 15 sessions per condition. Our recommendation.
8. Use hooks for rules and files for facts
Use a hook for a rule that code can check, and a file for a fact that code cannot. With the Stop hook, all 60 code-rule checks passed (95% interval 94% to 100%, pooled across 15 sessions). But 0 of 3 late-fee sessions used the right rate (0% to 56%). The median input token count was 1.6 times that of no memory (125,674 vs 78,455; calculation, 15 sessions each). See the memory study.
Basis: agent-memory, 15 sessions per condition and 3 for the late-fee task. The hook checks the same rules as the grader. Checks within a session are not independent. No-memory code-rule checks passed 57 of 60 (95% interval 86% to 98%); that interval overlaps the hook interval. Our recommendation.
9. Route with rules or a small decision model
Route with rules or a small model, not a large model behind a CLI. Rules took a median 1.42 µs (p95 2.33 µs, n = 20,000). Their model-call cost is $0 (calculation; host cost excluded). Sonnet 5.5 through the Claude Code CLI took a median 2.60 s (p95 4.30 s, n = 82). These spans run from median to p95, not confidence intervals. The recorded Jev run made 74 of 82 exact decisions (95% interval 82% to 95%). Sonnet made 77 of 82 (87% to 97%); the intervals overlap. See What does a router cost you?
Basis: routing-overhead and routing-jev-vs-llm. The rules ran in process and Sonnet through a CLI, so this compares two ways of routing, not two models. We revised the 82 cases against Jev answers and did not test the rules' accuracy. A later Jev live run measured a median 136.5 ms (p95 195.7 ms, n = 246 calls). It used a direct HTTPS route, not a CLI. The case revisions give Jev a home advantage. Our recommendation.
10. Use the API for tiny calls
Call the model API directly for a one-line job. The Codex CLI took a pooled median 3.9 s for a one-line answer (range 2.88 to 4.69 s, n = 15). The API took 1.1 s (0.65 to 2.23 s, n = 15). That is 3.5 times as long (calculation from the unrounded medians). These are observed ranges, not confidence intervals. Median input tokens were 19,551 against 17 (15 calls per route). See Claude Code vs Codex CLI vs the API.
Basis: cli-model-latency-tokens, 30 matched timed calls across GPT-6 Luna at none effort and GPT-6.1 Sol at low and high effort, one host. Each configuration holds 5 calls per route. The pooled ratio is directional. Our recommendation.
11. Count every attempt and show n
Report every run, failures too, with n and an interval. Agent resolved 25 of 33 SWE-bench Verified instances (76%, 95% interval 59% to 87%), with empty patches counted as failures. Eleven public runs resolved 21 to 28 of the same instances. Every interval overlaps, so this sample cannot rank Agent above or below any of them. See Why we count every failed attempt.
Basis: swe-bench-verified, 33 instances, one run each. Our recommendation.
12. Test on your own tasks
Run a small set of your own tasks before you pick a model. On our 8 hard tasks, 6 of 7 configurations passed every call, so pass rate could not separate those six. Four passed 24/24 (95% interval 86% to 100% each). Two passed 16/16 (81% to 100% each). This set still hits a ceiling. Make the tasks harder, or compare cost and time. See How to read AI benchmarks honestly.
Basis: hard-model-head-to-head, 24 or 16 calls per configuration. Our recommendation.
How we measured
The benchmark library holds 16 public studies; these rules use 10, linked in the Basis lines. The sources retain failed and blocked attempts, with exclusions identified. A strict pass means the whole reply passes a deterministic validator.
The studies publish protocols. However, the available hard-task, effort and caching protocol files have creation times after their first calls. These files cannot verify a claim that the protocols were written before inference. The memory protocol predates counted sessions; its later amendment and erratum are disclosed.
CLI token costs and repricing ratios are list-price calculations, not invoices. The CLI calls ran on subscriptions. The no-cache SWE-bench calculation holds the recorded tokens fixed; it does not predict another model's outcomes.
Caveats
- Small samples. Many cells hold 3 to 24 calls; a perfect 10 of 10 has an interval of 72% to 100%.
- Ceilings. Six of seven hard-set configurations passed every call; each task ran one turn with tools off. The hard and effort studies share reference calls, so they are not independent replications.
- One host, one repository. The CLI timings come from one host and network, and each CLI adds its own prompt. The memory results come from one small synthetic repository.
What to read next
Run the benchmark on your own work
Agent keeps receipts for your own tasks: model, route, tokens, time, cost and validation result. Try Agent.