• Best Practices
  • AI Coding Agents
  • Claude Code
  • Codex CLI

AI coding agent best practices: 12 rules, each backed by a measurement

12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.

TL;DR

Twelve rules, one measurement each. They are our recommendations, based on public data, not tests of every practice.

The numbered sections show sample sizes and uncertainty. All pass-rate intervals below are 95% Wilson intervals.

  1. Validate every answer. Haiku 4.5 passed 0 of 10 times (95% interval 0% to 28%).
  2. Grade format apart from content. 5 of 13 hard-set non-passes were format misses.
  3. Use the cheapest model that passes. Sonnet $0.0143 per strict pass, Opus $0.0282 (calculations); both 24/24 (95% interval 86% to 100% each).
  4. Start at low effort. All 11 effort configurations passed 16 of 16 (95% interval 81% to 100% each).
  5. Keep the prompt prefix stable. Turns 2 to 5 read 97% of input from the cache on average (6 sessions; range 88% to 99%).
  6. Keep memory short and curated. Team knowledge: 6 of 15 with no memory, 15 of 15 with an 11-line file (95% intervals 20%–64% and 80%–100%).
  7. Review what /init writes. 13 of 15 sessions with its file ran the broken README command (95% interval 62% to 96%).
  8. Use hooks for rules, files for facts. The Stop hook used 1.6 times the median input tokens (calculation). It used the right late-fee rate in 0/3 sessions (95% interval 0%–56%).
  9. Route with rules or a small decision model. Rules: median 1.42 µs, p95 2.33 µs (n = 20,000). Sonnet: median 2.60 s, p95 4.30 s (n = 82).
  10. Use the API for tiny calls. The Codex CLI took 3.5 times as long (calculation, n = 30) and sent a median 19,551 input tokens against 17.
  11. Count every attempt and show n. Agent resolved 25 of 33 SWE-bench instances (95% interval 59% to 87%).
  12. Test on your own tasks. 6 of 7 configurations passed every hard-set call (24/24 or 16/16; 95% intervals 86%–100% or 81%–100%).

1. Validate every answer

Check each answer against a known right answer. Haiku 4.5 gave the same wrong number (289, not 282) in 10 of 10 runs. That is 0 of 10 strict passes (95% interval 0% to 28%); Sonnet 5.5 passed 10 of 10 (95% interval 72% to 100%). See Same prompt, ten answers.

Basis: caching-consistency, 10 runs per model. Our recommendation.

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

2. Grade format apart from content

Score format and content as two checks. On eight hard tasks, 5 of 13 non-passes were right answers in the wrong format; the other 8 were wrong answers. Here, a format miss means a lenient extractor found text that passed the same validator. It does not prove general correctness. See When tasks get hard.

Basis: hard-model-head-to-head, 152 scored calls. All 13 scored non-passes came from Haiku 4.5. Every prompt said to use no code fence. Another 30 Codex attempts were blocked before inference and were not scored. Our recommendation.

3. Use the cheapest model that passes

Pick the cheapest model that passes your checks, and compare cost per strict pass, not price per token. Sonnet 5.5 and Opus 5.5 each passed all 24 hard-set calls (95% interval 86% to 100%), at $0.0143 and $0.0282 per strict pass (calculations). Haiku 4.5, the low-price model, passed 11 of 24 (95% interval 28% to 65%), at $0.0672 per pass (calculation). See Sonnet vs Opus.

Basis: hard-model-head-to-head, 24 calls per model. The provider index reports a 12.6-fold price spread for DeepSeek V4 Flash 0423 across 15 standard-tier providers. This is a calculation on third-party-reported prices from 2026-10-06, using a 3:1 input:output token mix. Endpoints can differ in precision and context. Our recommendation.

4. Start at low effort

Start at the lowest effort and raise it only when a check fails. All 11 effort configurations passed all 16 hard-set calls (95% interval 81% to 100% each). For Opus, high effort cost more than low effort but passed no more calls here. It cost $0.0212 per strict pass at low effort and $0.0337 at high (a calculation). See Does reasoning effort buy quality?

Basis: effort-ladder, 16 calls per configuration. Five cells reuse earlier calls; they are not new runs. The set hits a ceiling. The 16/16 interval reaches about 19 percentage points below 100% (calculation). This does not prove low effort is enough on harder work. Our recommendation.

5. Keep the prompt prefix stable

Keep the start of the prompt fixed and put changing text last. In Claude Code sessions, turns 2 to 5 read 97% of their input from the cache on average (24 later turns across 6 sessions; observed share range 88% to 99%). See How much does prompt caching save?

Basis: caching-consistency, 6 Claude Code sessions. The session-restart probe did not reuse the earlier session's growing cache. We did not test prefix changes, so this recommendation is not a measured effect of keeping a prefix stable. For the recorded tokens from 33 SWE-bench attempts, removing caching changes $87.23 to $343.33, about 3.9 times as much (calculation). The cost study excludes compaction calls from this repricing. Our recommendation.

  • Claude Sonnet 5.5 · Claude Code
  • Claude Opus 5.5 · Claude Code
  • GPT-6.1 Sol (medium) · Codex CLI*

* The note below the chart says what this route does not report.

5 turn in the sessions, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (medium) · Codex CLI. Claude Sonnet 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3). Claude Opus 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3).

Notesn = 3 per row

Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each

Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

6. Keep memory short and curated

Write down what your team knows and the code does not show. With no memory file, Sonnet 5.5 passed only 6 of 15 team-knowledge checks. With an 11-line curated file it passed 15 of 15 (95% intervals 20% to 64% and 80% to 100%, no overlap). See Does CLAUDE.md help?

Basis: agent-memory, 15 checks per condition, one small synthetic repository. A 211-line handbook also passed 15 of 15 checks (95% interval 80% to 100%). Its median input was 85,223 tokens, against 65,045 with curated memory (15 sessions each). These totals include cache reads. Team-knowledge checks within a session are not independent. The curated file is an upper bound: its author knew the tasks. Our recommendation.

Rules the code already shows

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Rules the folders hint at

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Team knowledge only

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 3 series: Rules the code already shows, Rules the folders hint at, Team knowledge only. Rules the code already shows: highest Curated, 11 lines 100% (95% interval 94%–100%, n 60). Lowest /init CLAUDE.md 95% (95% interval 86%–98%, n 60). All intervals overlap. Rules the folders hint at: all at 100%.

NotesWhiskers: 95% Wilson intervaln 15–60 per row6 of 8 (Rules the code already shows) at 100%: this task set cannot separate them.

Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

Source: Agent memory study: 8 kinds of project memory on Claude Code

7. Review what /init writes

Read the file that /init writes before you trust it. The README held a test command that fails on Node 25. The /init file repeated it. 13 of 15 Sonnet 5.5 sessions with that file ran it (95% interval 62% to 96%). With no file, 12 of 15 did (55% to 93%); with the curated file, 0 of 15 (0% to 20%). The /init and no-file intervals overlap. See the memory study.

Basis: agent-memory, 15 sessions per condition. Our recommendation.

8. Use hooks for rules and files for facts

Use a hook for a rule that code can check, and a file for a fact that code cannot. With the Stop hook, all 60 code-rule checks passed (95% interval 94% to 100%, pooled across 15 sessions). But 0 of 3 late-fee sessions used the right rate (0% to 56%). The median input token count was 1.6 times that of no memory (125,674 vs 78,455; calculation, 15 sessions each). See the memory study.

Basis: agent-memory, 15 sessions per condition and 3 for the late-fee task. The hook checks the same rules as the grader. Checks within a session are not independent. No-memory code-rule checks passed 57 of 60 (95% interval 86% to 98%); that interval overlaps the hook interval. Our recommendation.

9. Route with rules or a small decision model

Route with rules or a small model, not a large model behind a CLI. Rules took a median 1.42 µs (p95 2.33 µs, n = 20,000). Their model-call cost is $0 (calculation; host cost excluded). Sonnet 5.5 through the Claude Code CLI took a median 2.60 s (p95 4.30 s, n = 82). These spans run from median to p95, not confidence intervals. The recorded Jev run made 74 of 82 exact decisions (95% interval 82% to 95%). Sonnet made 77 of 82 (87% to 97%); the intervals overlap. See What does a router cost you?

Basis: routing-overhead and routing-jev-vs-llm. The rules ran in process and Sonnet through a CLI, so this compares two ways of routing, not two models. We revised the 82 cases against Jev answers and did not test the rules' accuracy. A later Jev live run measured a median 136.5 ms (p95 195.7 ms, n = 246 calls). It used a direct HTTPS route, not a CLI. The case revisions give Jev a home advantage. Our recommendation.

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

10. Use the API for tiny calls

Call the model API directly for a one-line job. The Codex CLI took a pooled median 3.9 s for a one-line answer (range 2.88 to 4.69 s, n = 15). The API took 1.1 s (0.65 to 2.23 s, n = 15). That is 3.5 times as long (calculation from the unrounded medians). These are observed ranges, not confidence intervals. Median input tokens were 19,551 against 17 (15 calls per route). See Claude Code vs Codex CLI vs the API.

Basis: cli-model-latency-tokens, 30 matched timed calls across GPT-6 Luna at none effort and GPT-6.1 Sol at low and high effort, one host. Each configuration holds 5 calls per route. The pooled ratio is directional. Our recommendation.

11. Count every attempt and show n

Report every run, failures too, with n and an interval. Agent resolved 25 of 33 SWE-bench Verified instances (76%, 95% interval 59% to 87%), with empty patches counted as failures. Eleven public runs resolved 21 to 28 of the same instances. Every interval overlaps, so this sample cannot rank Agent above or below any of them. See Why we count every failed attempt.

Basis: swe-bench-verified, 33 instances, one run each. Our recommendation.

GPT 5.2 (high)
Gemini 3 Flash (high)
GLM 5 (high)
Agent (Sonnet 5.5, full pipeline)
Claude 4.5 Sonnet (high)
Claude 4.5 Haiku (high)
Claude 4.5 Opus (high)
DeepSeek V3.2 (high)
MiniMax M2.5 (high)
Claude 4.6 Opus
Kimi K2.5 (high)
GPT 5 mini

Every interval overlaps every other: this chart does not order these rows.

12 rows. Highest GPT 5.2 (high) 85% (95% interval 69%–93%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 33 per row

Agent vs 11 public mini-SWE-agent v2 runs, one attempt each

Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

12. Test on your own tasks

Run a small set of your own tasks before you pick a model. On our 8 hard tasks, 6 of 7 configurations passed every call, so pass rate could not separate those six. Four passed 24/24 (95% interval 86% to 100% each). Two passed 16/16 (81% to 100% each). This set still hits a ceiling. Make the tasks harder, or compare cost and time. See How to read AI benchmarks honestly.

Basis: hard-model-head-to-head, 24 or 16 calls per configuration. Our recommendation.

How we measured

The benchmark library holds 16 public studies; these rules use 10, linked in the Basis lines. The sources retain failed and blocked attempts, with exclusions identified. A strict pass means the whole reply passes a deterministic validator.

The studies publish protocols. However, the available hard-task, effort and caching protocol files have creation times after their first calls. These files cannot verify a claim that the protocols were written before inference. The memory protocol predates counted sessions; its later amendment and erratum are disclosed.

CLI token costs and repricing ratios are list-price calculations, not invoices. The CLI calls ran on subscriptions. The no-cache SWE-bench calculation holds the recorded tokens fixed; it does not predict another model's outcomes.

Caveats

  • Small samples. Many cells hold 3 to 24 calls; a perfect 10 of 10 has an interval of 72% to 100%.
  • Ceilings. Six of seven hard-set configurations passed every call; each task ran one turn with tools off. The hard and effort studies share reference calls, so they are not independent replications.
  • One host, one repository. The CLI timings come from one host and network, and each CLI adds its own prompt. The memory results come from one small synthetic repository.

Run the benchmark on your own work

Agent keeps receipts for your own tasks: model, route, tokens, time, cost and validation result. Try Agent.

The data behind this post

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.