• LLM pricing
  • Claude Code
  • Prompt Caching
  • Model Routing

The cheapest way to run an AI coding agent: 7 levers from measured runs

7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.

TL;DR

  • 1. Keep the prompt cache warm. The same agent tokens cost $87.23 with caching and $343.33 without, on Sonnet 5.5 (n = 33 attempts): 3.9 times at list price (calculation).
  • 2. Pay for the cheapest model that passes. Cost per strict pass on 8 hard tasks: Sonnet $0.01435, Fable $0.09331 (n = 24 each, all passed; 95% interval 86% to 100%; calculation).
  • 3. Start at low effort. Sonnet passed 16/16 at every effort (95% interval 81% to 100%). Low cost 27% less per pass, a direction, not a ranking (calculation).
  • 4. Do not pay an LLM to route. Per 1,000 decisions, rules cost $0 and Sonnet $7.324 via the Claude Code CLI (n = 82; calculation from CLI cost estimates).
  • 5. Keep memory short. Per session, an 11-line curated file cost Sonnet $0.0818 and no memory $0.0832. Cost per full pass was 41% lower (calculation): 15/15 against 9/15. Their 95% intervals overlap (79.6% to 100% and 35.7% to 80.2%).
  • 6. Watch the CLI context tax. Median input per call, tools off, safe mode: Codex CLI 12,124 tokens (n = 40, range 12,083 to 12,163), Claude Code 2,130 (n = 90, range 2,017 to 3,828). These are ranges, not intervals.
  • 7. Shop providers for open models only. Closed models had one standard-tier price in this snapshot. Some open models differed by up to 12.6 times (calculation, OpenRouter snapshot 2026-10-06).

What is the cheapest way to run an AI coding agent, and how do you reduce Claude Code costs? Dollar figures are calculations, not bills. Most use reported tokens and list prices. Memory sessions and Claude router costs use the CLI’s cost estimates. Our calls ran on flat subscriptions. The levers overlap, so do not multiply them.

1. Keep the prompt cache warm

Effect: about 3.9 times (calculation). Our agent's 33 SWE-bench runs used 162.9M input tokens, 94.0% from the cache. At Sonnet 5.5 list price, that work costs $87.23 with caching and $343.33 without.

Across 3 Claude Code sessions of 5 turns each, Sonnet cost $0.1350 in total with the cache and $0.2698 without: 50% less (calculation). Turn 1 costs more, because a 1-hour cache write costs twice the input price. Turn 2 paid that back. The cache showed no clear speed effect. A new session did not reuse an older cache (0 of 4 later Claude sessions did, 95% interval 0% to 49%; we did not test why).

Do this: put the fixed part of the prompt first and the changing part last. See how much prompt caching saves and prompt caching, explained.

2. Pay for the cheapest model that passes

Effect: Fable cost 6.5 times Sonnet per strict pass in this sample (calculation). On 8 hard tasks, cost per strict pass was Sonnet 5.5 $0.01435 and GPT-6.1 Sol (high) $0.01514. Haiku 4.5 was $0.0672 and Fable 5.1 $0.09331 (calculations). Haiku costs less per token than Sonnet, but it passed only 11/24 (46%, 95% interval 28% to 65%). Its cost per pass was 4.7 times Sonnet's (calculation). These Claude rows used default effort; Sol (high) used Codex CLI at high effort.

These tasks hit a ceiling. Pass rate did not separate Sonnet, Opus and Fable. Each passed 24/24 (95% interval 86% to 100%), and Sol (high) passed 16/16 (81% to 100%). Cost per pass has no interval, so a cost ranking is unclear. Per-call costs overlap: Sonnet $0.0051 to $0.0424; Sol (high) $0.0092 to $0.0307 (calculation; ranges, not intervals). See the hard head-to-head.

The study also kept 30 Codex attempts blocked before any model call. They are reported separately, outside the 152 scored calls.

Do this: write a validator for your task. Run two or three models through it. Compare cost per pass, not price per token.

3. Start at low effort

Effect: on Sonnet, low cost 27% less per pass than high in this run (calculation). All 11 configurations in the effort ladder passed 16/16 (95% interval 81% to 100% each). Sonnet cost per strict pass was $0.01219 at low, $0.01352 at medium and $0.01671 at high (calculations).

Each cell has 16 calls, and the per-call costs overlap ($0.0051 to $0.0261 at low, $0.0056 to $0.0453 at high; calculation). Read 27% as a direction from this run, not a tested ranking. The set also hits a ceiling. It cannot rule out a gap of up to about 19 points, and it says nothing about harder tasks.

Do this: start at low. Raise effort when your own check fails.

4. Do not pay an LLM to route

Effect: $362.52 per 1,000 tasks at the median call count, if Sonnet routes every call through Claude Code (calculation, not measured task spend). Per 1,000 routing decisions, rules have $0 in model charges. This excludes their operating cost. Jev 1.13, a small routing model, cost $0.0337 (calculation: about 803 input tokens per decision at its published price, n = 246 live calls). Sonnet cost $7.324 and Haiku $8.924 (n = 82 decisions each; calculations from CLI cost estimates). Sonnet’s 82 receipts total $0.6005418. Per-decision estimates range from $0.00598 to $0.012472 for Sonnet and $0.003816 to $0.025989 for Haiku (ranges, not intervals). Haiku’s 82 receipts total $0.731801. Both Claude routers ran through the Claude Code CLI, Sonnet at low effort and Haiku with default thinking. The CLI adds tokens a direct API call would not, and we did not measure a direct call.

A median task makes 49.5 model calls (range 13 to 73; n = 48). Routing was off in these tasks. If a router decides every call, the median call count gives $0 with rules, $1.67 with Jev and $362.52 with Sonnet per 1,000 tasks (calculations).

Accuracy does not separate them. Jev got 221 of 246 live calls exactly right (3 repeats of 82 decisions) and Sonnet 77 of 82. The 95% case-level intervals (82% to 95% and 87% to 97%) overlap. Jev’s interval uses n = 82 cases, not 246 independent calls.

Routing pays only if the model it picks saves more than the router costs. We did not measure that saving. See what a router costs and rules vs Sonnet.

Do this: decide with rules wherever a rule can decide. For the rest, price the router first.

5. Keep memory short and curated

Calculation
  • Claude Sonnet 5.5
  • Claude Haiku 4.5 (square)
Sorted by gap, largest first.
/init CLAUDE.md
No memory
Handbook, 210 lines
Curated, 11 lines
Dreamed notes
Raw notes, 60 lines
Curated + hook
Stop hook only

Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).

List-price calculation, not a run. 8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest No memory $0.14 (n 9). Lowest Curated, 11 lines $0.082 (n 15). Claude Haiku 4.5: highest /init CLAUDE.md $0.43 (n 2). Lowest Curated + hook $0.096 (n 9).

Notesn 2–15 per row

Sum of the CLI's cost estimates for a condition, divided by its full passes

Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Effect: a curated file cost about the same per session as no memory (calculation, one small test repository). In the memory study, Sonnet cost $0.0818 per session with an 11-line curated file and $0.0832 with none (n = 15 each). Session costs ranged from $0.0476 to $0.1288 with curated memory and $0.0434 to $0.1431 without (calculations; ranges, not intervals).

Cost per full pass was 41% lower with the file, $0.0818 against $0.1386 (calculation). This arithmetic reflects the observed pass counts: 15/15 against 9/15. It does not prove that the file caused the difference. The 95% intervals (79.6% to 100% and 35.7% to 80.2%) overlap, so this is not a proven quality win.

A 210-line handbook cost $0.1006 per session (15/15 passes, 95% interval 79.6% to 100%), 23% above the curated file (calculation). Handbook costs ranged from $0.0627 to $0.1546 per session (calculation). The session-cost ranges overlap, so read this as a direction too.

Do this: write only the facts the code cannot show. We wrote the curated file knowing the tasks, so this is an upper bound.

6. Watch the CLI context tax

Effect: extra context on these short calls. In the five-task head-to-head, the Codex CLI sent a median 12,124 input tokens per call. Claude Code sent 2,130 (n = 40 Codex calls and 90 Claude Code calls). Input ranged from 12,083 to 12,163 for Codex and 2,017 to 3,828 for Claude Code (ranges, not intervals). Each row is a CLI plus a model.

The study says most of that load is the CLI's own context. Cache counters differ by route, so we do not size the dollar gap. We ran with tools and MCP servers off, in safe mode, so we did not measure what they add.

Do this: log input tokens per call on your own setup. Remove tools, servers and files the agent does not use. See the context tax.

7. Shop providers for open models, not closed ones

Calculation
ModelSpread
  1. DeepSeek V4 Flash 0423n = 15 providers
  2. DeepSeek V4 Pro 0423n = 15 providers
  3. gpt-oss-120bn = 20 providers
  4. Llama 3.3 70B Instructn = 10 providers
  5. GLM 5.3n = 32 providers
  6. Kimi K3n = 19 providers
  7. Llama 4 Maverickn = 3 providers
  8. 15 more models: the same price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
    1x

List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).

Notesn 2–32 per row

Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider

A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Effect: 1.0 times for closed models, 12.6 times for DeepSeek V4 Flash (calculation; OpenRouter's API, snapshot 2026-10-06, third-party-reported). All 15 closed models with two or more providers had one standard-tier price (15/15; 95% interval 80% to 100%). Some open-weight models differed (blended 3 input : 1 output, 10 to 20 providers each): DeepSeek V4 Flash 12.6 times, DeepSeek V4 Pro 11.2, gpt-oss-120b 6.9, Llama 3.3 70B 6.7 (calculations).

For 10 of 10 models with a vendor list price, OpenRouter's per-token price equaled it (95% interval 72% to 100%), and its Standard plan charges a 5.5% credit fee ($0.80 minimum by card; third-party-reported). These counts describe the catalog snapshot, not a random sample of all providers. The cheapest endpoint may run lower precision or a shorter context; the cheapest DeepSeek V4 Flash endpoint reports fp8. We measured no quality or speed.

Do this: for a closed model, compare fees and confirm its tier and region. For an open model, price your own token mix and check precision and context first. See all providers and OpenRouter vs Anthropic.

How we measured

  • Cost: reported tokens times vendor list price, or the CLI’s estimates for memory sessions and Claude routing decisions. Cost per pass divides total cost, failures included, by strict passes.
  • Data: 152 hard-set calls and 176 effort-ladder calls (80 reused from the hard set). Also 9 cache sessions, 200 memory sessions, 48 routing tasks, 246 live Jev calls and 265 provider endpoints. This post made no new model calls.
  • Rules: rates show n and a 95% Wilson interval. A side is ahead only when intervals do not overlap.

Caveats

  • Ceilings. 6 of 7 hard-set configurations passed every call, so pass rate cannot separate those six.
  • Scope. The hard-set Claude rows use Claude Code; the Sol rows use Codex CLI. They compare model plus route, on different days. These are short tasks, not full coding-agent jobs.
  • Small samples. The cache result rests on 3 sessions per model and the effort ladder on 16 calls per cell. The memory result rests on 15 sessions per condition in one repository.
  • Routing cost correction. The routing-overhead study’s older Sonnet projection uses $4.996 per 1,000 decisions. The receipts instead give $7.324 (calculation). This post uses the receipt total, including one-hour cache-write charges. Neither estimate is an invoice.
  • Routing. We tuned the 82 test cases on Jev's answers, which gives Jev a home advantage. Jev's 246 calls are 3 repeats of the same 82 decisions, so they are not independent.

See your own cost per pass

Agent records the model, route, tokens and cost of each call. Try Agent to see your own.

The data behind this post

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.