• Head to head
  • Claude Haiku
  • Claude Sonnet
  • Hard Tasks

Claude Haiku 4.5 vs Sonnet 5.5: all 80 comparison rows, and where the small model loses

Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.

TL;DR

  • Sonnet 5.5 was ahead of Haiku 4.5 on 14 of 80 comparison rows. Haiku was ahead on none. 31 rows tie and 35 are unclear. The 14 is 13 rows plus one memory row where Haiku did worse (calculation). 64 rows are measurements and 16 are list-price calculations.
  • Hard tasks: Sonnet passed 24 of 24 (95% interval 86% to 100%). Haiku passed 11 of 24 (46%, 28% to 65%). On five easy tasks the two tie.
  • Speed: Sonnet was ahead on all 6 speed rows that have a winner. As a router, Haiku's median decision took 12,543 ms and Sonnet's took 2,597 ms.
  • Price: Haiku lists at half of Sonnet's price per token (calculation). On the hard set it still cost 4.7 times as much per strict pass (calculation).
  • When Haiku fits: no measured row puts it ahead. Test it on your own tasks before you switch.

The side-by-side rows: /compare/claude-haiku-4-5-vs-claude-sonnet-5-5. Model pages: Claude Haiku 4.5 and Claude Sonnet 5.5.

Live story · 44 sHaiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.

Transcript
  1. Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
  2. 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  3. Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  4. Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  5. Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  6. Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  7. Open benchmarks: intervals, sources and every failure kept.

The 80 rows at a glance

A row goes to one model only when the 95% intervals, ranges or p50 to p95 bands do not overlap. A tie means the sample cannot separate the models. Unclear means the row has no interval, or its ranges overlap. Our count by study:

StudyRowsSonnet aheadHaiku aheadTieUnclear
Hard tasks62004
Easy tasks80017
Same prompt, 10 times93033
Agent memory40402016
Routing (two studies)175075
All801403135

The compare page lists one memory row as a Haiku lead because it reads a higher rate as better. That row is a bad outcome for Haiku (section 4), so we count it for Sonnet.

1. Hard tasks: Sonnet 24/24, Haiku 11/24

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Each model made 24 calls (3 repeats of 8 tasks). Sonnet passed all 24 (95% interval 86% to 100%). Haiku passed 11 (46%, 28% to 65%). The intervals do not overlap, so Sonnet is ahead.

It is also ahead on the lenient reading, which counts a right answer in the wrong format: Haiku 16/24 (67%, 47% to 82%). Study: /benchmarks/hard-model-head-to-head.

Haiku had 5 format misses (CSV parser 1, room schedule 2, SQLite query 2) and 8 wrong answers. By task, with 3 calls each, it passed 0/3 on event-loop order, room schedule and SQLite query. It passed 1/3 on DST day length, 2/3 on CSV parser and refactor, and 3/3 on interval merge and SemVer regex. Sonnet passed 3/3 on all eight tasks.

Sonnet's 24/24 is this set's ceiling. It shows that Haiku falls short, not how far Sonnet reaches.

2. Easy tasks: a tie

Claude Fable 5.1
Claude Sonnet 5.5
Claude Opus 5.5 (high)
Claude Opus 5.5
Claude Opus 5.5 (low)
Claude Haiku 4.5
GPT-6.1 Sol (high)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (low)

Every interval overlaps every other: this chart does not order these rows.

9 rows. Highest Claude Fable 5.1 · Claude Code 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row8 of 9 at 100%: this task set cannot separate them.

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

On five short tasks, Haiku passed 15/15 (95% interval 80% to 100%) and Sonnet 12/15 (80%, 55% to 93%). The intervals overlap, so this is a tie. Haiku's 15/15 is the set's ceiling. All three Sonnet misses were format misses on one arithmetic task. Each reply ended on the right answer, 2292, but added working lines that the exact-text validator rejects. Study: /benchmarks/model-head-to-head.

3. The same prompt, 10 times

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

On the exact-number prompt, Haiku passed 0/10 (0% to 28%). It gave the same wrong answer every time: 289, not the right 282. Sonnet passed 10/10 (72% to 100%).

On the JSON prompt, Haiku passed 1/10 (2% to 40%), with 9 format misses, and Sonnet passed 10/10. Sonnet is ahead on both rows. On the code-fix prompt both passed 10/10: a tie. Study: /benchmarks/caching-consistency.

4. Memory: Sonnet ahead on the handbook and /init files

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest No memory 40% (95% interval 20%–64%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest Curated + hook 100% (95% interval 72%–100%, n 10). Lowest No memory 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row5 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Changelog rule and late-fee rate, pooled · 95% Wilson intervals

Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Both models got the same memory files (Haiku 10 sessions per condition, Sonnet 15). The score is team knowledge followed: a changelog rule and a late-fee rate. Sonnet is ahead on two conditions and ties on the other six.

  • 211-line handbook: Haiku 3/10 (11% to 60%), Sonnet 15/15 (80% to 100%).
  • /init file: Haiku 1/10 (2% to 40%), Sonnet 10/15 (42% to 85%).
  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest /init CLAUDE.md 87% (95% interval 62%–96%, n 15). Lowest Curated + hook 0% (95% interval 0%–20%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest No memory 100% (95% interval 72%–100%, n 10). Lowest Curated + hook 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row

Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals

The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.

Source: Agent memory study: 8 kinds of project memory on Claude Code

One row reads the other way. This chart shows sessions that ran a stale README test command, which fails on Node 25, so a higher rate is worse. With 60 lines of raw notes, Haiku ran it in 10/10 sessions (72% to 100%) and Sonnet in 0/15 (0% to 20%). Haiku did worse, so we count this row for Sonnet. Study: /benchmarks/agent-memory.

5. Speed: Haiku was slower on every decided row

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

As a router, Haiku's median decision took 12,543 ms (p50 to p95: 12,543 to 34,481 ms; n = 82). Sonnet at low effort took 2,597 ms (2,597 to 4,298 ms; n = 82). Haiku's median is above Sonnet's p95, so Sonnet is ahead. A p50 to p95 band is not a confidence interval. Study: /benchmarks/routing-overhead.

Six speed rows have a winner, all Sonnet. Five are router rows, which are not independent tests. The sixth is the same-prompt code fix: 5.95 s against 2.67 s (ranges 4.89 to 7.33 s and 2.32 to 4.34 s; n = 10).

Other speed rows are unclear. On the hard set the medians were 39.01 s and 7.75 s, but the ranges overlap (15.27 to 75.13 s and 2.26 to 34.79 s). On the exact-number prompt Haiku's median was lower (5.06 s against 6.89 s), and those ranges overlap too. A range is not an interval.

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).

Notesn 16–24 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

On the hard set, Haiku's median call wrote 5,064 output tokens, 4,556 of them reasoning. Sonnet wrote 1,050 (585 reasoning). Time and tokens come from the same calls, so they show a link, not a cause.

6. Price: half the rate, not half the cost

List prices per million tokens:

Model (list price, calculation)InputCache readOutput
Claude Haiku 4.5$1$0.10$5
Claude Sonnet 5.5$2$0.20$10

Haiku lists at half of Sonnet's price at all three points (calculation). The charts below price the recorded tokens at these rates. They are calculations, not bills: the calls ran on subscriptions. Study: /benchmarks/cost-thought-experiments.

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

On the hard set, one strict pass cost $0.0672 with Haiku and $0.01435 with Sonnet (calculation): about 4.7 times (calculation). A failed call still costs, so Haiku's 13 misses raise its figure. The row is unclear: the dataset holds no interval for either cost.

Calculation
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 9 rows. Highest Claude Fable 5.1 · Claude Code $0.021 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0062 (n 15).

Notesn 10–15 per row

All calls in a configuration, failures included, divided by its passes

Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.

Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

On the easy set, one pass cost $0.00836 with Haiku and $0.00624 with Sonnet (calculation): about 1.3 times (calculation). Sonnet's 3 format misses count against it.

Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

As a router, 1,000 decisions cost $8.924 with Haiku and $4.996 with Sonnet (calculation): about 1.8 times (calculation). Haiku wrote 1,419 output tokens per decision and Sonnet 107. Accuracy tied: Haiku 73/82 exact (89%, 80% to 94%), Sonnet 77/82 (94%, 87% to 97%). Study: /benchmarks/routing-jev-vs-llm.

7. When Haiku fits

No measured row puts Haiku ahead. Half the price per token is its only lead in our data (calculation). All 14 cost rows show a higher figure for Haiku (calculations). All 14 are unclear: one has a range that overlaps, and the other 13 have no interval.

  • Test on your own tasks. Compare cost per pass, not price per token.
  • Check the output format. Haiku had 5 format misses on the hard set and 9 on the JSON prompt. Sonnet had none on those two, and 3 on the easy arithmetic task.

How we measured

  • Task sets: Claude Code, one turn, tools off, deterministic validators. n = 24 (hard), 15 (easy), 10 per prompt. Rates carry 95% Wilson intervals.
  • Memory: Claude Code sessions with tools on, one small repository, 8 conditions. Routing: 82 decisions per model, one CLI call each.
  • Times: medians with ranges, or p50 to p95 bands. Costs: reported tokens times list price (calculations). Ratios and counts are our arithmetic.

Caveats

  • Settings differ. Haiku ran with the CLI's default thinking in this pair. The router rows compare Haiku with thinking on against Sonnet at low effort. CLI timings include start-up and the CLI's own prompt.
  • Small samples. n is 10 to 82 per cell (194 for per-question accuracy). A hard-task cell holds 3 calls. The memory cost-per-pass rows divide by as few as 2 passes.
  • The 14 rows are not 14 tests. The strict and lenient hard rows share 24 calls. The two handbook rows show the same counts. The five router rows overlap.
  • Long agent runs are not in this pair. On 33 SWE-bench Verified tasks, a public Claude 4.5 Haiku run (bash-only harness) resolved 25 (76%, 95% interval 59% to 87%). So did Agent's full pipeline on Sonnet 5.5. These are two systems, so no row compares them: /benchmarks/swe-bench-verified.

Test the cheaper model on your own work

Agent records the model, the tokens and the result of every step, so you can see where a smaller model fails. Try Agent and compare on your own work.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.