• Thought experiment
  • Model Routing
  • LLM pricing
  • Classification

The cheapest LLM for classification: 100,000 decisions a day, and why Haiku cost more than Sonnet

Cost calculations for 100,000 routing decisions a day: Jev $3.37, Sonnet $732.40, Haiku $892.40. Tested on 82 cases; not general classification.

TL;DR

  • Jev had the lowest calculated cost among these three router configurations. At 100,000 decisions a day, Jev 1.13 costs $3.37, Claude Sonnet 5.5 $732.40 and Claude Haiku 4.5 $892.40 (calculations). Over 30 days: $101.10, $21,972 and $26,772.
  • Accuracy does not pick the winner. Jev scored 221/246 across three repeats of 82 cases; Haiku scored 73/82 and Sonnet 77/82. All 95% intervals overlap (see below).
  • Haiku cost 1.2 times as much as Sonnet per decision (calculation). It wrote 1,419 output tokens a decision against 107. Haiku ran with default thinking and Sonnet at low effort.
  • Time grows with volume. Median time per decision: Jev 136.5 ms (separate live run), Sonnet 2.60 s and Haiku 12.67 s (CLI). Using median time per call, 100,000 sequential decisions project to 3.8 hours with Jev, 72.2 hours with Sonnet and 14.7 days with Haiku (calculations).
  • Limits: routing decisions, not general text classification. The cases favour Jev. Claude ran through a CLI.

This is a thought experiment: we did not make 100,000 calls. We scaled the calculated cost and median time per decision (no volume discounts or batching) from routing-jev-vs-llm and routing-overhead. Scaled figures say "calculation".

Accuracy first: no clear difference

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Three routers answered the same 82 typed routing decisions in four sets: failure class, message intent, is it a rule, and context shape. A decision is exact when every scored question is right.

Exact decisions (95% intervals):

  • Jev 1.13: 221/246 live calls (89.8%; rounded-count approximation 81.9% to 95.0%, n = 82 cases).
  • Haiku 4.5: 73/82 (89.0%; 80.4% to 94.1%).
  • Sonnet 5.5: 77/82 (93.9%; 86.5% to 97.4%).

Jev’s chart interval uses Wilson bounds for 74/82, after rounding 221/3 successes to 74. It approximates uncertainty on the case scale; it is not an interval from a repeat-aware statistical method. The repeats are not independent samples.

The intervals overlap, so no router is ahead. Jev’s earlier production run scored 74/82 (90.2%; 95% interval 81.9% to 95.0%). The paired exact McNemar tests find no clear difference: Haiku against Jev p = 1; Sonnet against Jev p = 0.375. No clear difference on 82 cases does not prove equal skill. Test your own cases before choosing on cost and time.

What 100,000 decisions a day cost (calculation)

Cost per 1,000 decisions: Jev $0.0337 (calculation: mean input tokens × list price, n = 246 calls), Sonnet $7.324 and Haiku $8.924 (list-price calculations from reported tokens, n = 82 calls each). These figures price all calls, including wrong decisions.

Sonnet’s receipts report 1-hour cache writes. We use their listed rate: $4 per million tokens. The study chart uses $2.50 for those writes, the 5-minute rate. That gives its lower $4.996 per 1,000 figure; this post corrects that calculation. All scaled costs below use the rounded per-1,000 figures.

The cost correction uses Sonnet’s 82-call receipt summary: mean 2 uncached input tokens, 230.8 cache-read tokens, 1,551.9 cache-write tokens and 106.6 output tokens. Using the unrounded token means gives $7.32368 per 1,000 decisions before rounding.

A rule-based policy has $0 model-call cost, but we did not score rules, so we claim no accuracy for them.

Cost per day (calculation)1,000 a day10,000 a day100,000 a day
Jev 1.13$0.034$0.34$3.37
Claude Sonnet 5.5$7.32$73.24$732.40
Claude Haiku 4.5$8.92$89.24$892.40
Cost per 30-day month (calculation)1,000 a day10,000 a day100,000 a day
Jev 1.13$1.01$10.11$101.10
Claude Sonnet 5.5$219.72$2,197.20$21,972
Claude Haiku 4.5$267.72$2,677.20$26,772

Sonnet costs 217.3 times as much as Jev, and Haiku 264.8 times as much (calculations from rounded costs). At 100,000 a day, Haiku costs $160.00 more than Sonnet each day (calculation).

Why the small model was not the cheap one

Haiku 4.5 is the smaller model, yet it cost more per decision. The reported token mix gives the higher list-price calculation.

Tokens per decisionInput (incl. cache)Output
Jev 1.13 (live run, n = 246)803148
Claude Sonnet 5.5 (low effort, n = 82)1,785107
Claude Haiku 4.5 (default thinking, n = 82)1,8311,419

Input is similar for both Claude routers, but output is 13.3 times larger for Haiku (1,419 ÷ 107, calculation). Output tokens cost more than input tokens at list price, and thinking tokens count as output. We did not separate the model from the setting: no Haiku run with thinking off is in these figures. Jev’s live run reported a mean of 148.4 output tokens; the listed output price is zero. Sonnet’s calculation also includes its 1-hour cache-write charge.

Time at scale (calculation)

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Routern callsMedianp95Observed range (not a 95% interval)100,000 sequential decisions at the median (calculation)
Jev 1.13 (live run)246136.5 ms195.7 ms100.9–297.3 ms3.8 hours
Claude Sonnet 5.5 (CLI)822.60 s4.30 s1.993–5.583 s72.2 hours
Claude Haiku 4.5 (CLI, default thinking)8212.67 s34.41 s5.857–51.278 s14.7 days

The deterministic policy took a median 1.42 µs (p95 2.33 µs, n = 20,000; maximum 2,538.21 µs). Its summary gives no minimum. It excludes database reads and logging. Multiplying its median by 100,000 gives 0.14 s (calculation). This is not a measured batch time.

At 100,000 decisions a day, the arrival rate is 1.16 a second (calculation). Mean call times are Jev 142.1 ms, Sonnet 2.719 s and Haiku 15.810 s. Rate × mean call time gives mean calls in flight: Jev 0.16, Sonnet 3.1 and Haiku 18.3 (calculations).

Peaks, rate limits and queues can raise the capacity you need. We did not test concurrent load. Median-based totals above are illustrations, not throughput measurements.

The chart includes Jev’s direct HTTPS run from 2026-10-06. It made 246 calls from one Mac over a home network, one at a time, with 0 errors (3 repeats of 82 decisions). Its 221/246 exact calls give 89.8%; the study shows an approximate 95% interval of 81.9% to 95.0%, from rounded 74/82 successes. Repeats are not independent samples. The run covers one 35-second window and reports no server time.

We recomputed the table from the per-call logs: medians average the two middle observations; p95 uses linear interpolation. The published chart uses nearest-rank quantiles, including the lower middle observation for p50. Its Haiku p50 is 12.543 s, while the table’s median is 12.6735 s. Its Haiku p95 is 34.481 s; the table’s interpolated p95 is 34.4133 s. The sequential totals use unrounded medians. These routes compare caller wait time, not model compute time.

Where all three struggle

Jev 1.13 (TypeSafe)

Failure class
Message intent
Is it a rule?
Context shape

Claude Haiku 4.5

Failure class
Message intent
Is it a rule?
Context shape

Claude Sonnet 5.5

Failure class
Message intent
Is it a rule?
Context shape

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5, Claude Sonnet 5.5. Jev 1.13 (TypeSafe): highest Failure class 100% (95% interval 82%–100%, n 18). Lowest Context shape 74% (95% interval 58%–87%, n 32). All intervals overlap. Claude Haiku 4.5: highest Message intent 100% (95% interval 84%–100%, n 20). Lowest Context shape 75% (95% interval 58%–87%, n 32). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 12–32 per row3 of 4 (Jev 1.13 (TypeSafe)) at 100%: this task set cannot separate them.

A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

The three easy sets hit a ceiling. Failure class (n = 18 cases each) scores 94% to 100%, with 95% intervals 74%–99% or 82%–100%. Message intent (n = 20 cases each) is 100% for all three (95% interval 84%–100%). Is-it-a-rule (n = 12 cases each) is 100% for all three (95% interval 76%–100%). These sets cannot separate the routers.

The weak set is context shape (n = 32 cases each):

  • Jev: 71/96 calls across three repeats (74%; approximate 95% interval 58% to 87%).
  • Haiku: 24/32 (75%; 95% interval 58% to 87%).
  • Sonnet: 27/32 (84%; 95% interval 68% to 93%).

Jev’s chart interval rounds 71/3 successes to 24 and uses Wilson bounds for 24/32. This is a rounded-count approximation, not a repeat-aware interval. The intervals overlap. Wrong decisions can add downstream costs that this calculation excludes. Test your hardest decision type before you scale.

What to do with this

  1. Price your volume: decisions a day × cost per 1,000 ÷ 1,000.
  2. Count output tokens. Thinking tokens are output tokens.
  3. Test a rule first, on your own labelled cases (we used 82). Compare two or three options. Overlapping intervals leave the ranking unclear.
  4. Check time. At volume, time sets how many workers you need.

How we measured

  • Cases. 82 typed routing decisions, 194 scored questions; 95% Wilson intervals (rounded-count approximations for Jev’s live repeats); exact McNemar test.
  • Jev. A recorded production run on 2026-10-05 supplied the paired accuracy comparison. The live run on 2026-10-06 supplies Jev’s headline accuracy, tokens and timing. The list price (2026-09-23) charges for input tokens only; the live run's calculation gives the same $0.0337.
  • Claude. Claude Code CLI on a subscription account, one call per decision. Cost is a calculation: list price (2026-09-21) × reported tokens, including cache reads and 1-hour cache writes.

Caveats

  • Routing decisions, not general classification. Other labels or longer text change cost.
  • Home advantage. We revised the case sets against Jev answers on 2026-10-04 and 2026-10-05.
  • CLI overhead. The summary reports 973 ms p50 CLI overhead for Sonnet (n = 82; range 826–3,673 ms, not a 95% interval). This is CLI and harness overhead, not isolated start-up time. The CLI also adds tool-schema tokens, so the time gap is not pure model speed. To match Jev, a direct call would need to cut Sonnet's token cost by 99.5% and Haiku's by 99.6% (calculation). We did not measure a direct Claude API call.
  • Limited samples. Claude made one call per decision; Jev repeated the same 82 cases three times. The intervals are wide. We did not compare confidence-gated coverage or retries.
  • List prices. Claude ran on subscriptions; these amounts are calculations, not invoices. These prices come from dated snapshots, not current quotes.
  • Sources. Routing receipts, routing-overhead summary, Jev live summary and Jev calls.

Price your own decisions

Agent records the cost of each model call, so you can price your own volume. Try Agent.

The data behind this post

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.