The cheapest LLM for classification: 100,000 decisions a day, and why Haiku cost more than Sonnet
Cost calculations for 100,000 routing decisions a day: Jev $3.37, Sonnet $732.40, Haiku $892.40. Tested on 82 cases; not general classification.
TL;DR
- Jev had the lowest calculated cost among these three router configurations. At 100,000 decisions a day, Jev 1.13 costs $3.37, Claude Sonnet 5.5 $732.40 and Claude Haiku 4.5 $892.40 (calculations). Over 30 days: $101.10, $21,972 and $26,772.
- Accuracy does not pick the winner. Jev scored 221/246 across three repeats of 82 cases; Haiku scored 73/82 and Sonnet 77/82. All 95% intervals overlap (see below).
- Haiku cost 1.2 times as much as Sonnet per decision (calculation). It wrote 1,419 output tokens a decision against 107. Haiku ran with default thinking and Sonnet at low effort.
- Time grows with volume. Median time per decision: Jev 136.5 ms (separate live run), Sonnet 2.60 s and Haiku 12.67 s (CLI). Using median time per call, 100,000 sequential decisions project to 3.8 hours with Jev, 72.2 hours with Sonnet and 14.7 days with Haiku (calculations).
- Limits: routing decisions, not general text classification. The cases favour Jev. Claude ran through a CLI.
This is a thought experiment: we did not make 100,000 calls. We scaled the calculated cost and median time per decision (no volume discounts or batching) from routing-jev-vs-llm and routing-overhead. Scaled figures say "calculation".
Accuracy first: no clear difference
Three routers answered the same 82 typed routing decisions in four sets: failure class, message intent, is it a rule, and context shape. A decision is exact when every scored question is right.
Exact decisions (95% intervals):
- Jev 1.13: 221/246 live calls (89.8%; rounded-count approximation 81.9% to 95.0%, n = 82 cases).
- Haiku 4.5: 73/82 (89.0%; 80.4% to 94.1%).
- Sonnet 5.5: 77/82 (93.9%; 86.5% to 97.4%).
Jev’s chart interval uses Wilson bounds for 74/82, after rounding 221/3 successes to 74. It approximates uncertainty on the case scale; it is not an interval from a repeat-aware statistical method. The repeats are not independent samples.
The intervals overlap, so no router is ahead. Jev’s earlier production run scored 74/82 (90.2%; 95% interval 81.9% to 95.0%). The paired exact McNemar tests find no clear difference: Haiku against Jev p = 1; Sonnet against Jev p = 0.375. No clear difference on 82 cases does not prove equal skill. Test your own cases before choosing on cost and time.
What 100,000 decisions a day cost (calculation)
Cost per 1,000 decisions: Jev $0.0337 (calculation: mean input tokens × list price, n = 246 calls), Sonnet $7.324 and Haiku $8.924 (list-price calculations from reported tokens, n = 82 calls each). These figures price all calls, including wrong decisions.
Sonnet’s receipts report 1-hour cache writes. We use their listed rate: $4 per million tokens. The study chart uses $2.50 for those writes, the 5-minute rate. That gives its lower $4.996 per 1,000 figure; this post corrects that calculation. All scaled costs below use the rounded per-1,000 figures.
The cost correction uses Sonnet’s 82-call receipt summary: mean 2 uncached input tokens, 230.8 cache-read tokens, 1,551.9 cache-write tokens and 106.6 output tokens. Using the unrounded token means gives $7.32368 per 1,000 decisions before rounding.
A rule-based policy has $0 model-call cost, but we did not score rules, so we claim no accuracy for them.
| Cost per day (calculation) | 1,000 a day | 10,000 a day | 100,000 a day |
|---|---|---|---|
| Jev 1.13 | $0.034 | $0.34 | $3.37 |
| Claude Sonnet 5.5 | $7.32 | $73.24 | $732.40 |
| Claude Haiku 4.5 | $8.92 | $89.24 | $892.40 |
| Cost per 30-day month (calculation) | 1,000 a day | 10,000 a day | 100,000 a day |
|---|---|---|---|
| Jev 1.13 | $1.01 | $10.11 | $101.10 |
| Claude Sonnet 5.5 | $219.72 | $2,197.20 | $21,972 |
| Claude Haiku 4.5 | $267.72 | $2,677.20 | $26,772 |
Sonnet costs 217.3 times as much as Jev, and Haiku 264.8 times as much (calculations from rounded costs). At 100,000 a day, Haiku costs $160.00 more than Sonnet each day (calculation).
Why the small model was not the cheap one
Haiku 4.5 is the smaller model, yet it cost more per decision. The reported token mix gives the higher list-price calculation.
| Tokens per decision | Input (incl. cache) | Output |
|---|---|---|
| Jev 1.13 (live run, n = 246) | 803 | 148 |
| Claude Sonnet 5.5 (low effort, n = 82) | 1,785 | 107 |
| Claude Haiku 4.5 (default thinking, n = 82) | 1,831 | 1,419 |
Input is similar for both Claude routers, but output is 13.3 times larger for Haiku (1,419 ÷ 107, calculation). Output tokens cost more than input tokens at list price, and thinking tokens count as output. We did not separate the model from the setting: no Haiku run with thinking off is in these figures. Jev’s live run reported a mean of 148.4 output tokens; the listed output price is zero. Sonnet’s calculation also includes its 1-hour cache-write charge.
Time at scale (calculation)
| Router | n calls | Median | p95 | Observed range (not a 95% interval) | 100,000 sequential decisions at the median (calculation) |
|---|---|---|---|---|---|
| Jev 1.13 (live run) | 246 | 136.5 ms | 195.7 ms | 100.9–297.3 ms | 3.8 hours |
| Claude Sonnet 5.5 (CLI) | 82 | 2.60 s | 4.30 s | 1.993–5.583 s | 72.2 hours |
| Claude Haiku 4.5 (CLI, default thinking) | 82 | 12.67 s | 34.41 s | 5.857–51.278 s | 14.7 days |
The deterministic policy took a median 1.42 µs (p95 2.33 µs, n = 20,000; maximum 2,538.21 µs). Its summary gives no minimum. It excludes database reads and logging. Multiplying its median by 100,000 gives 0.14 s (calculation). This is not a measured batch time.
At 100,000 decisions a day, the arrival rate is 1.16 a second (calculation). Mean call times are Jev 142.1 ms, Sonnet 2.719 s and Haiku 15.810 s. Rate × mean call time gives mean calls in flight: Jev 0.16, Sonnet 3.1 and Haiku 18.3 (calculations).
Peaks, rate limits and queues can raise the capacity you need. We did not test concurrent load. Median-based totals above are illustrations, not throughput measurements.
The chart includes Jev’s direct HTTPS run from 2026-10-06. It made 246 calls from one Mac over a home network, one at a time, with 0 errors (3 repeats of 82 decisions). Its 221/246 exact calls give 89.8%; the study shows an approximate 95% interval of 81.9% to 95.0%, from rounded 74/82 successes. Repeats are not independent samples. The run covers one 35-second window and reports no server time.
We recomputed the table from the per-call logs: medians average the two middle observations; p95 uses linear interpolation. The published chart uses nearest-rank quantiles, including the lower middle observation for p50. Its Haiku p50 is 12.543 s, while the table’s median is 12.6735 s. Its Haiku p95 is 34.481 s; the table’s interpolated p95 is 34.4133 s. The sequential totals use unrounded medians. These routes compare caller wait time, not model compute time.
Where all three struggle
The three easy sets hit a ceiling. Failure class (n = 18 cases each) scores 94% to 100%, with 95% intervals 74%–99% or 82%–100%. Message intent (n = 20 cases each) is 100% for all three (95% interval 84%–100%). Is-it-a-rule (n = 12 cases each) is 100% for all three (95% interval 76%–100%). These sets cannot separate the routers.
The weak set is context shape (n = 32 cases each):
- Jev: 71/96 calls across three repeats (74%; approximate 95% interval 58% to 87%).
- Haiku: 24/32 (75%; 95% interval 58% to 87%).
- Sonnet: 27/32 (84%; 95% interval 68% to 93%).
Jev’s chart interval rounds 71/3 successes to 24 and uses Wilson bounds for 24/32. This is a rounded-count approximation, not a repeat-aware interval. The intervals overlap. Wrong decisions can add downstream costs that this calculation excludes. Test your hardest decision type before you scale.
What to do with this
- Price your volume: decisions a day × cost per 1,000 ÷ 1,000.
- Count output tokens. Thinking tokens are output tokens.
- Test a rule first, on your own labelled cases (we used 82). Compare two or three options. Overlapping intervals leave the ranking unclear.
- Check time. At volume, time sets how many workers you need.
How we measured
- Cases. 82 typed routing decisions, 194 scored questions; 95% Wilson intervals (rounded-count approximations for Jev’s live repeats); exact McNemar test.
- Jev. A recorded production run on 2026-10-05 supplied the paired accuracy comparison. The live run on 2026-10-06 supplies Jev’s headline accuracy, tokens and timing. The list price (2026-09-23) charges for input tokens only; the live run's calculation gives the same $0.0337.
- Claude. Claude Code CLI on a subscription account, one call per decision. Cost is a calculation: list price (2026-09-21) × reported tokens, including cache reads and 1-hour cache writes.
Caveats
- Routing decisions, not general classification. Other labels or longer text change cost.
- Home advantage. We revised the case sets against Jev answers on 2026-10-04 and 2026-10-05.
- CLI overhead. The summary reports 973 ms p50 CLI overhead for Sonnet (n = 82; range 826–3,673 ms, not a 95% interval). This is CLI and harness overhead, not isolated start-up time. The CLI also adds tool-schema tokens, so the time gap is not pure model speed. To match Jev, a direct call would need to cut Sonnet's token cost by 99.5% and Haiku's by 99.6% (calculation). We did not measure a direct Claude API call.
- Limited samples. Claude made one call per decision; Jev repeated the same 82 cases three times. The intervals are wide. We did not compare confidence-gated coverage or retries.
- List prices. Claude ran on subscriptions; these amounts are calculations, not invoices. These prices come from dated snapshots, not current quotes.
- Sources. Routing receipts, routing-overhead summary, Jev live summary and Jev calls.
What to read next
- Jev vs Claude Haiku and Sonnet as a router
- What does a router cost you?
- Jev vs Claude Haiku vs Sonnet: an honest comparison
- What is an LLM router?, the routing hub, Jev vs Haiku
Price your own decisions
Agent records the cost of each model call, so you can price your own volume. Try Agent.