• Statistics
  • Evaluation
  • Methodology
  • Model Routing

Why a paired test: comparing two LLMs on the same cases with McNemar

Only 5 to 7 of 82 routing cases split Claude and Jev. The exact McNemar test uses just those cases: p = 0.375 and p = 1. The method, with the arithmetic.

TL;DR

  • When two models answer the same cases, compare them case by case. The exact McNemar test ignores cases where both models have the same right-or-wrong score. It asks: do the disagreements split evenly?
  • Worked example from our routing study: Claude Haiku 4.5 and Jev 1.13 disagreed on 7 of 82 cases (3 to 4), exact p = 1. Claude Sonnet 5.5 and Jev disagreed on 5 of 82 (4 to 1), exact p = 0.375. Our hand calculation gives the same values.
  • Neither pair shows a statistically significant difference. The 95% intervals overlap too.
  • A property of the test (calculation): 5 disagreements give p = 0.0625 even at a 5 to 0 split. You need at least 6 disagreements to go below 0.05, and with exactly 6 they must all fall on one side.
  • Use McNemar for the same cases and a right-or-wrong outcome. For rates on different cases, use intervals.

Disclosure: I build Agent, which routes work between models. Jev is one of the routers here, and the case sets favour it (see Caveats).

Two ways to compare two models

Way 1: separate intervals. Give each model its rate and its 95% Wilson interval. If the intervals do not overlap, one side is ahead. For different case sets, differences in difficulty can confound the comparison. Intervals do not correct that.

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

For the paired example, we use Jev’s recorded production pass: 74 of 82 exact decisions (95% Wilson interval 81.9% to 95.0%). Claude Haiku 4.5 scored 73 of 82 (80.4% to 94.1%). Claude Sonnet 5.5 scored 77 of 82 (86.5% to 97.4%). All three 95% intervals overlap, so intervals cannot rank these routers. Exact means every scored question in a case was acceptable.

The chart uses a newer Jev live run: 221 of 246 calls across 3 repeats of those 82 cases. Its pooled rate is 89.8% (case-level 95% interval 81.9% to 95.0%). Repeats are not independent cases. The paired tables below use the recorded pass, not these live repeats.

Way 2: paired cases. Each router saw the same 82 cases, so we know which router was right on each case. Separate intervals ignore that. A paired test uses it. So overlapping intervals do not always mean a tie. A paired test can separate two models that agree on most cases.

The 2x2 table: only disagreements count

Claude Haiku 4.5 against Jev:

Jev rightJev wrong
Haiku right70 (both right)3 (only Haiku)
Haiku wrong4 (only Jev)5 (both wrong)

Claude Sonnet 5.5 against Jev:

Jev rightJev wrong
Sonnet right73 (both right)4 (only Sonnet)
Sonnet wrong1 (only Jev)4 (both wrong)

Each table holds 82 cases. The cells match the published scores: Haiku 70 + 3 = 73, Sonnet 73 + 4 = 77, and Jev 70 + 4 = 74 and 73 + 1 = 74.

Why do only the off-diagonal cells count? Call them b (only the first model right) and c (only the second model right). The accuracy gap is (b − c) / n. A case that both models got right adds one point to each score. A case that both got wrong adds nothing to either. Both kinds cancel out of the gap.

  • Haiku minus Jev: (3 − 4) / 82 = −1.2 points (calculation).
  • Sonnet minus Jev: (4 − 1) / 82 = +3.7 points (calculation).

The exact test, step by step

The null hypothesis says: when only one model is right, either model is equally likely to be that model. Conditional on d = b + c such cases, X counts those where the first model is right. Under the null, X follows a binomial distribution with probability 0.5. This assumes independent case pairs. Add the probabilities for 0 to k = min(b, c), double the sum and cap it at 1:

p = min(1, 2 × P(X ≤ k)), where X ~ Binomial(d, 0.5)

Haiku vs Jev (calculation). b = 3 and c = 4, so d = 7 and k = 3. Each specific assignment of the 7 disagreements has probability 0.5^7 = 1/128. The binomial coefficients count assignments for each split.

  • P(X ≤ 3) = (1 + 7 + 21 + 35) / 128 = 64/128 = 0.5.
  • p = 2 × 0.5 = 1. The study table says 1.

Sonnet vs Jev (calculation). b = 4 and c = 1, so d = 5 and k = 1. Each specific assignment of the 5 disagreements has probability 0.5^5 = 1/32.

  • P(X ≤ 1) = (1 + 5) / 32 = 6/32 = 0.1875.
  • p = 2 × 0.1875 = 0.375. The study table says 0.375.

How to read p = 0.375: suppose the routers were equally good where they differ. A split of 4 to 1 or worse, in either direction, would then occur 37.5% of the time. The test does not reject equal accuracy at the 0.05 level. It does not prove equal accuracy.

How many disagreements do you need?

The smallest p comes when every disagreement falls on one side. For d ≥ 1, p = min(1, 2 × 0.5^d). With no disagreements, p = 1.

Disagreements (d)Most lopsided splitSmallest two-sided p (calculation)
55 to 02 × 0.5^5 = 0.0625
66 to 02 × 0.5^6 = 0.03125

With 5 disagreements, no split can go below 0.05. You need at least 6, and with exactly 6 they must all fall on one side.

So the Sonnet pair could not show a difference at the 0.05 level, whatever the split. This is a property of the test, not a result about the routers. When two models agree on most cases, you need more cases to collect enough disagreements.

What the routing study says

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 97% (95% interval 94%–99%, n 194). Lowest Claude Haiku 4.5 94% (95% interval 90%–97%, n 194). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 194 per row

Each open question the router was asked; an unanswered question counts as wrong

Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

  • Exact decisions: neither pair shows a statistically significant difference (p = 1 and p = 0.375). The intervals overlap too.
  • Per-question accuracy: the chart uses Jev’s live run: 552 of 582 answers across 3 repeats (94.8%). Its case-set-scale 95% Wilson interval is 90.8% to 97.2%, taken at n = 194 questions, not 582 independent answers. Haiku scored 183 of 194 (95% interval 90.1% to 96.8%). Sonnet scored 189 of 194 (94.1% to 98.9%). Jev’s recorded pass scored 185 of 194 (91.4% to 97.5%). These intervals overlap.

The dataset has no paired counts per question, so we show intervals only. Questions within a case can also be related; these question-level intervals do not account for that.

Sonnet’s observed +3.7 percentage-point gap over the recorded Jev pass is a calculation. It rests on a 4 to 1 split of 5 cases. Neither the paired test nor the intervals establish a lead.

Accuracy does not rank these routers. A list-price calculation gives $0.0337 per 1,000 decisions for Jev, $7.32 for Sonnet and $8.92 for Haiku. Jev’s estimate uses 246 live calls; each Claude estimate uses 82 calls. These are token-based calculations, not invoices; the dataset gives no cost interval or range.

Sonnet’s calculation uses the one-hour cache writes in all 82 CLI receipts. It matches their reported list-price cost. The dataset’s $5.00 estimate uses the five-minute write rate and needs correction.

Jev used a direct API route. The Claude routers ran through a subscription CLI, which adds tool-schema and thinking tokens. This compares route-and-model configurations, not model prices alone. See Jev vs Claude Haiku and Sonnet and Jev 1.13 vs Claude Sonnet 5.5.

When to use McNemar

Use the exact McNemar test for two models, the same cases and a right-or-wrong outcome per case. Use something else here:

  • Different case sets. There are no pairs. Show rates with Wilson intervals, and disclose differences in case difficulty.
  • A number per case, such as cost or time. Use a paired test on the per-case differences, such as a sign test or a Wilcoxon signed-rank test.
  • Three or more models. Preselect the comparisons or adjust for multiple tests. More tests give more chances of a false alarm.

Report n, the interval and the raw counts with every test. See how to read benchmarks honestly.

How we measured

  • Data: the "Same cases, two routers" table of the routing study: four cells and an exact p value per pair, 82 cases.
  • Routers: Jev numbers come from a recorded production run with the same case versions and runner. We did not re-measure Jev for this post. The Claude routers ran through the Claude Code CLI, one call per decision. Haiku 4.5 ran with the CLI default extended thinking. Sonnet 5.5 ran at effort low.
  • Statistics: 95% Wilson intervals and the exact two-sided McNemar test. Our recomputed p values match the dataset. These two tests have no multiple-test correction.

Caveats

  • Home advantage. We revised the case sets and question wording against Jev answers on 2026-10-04 and 2026-10-05.
  • Small sample. 82 cases give only 7 and 5 disagreements, so the test has little power.
  • Ceiling in some case sets. Jev and both Claude routers got every message-intent and is-it-a-rule case right. Those sets cannot separate them on accuracy.
  • No difference is not the same as equal. A high p says the data cannot separate the routers. It does not prove they are equally accurate.
  • One sample per decision. Production asks a second sample when confidence is low. This test omits that.

Measure your own models with receipts

Agent records the model, the route, the tokens, the time and the result of every step. Your own n and intervals then come from real runs. Try Agent on your own work.

The data behind this post

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.