Explainer · McNemar test

The McNemar test: compare two models on the same cases

Definition

McNemar test, also called McNemar's test, compares two models on the same cases, such as two LLMs or two classifiers. It reads cases where exactly one model was right; the others cannot separate them. The exact version is a binomial test on those disagreements with a probability of 0.5.

Agent team · · 5 min read · Every number is from the public studies

Same cases, two routers (Jev: its recorded production run, one pass)

Every case once: both right, only one right, both wrong

Only the 7 discordant decisions count: 3 vs 4. Exact McNemar p = 1: no evidence of a difference.

70Both right3Only Claude Haiku 4.5 right4Only Jev 1.13 (TypeSafe) right5Both wrong
Both right
70 85%Agree and right: says nothing about which is better.
Only Claude Haiku 4.5 right
3 4%Discordant: Claude Haiku 4.5 right where Jev 1.13 (TypeSafe) was wrong.
Only Jev 1.13 (TypeSafe) right
4 5%Discordant: Jev 1.13 (TypeSafe) right where Claude Haiku 4.5 was wrong.
Both wrong
5 6%Agree and wrong: says nothing about which is better.

One square per paired decision (n = 82). Squares start in one grid and sort into the four outcomes; their order inside a quadrant carries no meaning.

Counts of cases from the table; the exact McNemar p reads only the discordant casesn = 82 cases per pair

Claude Haiku 4.5 vs Jev 1.13 (TypeSafe): 3 vs 4 discordant cases of 82, exact McNemar p = 1. Claude Sonnet 5.5 vs Jev 1.13 (TypeSafe): 4 vs 1 discordant cases of 82, exact McNemar p = 0.375.

Why compare the models case by case

A hard case can lower both models’ scores. Separate Wilson intervals ignore that link. A paired test keeps it.

Our routing study gives the example. Three routers answered the same 82 typed decisions. We pair Jev’s recorded production pass from 2026-10-05 with one Claude Code CLI pass per Claude router. An exact decision has every scored question acceptable. Counts with 95% Wilson intervals:

  • Jev 1.13: 74 of 82 (81.9% to 95.0%)
  • Claude Haiku 4.5: 73 of 82 (80.4% to 94.1%)
  • Claude Sonnet 5.5: 77 of 82 (86.5% to 97.4%)
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

This chart uses a later Jev live run: 221 of 246 calls across three repeats of the same 82 decisions. Its case-level 95% interval is 81.9% to 95.0%; repeats are not independent. The paired table uses the earlier pass.

All intervals overlap: no router is ahead. Does one win more disagreements?

The 2 by 2 table and a worked example

Sort cases into four cells: both right, only the first right, only the second right, both wrong. The middle cells are discordant pairs. The study’s "Same cases, two routers" table and public routing receipts give these counts for Jev’s recorded pass. The p values are calculations from the discordant cells.

PairBoth rightOnly firstOnly secondBoth wrongExact p
Claude Haiku 4.5 vs Jev 1.13703451
Claude Sonnet 5.5 vs Jev 1.13734140.375

Each row holds 82 cases. Haiku is right on 70 + 3 = 73 and Jev on 70 + 4 = 74. They differ by one case but disagree on 7; accuracies hide that.

Jev 1.13 (TypeSafe)

Failure class
Message intent
Is it a rule?
Context shape

Claude Haiku 4.5

Failure class
Message intent
Is it a rule?
Context shape

Claude Sonnet 5.5

Failure class
Message intent
Is it a rule?
Context shape

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5, Claude Sonnet 5.5. Jev 1.13 (TypeSafe): highest Failure class 100% (95% interval 82%–100%, n 18). Lowest Context shape 74% (95% interval 58%–87%, n 32). All intervals overlap. Claude Haiku 4.5: highest Message intent 100% (95% interval 84%–100%, n 20). Lowest Context shape 75% (95% interval 58%–87%, n 32). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 12–32 per row3 of 4 (Jev 1.13 (TypeSafe)) at 100%: this task set cannot separate them.

A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Both routers were right in 70 and 73 of the 82 cases; those cases cannot inform this test. Two of four suites hit a ceiling. All three routers got every message-intent case (20 of 20, 84% to 100%) and every is-it-a-rule case (12 of 12, 76% to 100%) exactly right. Those 32 cases (calculation: 20 + 12) are both-right in every pair and cannot separate them.

How the exact test works

The null hypothesis treats disagreements as independent coin flips: either router is equally likely to be the only one right. Let b count only-first-right cases, c only-second-right cases, and d = b + c. Then b follows a binomial distribution: d trials, probability 0.5. The two-sided p value is p = min(1, 2 × P(X ≤ min(b, c))), where X has that distribution.

Both table values follow (calculations):

  • Sonnet vs Jev: b = 4, c = 1, d = 5. P(X ≤ 1) = (1 + 5) / 32. So p = 2 × 6/32 = 0.375.
  • Haiku vs Jev: b = 3, c = 4, d = 7. P(X ≤ 3) = (1 + 7 + 21 + 35) / 128 = 0.5. So p = 2 × 0.5 = 1.

Both match the table. All disagreements on one side give the smallest p: p = min(1, 2 × 0.5^d) (calculation). With 5 disagreements that is 2/32 = 0.0625, above 0.05. With 6 it is 2/64 = 0.03125, below 0.05. So the test needs at least 6 disagreements, and exactly 6 must all fall on one side.

The chi-square form approximates many disagreements; with few, use the exact form. In Python, scipy.stats.binomtest(min(b, c), b + c, 0.5, alternative="two-sided").pvalue gives the same p for disagreements. With none, report p = 1. Both pairs were checked.

How to read p

  • p measures surprise. Under equal chances, p = 0.375 means a 4 to 1 or more lopsided split, either way, occurs 37.5% of the time (calculation: 12/32). That is common; it is not the probability that the routers are equal.
  • No difference found is not "the same". The Sonnet pair has 5 disagreements, so no split could give p below 0.05. The data cannot separate them or show they are equal.
  • p is not an effect size. Read the gap from the counts. Sonnet minus Jev is (4 − 1) / 82 = +3.7 points. Haiku minus Jev is (3 − 4) / 82 = −1.2 points. Both are calculations.

When to use it, and where it stops

Use it for right-or-wrong results on the same items. For independent case sets, use an unpaired test and show intervals; difficulty can differ. For numbers such as cost, use a paired test on differences.

Limits in our study:

  • One sample per item. One pass per router, including Jev’s recorded production pass. Production asks for a second sample at low confidence; we did not compare that.
  • Several pairs. Two pairs add chances of a false alarm. Neither passed; a correction could only raise the p values.
  • A home advantage. We revised the case sets against Jev answers on 2026-10-04 and 2026-10-05. A paired test cannot undo a case set that favours one side.
  • Few cases. 82 cases give wide intervals and only 7 and 5 disagreements. The test assumes independent cases; related cases can weaken that assumption.
  • Report the table, not only p. Show n, the four cells and each interval. Then anyone can redo the test.

Frequently asked questions

When should I use a McNemar test instead of comparing accuracies?

Use it for right-or-wrong answers to the same cases. Accuracies hide disagreements: Claude Haiku 4.5 and Jev 1.13 differ by one exact decision (73 and 74 of 82), but disagree on 7, split 3 to 4.

What are discordant pairs?

They are the cases where exactly one of the two models was right. Claude Haiku 4.5 and Jev 1.13 had 7 (3 + 4). Claude Sonnet 5.5 and Jev 1.13 had 5 (4 + 1).

Why can overlapping intervals still hide a difference?

Separate intervals omit the case-by-case link. A paired test can detect a difference despite overlap. Here both methods find no difference. We have no example where pairing separated models that intervals did not.

How many test cases do I need for a McNemar test?

Disagreements are the count that matters. The exact test needs at least 6 to reach p below 0.05; at that minimum, all must favour one side (calculation). Five disagreements never pass, not even a 5 to 0 split (p = 0.0625). Our 82 cases gave 7 and 5 disagreements. Without a power study, we cannot say how many cases a real gap needs.

Watch the data

Live story · 34 sJev vs Claude as a router: accuracy and cost

Jev vs Claude as a router: accuracy and cost

Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.

Transcript
  1. Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
  2. Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  3. Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
  4. Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  5. Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
  6. Open benchmarks: intervals, sources and every failure kept.

The data behind this explainer

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

More explainers

  • LLM router

    What is an LLM router?

    An LLM router picks the model, effort and context for each request. What routers exist, what a decision costs in time and money, and how to judge one.

  • Reading an AI benchmark honestly

    How to read AI benchmarks honestly

    A checklist for reading AI model benchmarks: sample size, intervals, ranges, failures, calculations vs runs, and why one overall score hides what was measured.

  • Wilson confidence interval

    Wilson confidence intervals for AI benchmarks

    A Wilson interval shows the range of true pass rates that fit a benchmark result. The formula, a worked example, and why 16 of 16 still means 81% to 100%.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.