Explainer · McNemar test
The McNemar test: compare two models on the same cases
Definition
McNemar test, also called McNemar's test, compares two models on the same cases, such as two LLMs or two classifiers. It reads cases where exactly one model was right; the others cannot separate them. The exact version is a binomial test on those disagreements with a probability of 0.5.
Agent team · · 5 min read · Every number is from the public studies
Same cases, two routers (Jev: its recorded production run, one pass)
Every case once: both right, only one right, both wrong
Only the 7 discordant decisions count: 3 vs 4. Exact McNemar p = 1: no evidence of a difference.
- Both right
- 70 85%Agree and right: says nothing about which is better.
- Only Claude Haiku 4.5 right
- 3 4%Discordant: Claude Haiku 4.5 right where Jev 1.13 (TypeSafe) was wrong.
- Only Jev 1.13 (TypeSafe) right
- 4 5%Discordant: Jev 1.13 (TypeSafe) right where Claude Haiku 4.5 was wrong.
- Both wrong
- 5 6%Agree and wrong: says nothing about which is better.
One square per paired decision (n = 82). Squares start in one grid and sort into the four outcomes; their order inside a quadrant carries no meaning.
| Pair | Both right | Only first right | Only second right | Both wrong | Exact McNemar p |
|---|---|---|---|---|---|
| Claude Haiku 4.5 vs Jev 1.13 (TypeSafe) | 70 | 3 | 4 | 5 | 1 |
| Claude Sonnet 5.5 vs Jev 1.13 (TypeSafe) | 73 | 4 | 1 | 4 | 0.375 |
Counts of cases from the table; the exact McNemar p reads only the discordant casesn = 82 cases per pair
Claude Haiku 4.5 vs Jev 1.13 (TypeSafe): 3 vs 4 discordant cases of 82, exact McNemar p = 1. Claude Sonnet 5.5 vs Jev 1.13 (TypeSafe): 4 vs 1 discordant cases of 82, exact McNemar p = 0.375.
Why compare the models case by case
A hard case can lower both models’ scores. Separate Wilson intervals ignore that link. A paired test keeps it.
Our routing study gives the example. Three routers answered the same 82 typed decisions. We pair Jev’s recorded production pass from 2026-10-05 with one Claude Code CLI pass per Claude router. An exact decision has every scored question acceptable. Counts with 95% Wilson intervals:
- Jev 1.13: 74 of 82 (81.9% to 95.0%)
- Claude Haiku 4.5: 73 of 82 (80.4% to 94.1%)
- Claude Sonnet 5.5: 77 of 82 (86.5% to 97.4%)
This chart uses a later Jev live run: 221 of 246 calls across three repeats of the same 82 decisions. Its case-level 95% interval is 81.9% to 95.0%; repeats are not independent. The paired table uses the earlier pass.
All intervals overlap: no router is ahead. Does one win more disagreements?
The 2 by 2 table and a worked example
Sort cases into four cells: both right, only the first right, only the second right, both wrong. The middle cells are discordant pairs. The study’s "Same cases, two routers" table and public routing receipts give these counts for Jev’s recorded pass. The p values are calculations from the discordant cells.
| Pair | Both right | Only first | Only second | Both wrong | Exact p |
|---|---|---|---|---|---|
| Claude Haiku 4.5 vs Jev 1.13 | 70 | 3 | 4 | 5 | 1 |
| Claude Sonnet 5.5 vs Jev 1.13 | 73 | 4 | 1 | 4 | 0.375 |
Each row holds 82 cases. Haiku is right on 70 + 3 = 73 and Jev on 70 + 4 = 74. They differ by one case but disagree on 7; accuracies hide that.
Both routers were right in 70 and 73 of the 82 cases; those cases cannot inform this test. Two of four suites hit a ceiling. All three routers got every message-intent case (20 of 20, 84% to 100%) and every is-it-a-rule case (12 of 12, 76% to 100%) exactly right. Those 32 cases (calculation: 20 + 12) are both-right in every pair and cannot separate them.
How the exact test works
The null hypothesis treats disagreements as independent coin flips: either router is equally likely to be the only one right. Let b count only-first-right cases, c only-second-right cases, and d = b + c. Then b follows a binomial distribution: d trials, probability 0.5. The two-sided p value is p = min(1, 2 × P(X ≤ min(b, c))), where X has that distribution.
Both table values follow (calculations):
- Sonnet vs Jev: b = 4, c = 1, d = 5. P(X ≤ 1) = (1 + 5) / 32. So p = 2 × 6/32 = 0.375.
- Haiku vs Jev: b = 3, c = 4, d = 7. P(X ≤ 3) = (1 + 7 + 21 + 35) / 128 = 0.5. So p = 2 × 0.5 = 1.
Both match the table. All disagreements on one side give the smallest p: p = min(1, 2 × 0.5^d) (calculation). With 5 disagreements that is 2/32 = 0.0625, above 0.05. With 6 it is 2/64 = 0.03125, below 0.05. So the test needs at least 6 disagreements, and exactly 6 must all fall on one side.
The chi-square form approximates many disagreements; with few, use the exact form. In Python, scipy.stats.binomtest(min(b, c), b + c, 0.5, alternative="two-sided").pvalue gives the same p for disagreements. With none, report p = 1. Both pairs were checked.
How to read p
- p measures surprise. Under equal chances, p = 0.375 means a 4 to 1 or more lopsided split, either way, occurs 37.5% of the time (calculation: 12/32). That is common; it is not the probability that the routers are equal.
- No difference found is not "the same". The Sonnet pair has 5 disagreements, so no split could give p below 0.05. The data cannot separate them or show they are equal.
- p is not an effect size. Read the gap from the counts. Sonnet minus Jev is (4 − 1) / 82 = +3.7 points. Haiku minus Jev is (3 − 4) / 82 = −1.2 points. Both are calculations.
When to use it, and where it stops
Use it for right-or-wrong results on the same items. For independent case sets, use an unpaired test and show intervals; difficulty can differ. For numbers such as cost, use a paired test on differences.
Limits in our study:
- One sample per item. One pass per router, including Jev’s recorded production pass. Production asks for a second sample at low confidence; we did not compare that.
- Several pairs. Two pairs add chances of a false alarm. Neither passed; a correction could only raise the p values.
- A home advantage. We revised the case sets against Jev answers on 2026-10-04 and 2026-10-05. A paired test cannot undo a case set that favours one side.
- Few cases. 82 cases give wide intervals and only 7 and 5 disagreements. The test assumes independent cases; related cases can weaken that assumption.
- Report the table, not only p. Show n, the four cells and each interval. Then anyone can redo the test.
Frequently asked questions
When should I use a McNemar test instead of comparing accuracies?
Use it for right-or-wrong answers to the same cases. Accuracies hide disagreements: Claude Haiku 4.5 and Jev 1.13 differ by one exact decision (73 and 74 of 82), but disagree on 7, split 3 to 4.
What are discordant pairs?
They are the cases where exactly one of the two models was right. Claude Haiku 4.5 and Jev 1.13 had 7 (3 + 4). Claude Sonnet 5.5 and Jev 1.13 had 5 (4 + 1).
Why can overlapping intervals still hide a difference?
Separate intervals omit the case-by-case link. A paired test can detect a difference despite overlap. Here both methods find no difference. We have no example where pairing separated models that intervals did not.
How many test cases do I need for a McNemar test?
Disagreements are the count that matters. The exact test needs at least 6 to reach p below 0.05; at that minimum, all must favour one side (calculation). Five disagreements never pass, not even a 5 to 0 split (p = 0.0625). Our 82 cases gave 7 and 5 disagreements. Without a power study, we cannot say how many cases a real gap needs.
Watch the data
Jev vs Claude as a router: accuracy and cost
Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.
Transcript
- Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
- Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
- Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
- Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
- Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
- Open benchmarks: intervals, sources and every failure kept.