Explainer · LLM as judge
LLM as judge: how blind review works, and where it fails
Definition
LLM as judge (or LLM-as-a-judge) is an evaluation method in which one or more language models grade or compare outputs, such as answers, essays or code changes, in place of human reviewers or fixed tests. In a blind review, the judge does not know which output came from which source, and in a pairwise setup it sees both outputs in both orders, so that labels and position cannot decide the verdict. It scales to tasks that have no exact answer, but its verdicts carry the judge's own biases.
Agent team · · 4 min read · Every number is from the public studies
Interactive
A blind review panel, one task at a time
Each critic reads both changes without knowing which one an AI wrote, and picks the better one. Then the labels come off.
4 of 8 critics preferred the AI change, 4 the merged human change. All tasks: 69% (91/132) of single verdicts preferred the AI change (95% CI 61–76%).
| Task | Prefers AI change | Prefers human change | Tie | n (critics) |
|---|---|---|---|---|
| h3js/h3#1533 | 8 | 0 | 0 | 8 |
| private-go-a (private Go) | 8 | 0 | 0 | 8 |
| open-telemetry/opentelemetry-go#8706 | 8 | 0 | 0 | 8 |
| fastify/session#348 | 8 | 0 | 0 | 8 |
| private-go-d (private Go) | 6 | 0 | 0 | 6 |
| redis/go-redis#3914 | 8 | 0 | 0 | 8 |
| private-go-g (private Go) | 6 | 0 | 0 | 6 |
| private-go-f (private Go) | 6 | 0 | 0 | 6 |
| private-go-e (private Go) | 6 | 0 | 0 | 6 |
| aio-libs/aiohttp#13122 | 4 | 4 | 0 | 8 |
| private-go-c (private Go) | 0 | 6 | 0 | 6 |
| private-go-b (private Go) | 0 | 6 | 0 | 6 |
aio-libs/aiohttp#13122: 4 of 8 critics preferred the AI change, 4 the merged human change. Across all scored pairs: 69% (91/132) of single critic verdicts preferred the AI change.
The human change is the one the maintainers merged upstream. Private tasks are from one private Go service and carry neutral labels. The left-right order here is for illustration. Pairs flagged for position bias: 4 of 20.
Source: Blind review study
When to use a judge
Use a deterministic check whenever one exists: a unit test, an exact answer, a schema. A judge is for what a check cannot see: whether a code change fixes the root cause, whether it is simple, whether it fits the project's conventions, whether the description is clear.
A judge also does not replace the real outcome. A critic panel's preference is not a merge decision and not a proof of correctness.
How a blind pairwise panel works
Our blind review study compared pull requests from the Agent worker with the change the maintainers actually merged for the same issue:
- A real task. Each task is a merged pull request: the issue as the ticket, the repository at the base commit, the merged change as the human reference.
- An independent attempt. The worker solves the ticket in a sandbox without seeing the reference.
- No labels. The critics see change A and change B, with nothing that says which is human.
- Both orders. Every critic reviews the pair once in each order. A critic that flips its verdict when the order swaps shows position bias.
- A panel and a rule. Several critic models vote. The AI change wins on a strict majority of verdicts, and a serious issue flagged by at least two critics is a veto.
What the panel found
- On the first scored attempt, the panel preferred the AI change on 6 of 12 tasks (50%, 95% interval 25% to 75%).
- On the latest attempt, it preferred the AI change on 9 of 12 (75%, 47% to 91%).
- Across 132 single verdicts, 69% preferred the AI change (61% to 76%).
The first-attempt figure is the cleaner estimate: later attempts had lessons, review replays and some operator answers from earlier attempts on the same task. With 12 tasks, every interval is wide.
Known biases of LLM judges
- Position bias. A judge may prefer the first or the second output. Swapping the order and counting flips exposes it. In our panel, 4 of 20 pairs were flagged for position bias.
- Self-preference. A judge may prefer outputs from its own model family. Most of our critics are Anthropic models, and the worker runs on an Anthropic model.
- Length and style. Judges often reward longer, more polished text, more tests and more explanation, whether or not the extra material helps.
- Rubric drift. A judge with a vague rubric scores inconsistently. A rubric with named dimensions and a decision rule makes verdicts comparable.
The per-critic chart tests the family question:
The OpenAI critics preferred the AI change at least as often as the Anthropic critics (Codex: 8 of 10 verdicts; GPT 5.5: 2 of 2), but they judged far fewer pairs, so their intervals are wide. Verdicts within one pair are not independent, which makes these intervals optimistic.
How to run a fair judge
- Hide the source. Strip names, branch labels, commit messages and anything else that tells the judge which side is which.
- Swap the order and count disagreement.
- Use more than one model family as critics, and report each critic's share.
- Score dimensions, then decide by a rule. Write the decision rule before you look at results.
- Keep every scored attempt, and report the first attempt as well as the latest.
- Report n and intervals. Twelve tasks give intervals 40 to 50 points wide.
Frequently asked questions
Is LLM-as-a-judge reliable?
It is useful for qualities that tests cannot check, but it has known biases: position, self-preference and length. Blind labels, order swaps, several critic families and a fixed decision rule reduce them. Treat a judge's verdict as evidence about quality, not as proof of correctness.
What is position bias in LLM evaluation?
It is a judge's tendency to prefer an output because of where it appears, first or second. You detect it by asking for each comparison in both orders; a verdict that flips with the order is position-biased. Our panel flagged 4 of 20 pairs.
Did the critics prefer AI pull requests over human ones?
On the latest attempt, the blind panel preferred the AI change on 9 of 12 tasks (95% interval 47% to 91%); on the first scored attempt, 6 of 12 (25% to 75%). The human change was merged and shipped, and a critic preference says nothing about long-term maintenance cost.
Can a model judge its own outputs fairly?
It may prefer its own family's style. Use critics from more than one vendor and report each critic's share, so a reader can see whether the verdict depends on the judge's family.
Watch the data
AI pull requests vs merged human pull requests, judged blind
Blind critics preferred the AI change on 9 of 12 tasks at the latest attempt and 6 of 12 at the first. Same-family bias disclosed.
Transcript
- Blind review · 12 real merged pull requests. AI change vs the merged human change. Critic models review both, unlabelled, once in each order.
- Latest attempt: the panel preferred the AI change on 9 of 12 tasks. First attempt: 6 of 12. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.
- The first attempt is the cleaner estimate: 50%, with a 95% interval from 25% to 75%. Chart: First attempt vs latest attempt (n = 5–12 each). Caveat: n = 12 tasks: intervals are wide.
- Vote by vote: 9 of 12 tasks went unanimously to the AI change, 2 unanimously to the human one. Chart: Blind panel votes per task (latest attempt) (n = 6–8 each). Caveat: 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.
- Cross-family check: OpenAI critics preferred the AI change in 10 of 12 verdicts. Small n, wide intervals. Chart: Does the judge’s model family matter? (n = 2–40 each). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
- Open benchmarks: intervals, sources and every failure kept.