Do blind AI critics prefer AI pull requests over human ones?
Blind Claude and GPT critics preferred an AI worker's change over the merged human change on 9 of 12 real tasks, and 6 of 12 on the first try. Caveats inside.
TL;DR
- We paired an AI worker's pull request with the change that maintainers actually merged, on 12 real tasks. Critic models judged each pair blind, in both orders.
- On the latest attempt per task, the panel preferred the AI change on 9 of 12 tasks (75%, 95% interval 47% to 91%).
- On the first scored attempt, it was 6 of 12 (50%, interval 25% to 75%). That is the cleaner estimate, because later attempts learned from earlier ones.
- Across 132 single verdicts, 69% preferred the AI change.
- Most critics are Anthropic models and so is the worker. Same-family preference is possible, and a critic's preference is not a merge decision.
Every pair and every vote: /benchmarks/blind-review-head-to-head.
The question
"Is AI code as good as human code?" is too vague to test. So we asked a narrower one: when critics cannot see which change came from a person, which one do they prefer?
Each task is a real merged pull request. The issue becomes the ticket. The repository is frozen at the base commit. The merged change is the human reference. Agent's worker solves the ticket in a sandbox without seeing that reference. Its pull request and the merged one form a pair.
Then a panel of critic models reviews both changes without labels. Claude Opus 5, Claude Fable 5 and Claude Sonnet judged every pair. On some pairs, Codex (GPT) or GPT 5.5 joined. Each critic judges each pair twice, once in each order, so position effects can show up.
The decision rule is strict. The AI change "wins" a pair only on a strict majority of verdicts and with no panel veto. A veto is a serious issue flagged by at least two critics.
First try vs latest try
AI pull requests vs merged human pull requests, judged blind
Blind critics preferred the AI change on 9 of 12 tasks at the latest attempt and 6 of 12 at the first. Same-family bias disclosed.
Transcript
- Blind review · 12 real merged pull requests. AI change vs the merged human change. Critic models review both, unlabelled, once in each order.
- Latest attempt: the panel preferred the AI change on 9 of 12 tasks. First attempt: 6 of 12. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.
- The first attempt is the cleaner estimate: 50%, with a 95% interval from 25% to 75%. Chart: First attempt vs latest attempt (n = 5–12 each). Caveat: n = 12 tasks: intervals are wide.
- Vote by vote: 9 of 12 tasks went unanimously to the AI change, 2 unanimously to the human one. Chart: Blind panel votes per task (latest attempt) (n = 6–8 each). Caveat: 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.
- Cross-family check: OpenAI critics preferred the AI change in 10 of 12 verdicts. Small n, wide intervals. Chart: Does the judge’s model family matter? (n = 2–40 each). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
- Open benchmarks: intervals, sources and every failure kept.
- First scored attempt: 6 of 12 tasks, 50%.
- Latest attempt: 9 of 12 tasks, 75%.
- Public OSS tasks, latest: 4 of 5.
- Private tasks, latest: 5 of 7.
Why show two numbers? Because the latest attempts are not independent. They had the lessons, review replays and, on some tasks, operator answers from earlier attempts on the same task. That is how a real worker improves on a codebase. It is not how you measure a first try.
If you want one honest number, take the first attempt: a coin flip against merged human work, with a wide interval. That is still a strong result for an unassisted first try at real tickets. It is not "AI beats humans".
Pair by pair
Most pairs are not close. All nine tasks the AI change won at the latest attempt were unanimous, including fastify/session#348, h3js/h3#1533, redis/go-redis#3914 and open-telemetry/opentelemetry-go#8706. Two private Go tasks got unanimous votes for the human change. One, aio-libs/aiohttp#13122, split 4 to 4, and a tie does not count as a win under the rule.
The every-pair table on the study page shows how scores moved. redis/go-redis#3914 went from 1-5 against the AI change on the first attempt to 8-0 for it on the second. One private Go task stayed with the human change over three attempts, at 1-5, 0-6 and 0-6. Some tickets are hard, and more attempts do not always help.
What the critics scored
Each critic also scored both changes from 1 to 5 on 14 dimensions.
The AI change scored higher on every dimension, on average over the latest attempts. The largest gaps were:
- Tests: 4.51 vs 3.29
- Communication: 4.59 vs 3.58
- Production readiness: 4.03 vs 3.21
- Completeness: 4.22 vs 3.34
- Docs: 4.04 vs 3.20
The smallest gap was security, 4.38 vs 4.14.
There is a pattern there. The AI change tends to bring more tests, more docs and a fuller write-up. Critics reward that. Maintainers often merge a smaller, tighter change on purpose, and they know context a critic does not. So read "higher score" as "what a reviewer sees on the page", not as "what the project needed".
Does the judge's family matter?
This is the caveat we worry about most. The worker runs on an Anthropic model, and most critics are Anthropic models. LLM judges can prefer text that looks like their own.
Share of single verdicts that preferred the AI change, per critic:
- Claude Fable 5: 65%, n = 40
- Claude Sonnet: 67.5%, n = 40
- Claude Opus 5: 70%, n = 40
- Codex (GPT): 80%, n = 10
- GPT 5.5: 100%, n = 2
The OpenAI critics did not prefer the AI change less. That is mildly reassuring. But they judged far fewer pairs, so their intervals are wide, and verdicts within one pair are not independent. We cannot rule out family bias with this data. We can only say we looked, and the small OpenAI sample points the same way.
We also checked position bias. In 4 of 20 pairs, a critic flipped its preference when the two changes swapped places. Judging both orders is there to catch exactly that.
How we measured
- Tasks. 12 real merged pull requests: 5 public OSS tasks and 7 from one private Go service. 20 scored pairs in total, counting repeat attempts.
- Worker. Agent's AI worker in a sandbox, no access to the reference change.
- Critics. Claude Opus 5, Claude Fable 5 and Claude Sonnet on every pair, with Codex (GPT) or GPT 5.5 on some. Blind, both orders.
- Decision rule. The AI change wins on a strict majority of verdicts and no panel veto.
- Kept. Every scored attempt. A re-scoring of the same pair under the current rule replaces the earlier scoring.
- Spend. A notional $979 for worker runs and $332 for the critic panel, over all 20 scored pairs, at list prices.
Caveats
- n = 12 tasks. The intervals are wide.
- Same-family judges. Most critics and the worker are Anthropic models.
- Latest attempts had help. They had lessons, review replays and some operator answers from earlier attempts.
- Private tasks. 7 of the 12 tasks come from one private Go service. They carry neutral labels, and we publish only their language and kind.
- Position bias. 4 pairs were flagged.
- A preference is not a merge. The human change was merged, reviewed and shipped. A critic's preference says nothing about long-term maintenance cost.
What to read next
- SWE-bench Verified: Agent vs 11 public models on the same tasks
- Why we count every failed attempt
- Harness vs model: where the gains come from
Put it in front of your own reviewers
The best judge of a pull request is the team that has to live with it. Agent opens reviewed pull requests with tests and a clear write-up, and your reviewers decide. Try Agent on a real ticket.