AI pull requests vs merged human pull requests, judged blind
When critics cannot see which change came from a person, do they prefer the AI worker’s pull request or the one the maintainers merged?
Published · Updated · 4 charts · Download the data or a carousel
75%
The answer
On the latest attempt per task, the blind panel preferred the AI change on 9 of 12 tasks (75%, 95% interval 47% to 91%). On the first scored attempt it was 6 of 12 (25% to 75%). Later attempts learned from earlier ones, so the first-attempt figure is the cleaner estimate. Public OSS tasks: 4 of 5; private Go tasks: 5 of 7. Across 132 single verdicts, 69% preferred the AI change. A critic panel's preference is not a merge decision and not a correctness proof.
Live story
Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.
AI pull requests vs merged human pull requests, judged blind
Blind critics preferred the AI change on 9 of 12 tasks at the latest attempt and 6 of 12 at the first. Same-family bias disclosed.
Transcript
- Blind review · 12 real merged pull requests. AI change vs the merged human change. Critic models review both, unlabelled, once in each order.
- Latest attempt: the panel preferred the AI change on 9 of 12 tasks. First attempt: 6 of 12. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.
- The first attempt is the cleaner estimate: 50%, with a 95% interval from 25% to 75%. Chart: First attempt vs latest attempt (n = 5–12 each). Caveat: n = 12 tasks: intervals are wide.
- Vote by vote: 9 of 12 tasks went unanimously to the AI change, 2 unanimously to the human one. Chart: Blind panel votes per task (latest attempt) (n = 6–8 each). Caveat: 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.
- Cross-family check: OpenAI critics preferred the AI change in 10 of 12 verdicts. Small n, wide intervals. Chart: Does the judge’s model family matter? (n = 2–40 each). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
- Open benchmarks: intervals, sources and every failure kept.
Key numbers
50% (6/12)
Tasks where the panel preferred the AI change (first scored attempt)
95% CI 25%–75% · n = 12
60% (12/20)
All scored pairs where the panel preferred the AI change
95% CI 39%–78% · n = 20
80% (4/5)
Public OSS tasks, latest attempt
95% CI 38%–96% · n = 5
69% (91/132)
Single critic verdicts that preferred the AI change
95% CI 61%–76% · n = 132
$979 + $332
Notional spend: worker runs + critic panel
n = 20
4 of 20
Pairs flagged for position bias
n = 20
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
Every interval overlaps every other: this chart does not order these rows.
| Item | AI preferred | 95% interval | n |
|---|---|---|---|
| First scored attempt | 50% | 25%–75% | 12 |
| Latest attempt | 75% | 47%–91% | 12 |
| Public OSS tasks, latest | 80% | 38%–96% | 5 |
| Private tasks, latest | 71% | 36%–92% | 7 |
4 rows. Highest Public OSS tasks, latest 80% (95% interval 38%–96%, n 5). Lowest First scored attempt 50% (95% interval 25%–75%, n 12). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 5–12 per row
Share of tasks where the blind panel preferred the AI change
Later attempts had the earlier attempts’ lessons, review replays and, on some tasks, operator answers. They are not independent first tries.
Source: Blind review panel: AI worker change vs merged human change
- Prefers AI change
- Prefers human change
- Tie
One square per verdict; counts at the right are exact and in legend order.
| Item | Prefers AI change | Prefers human change | Tie | n |
|---|---|---|---|---|
| h3js/h3#1533 | 8 | 0 | 0 | 8 |
| private-go-a (private Go) | 8 | 0 | 0 | 8 |
| open-telemetry/opentelemetry-go#8706 | 8 | 0 | 0 | 8 |
| fastify/session#348 | 8 | 0 | 0 | 8 |
| private-go-d (private Go) | 6 | 0 | 0 | 6 |
| redis/go-redis#3914 | 8 | 0 | 0 | 8 |
| private-go-g (private Go) | 6 | 0 | 0 | 6 |
| private-go-f (private Go) | 6 | 0 | 0 | 6 |
| private-go-e (private Go) | 6 | 0 | 0 | 6 |
| aio-libs/aiohttp#13122 | 4 | 4 | 0 | 8 |
| private-go-c (private Go) | 0 | 6 | 0 | 6 |
| private-go-b (private Go) | 0 | 6 | 0 | 6 |
12 rows, 3 series: Prefers AI change, Prefers human change, Tie. Prefers AI change: highest h3js/h3#1533 8 (n 8). Lowest private-go-b (private Go) 0 (n 6). Prefers human change: highest private-go-c (private Go) 6 (n 6). Lowest private-go-e (private Go) 0 (n 6).
Notesn 6–8 per row
Each critic judges both orders without labels
The human change is the one the maintainers merged upstream. Private tasks are from one private Go service and carry neutral labels.
Source: Blind review panel: AI worker change vs merged human change
- AI change
- Merged human change (square)
Gap labels, Merged human change vs AI change: Merged human change is x points higher (+) or lower (−) than AI change, calculated from the two values shown.
| Item | AI change | Merged human change | n |
|---|---|---|---|
| Correctness | 4.4 | 3.86 | 12 |
| Root cause | 4.46 | 3.97 | 12 |
| Edge cases | 3.95 | 3.49 | 12 |
| Tests | 4.51 | 3.29 | 12 |
| Completeness | 4.22 | 3.34 | 12 |
| Compatibility | 4.39 | 3.74 | 12 |
| Security | 4.38 | 4.14 | 12 |
| Maintainability | 4.19 | 3.64 | 12 |
| Simplicity | 4.19 | 3.85 | 12 |
| Conventions | 4.36 | 3.52 | 12 |
| Docs | 4.04 | 3.2 | 12 |
| Communication | 4.59 | 3.58 | 12 |
| Production readiness | 4.03 | 3.21 | 12 |
| Overall | 4.14 | 3.2 | 12 |
14 rows, 2 series: AI change, Merged human change. AI change: highest Communication 4.59 (n 12). Lowest Edge cases 3.95 (n 12). Merged human change: highest Security 4.14 (n 12). Lowest Overall 3.2 (n 12).
Notesn = 12 per row
Mean 1-5 score per dimension, latest attempt of 12 tasks
Unweighted mean over tasks of each pair’s mean critic score. Critics were not told which change came from a person.
Source: Blind review panel: AI worker change vs merged human change
Every interval overlaps every other: this chart does not order these rows.
| Item | Critic model | 95% interval | n |
|---|---|---|---|
| GPT 5.5 (OpenAI) | 100% | 34%–100% | 2 |
| Codex (GPT) (OpenAI) | 80% | 49%–94% | 10 |
| Claude Opus 5 (Anthropic) | 70% | 55%–82% | 40 |
| Claude Sonnet (Anthropic) | 68% | 52%–80% | 40 |
| Claude Fable 5 (Anthropic) | 65% | 50%–78% | 40 |
5 rows. Highest GPT 5.5 (OpenAI) 100% (95% interval 34%–100%, n 2). Lowest Claude Fable 5 (Anthropic) 65% (95% interval 50%–78%, n 40). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 2–40 per row
Share of single verdicts preferring the AI change, per critic model, all scored pairs
Whiskers are 95% Wilson intervals over single verdicts (verdicts within one pair are not independent). The worker model is Anthropic; OpenAI critics joined later and judged fewer pairs.
Source: Blind review panel: AI worker change vs merged human change
Tables
Every scored pair
| Task | Language | Kind | Attempt | Panel decision | Votes AI-human | Critics | Operator answers | Run cost (notional) |
|---|---|---|---|---|---|---|---|---|
| aio-libs/aiohttp#13122 | Python | feature | 3 | Human preferred | 4-4 | Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet | 5 | $78.28 |
| redis/go-redis#3914 | Go | ambiguous | 1 | Human preferred | 1-5 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 0 | $21.05 |
| redis/go-redis#3914 | Go | ambiguous | 2 | AI preferred | 8-0 | Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet | 1 | $31.50 |
| h3js/h3#1533 | TypeScript | refactor | 2 | AI preferred | 8-0 | Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet | 0 | $21.77 |
| open-telemetry/opentelemetry-go#8706 | Go | bug | 1 | Human preferred | 3-3 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 2 | $35.17 |
| open-telemetry/opentelemetry-go#8706 | Go | bug | 2 | AI preferred | 8-0 | Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet | 3 | $80.48 |
| private-go-a (private Go) | Go | bug | 1 | AI preferred | 6-0 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 1 | $72.29 |
| private-go-a (private Go) | Go | bug | 3 | AI preferred | 6-0 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 3 | $54.51 |
| private-go-a (private Go) | Go | bug | 7 | AI preferred | 8-0 | Claude Opus 5, Claude Fable 5, GPT 5.5, Claude Sonnet | 0 | $47.65 |
| private-go-b (private Go) | Go | bug | 1 | Human preferred | 1-5 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 0 | $74.97 |
| private-go-b (private Go) | Go | bug | 2 | Human preferred | 0-6 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 2 | $74.76 |
| private-go-b (private Go) | Go | bug | 3 | Human preferred | 0-6 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 4 | $70.28 |
| private-go-c (private Go) | Go | bug | 1 | Human preferred | 0-6 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 3 | $73.13 |
| private-go-d (private Go) | Go | refactor | 1 | AI preferred | 6-0 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 3 | $43.39 |
| private-go-e (private Go) | Go | feature | 1 | Human preferred | 0-6 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 2 | $22.93 |
| private-go-e (private Go) | Go | feature | 2 | AI preferred | 6-0 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 1 | $32.75 |
| private-go-e (private Go) | Go | feature | 3 | AI preferred | 6-0 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 3 | $26.81 |
| private-go-f (private Go) | Go | feature | 1 | AI preferred | 6-0 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 2 | $53.93 |
| private-go-g (private Go) | Go | feature | 2 | AI preferred | 6-0 | Claude Opus 5, Claude Fable 5, Claude Sonnet | 1 | $41.35 |
| fastify/session#348 | JavaScript | security | 1 | AI preferred | 8-0 | Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet | 0 | $21.53 |
Method
- Each task is a real merged pull request: the issue as the ticket, the repository at the base commit, the merged change as the human reference.
- The Agent worker solves the ticket in a sandbox without seeing the reference. Its pull request and the merged one form a pair.
- Critic models (Claude Opus 5, Claude Fable 5, Claude Sonnet, and on some pairs Codex or GPT 5.5) review both changes without labels, once in each order.
- Decision rule 2: the AI change wins on a strict majority of verdicts and no panel veto (serious issues flagged by at least two critics).
- Every scored attempt is kept. A re-scoring of the same pair under the current rule replaces the earlier scoring.
Caveats
- n = 12 tasks: intervals are wide.
- Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
- Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.
- 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.
- 4 pairs were flagged for position bias (a critic flipped when the order swapped).
- The human change was merged, reviewed and shipped. A critic preference says nothing about long-term maintenance cost.
Sources
Blind review panel: AI worker change vs merged human change
Each pair is judged by 3 or 4 critic models without labels, in both orders. 5 public OSS tasks and 7 private tasks. Private task names are replaced by neutral labels.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “AI pull requests vs merged human pull requests, judged blind”, updated October 5, 2026, https://agent.sasid.ai/benchmarks/blind-review-head-to-head.
Explainers that cite this study
Read the methods and terms in the context of these recorded results.
Models and comparisons in this study
Write-ups on this study
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
Do blind AI critics prefer AI pull requests over human ones?
Blind Claude and GPT critics preferred an AI worker's change over the merged human change on 9 of 12 real tasks, and 6 of 12 on the first try. Caveats inside.
Harness vs model: where do AI coding agent gains really come from?
Is it the model or the harness? Our data shows the harness clearly moves speed, tokens and cost. Whether it moves accuracy, our samples cannot yet say.
Why we count every failed attempt: our rules for honest AI benchmarks
How we benchmark AI agents: declare the protocol first, keep every failure, show intervals, publish our own confounds and label interim results.
More studies
All benchmarksDoes a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.
Does a new Claude Code session reuse the prompt cache of an earlier one?
30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.