• Code Review
  • AI vs human
  • LLM as judge
  • Pull Requests

AI pull requests vs merged human pull requests, judged blind

When critics cannot see which change came from a person, do they prefer the AI worker’s pull request or the one the maintainers merged?

Published · Updated · 4 charts · Download the data or a carousel

75%

95% CI 47%–91% · n = 12

9/12 · Tasks where the panel preferred the AI change (latest attempt)

Latest attempts come after earlier attempts on the same task; see caveats.

The answer

On the latest attempt per task, the blind panel preferred the AI change on 9 of 12 tasks (75%, 95% interval 47% to 91%). On the first scored attempt it was 6 of 12 (25% to 75%). Later attempts learned from earlier ones, so the first-attempt figure is the cleaner estimate. Public OSS tasks: 4 of 5; private Go tasks: 5 of 7. Across 132 single verdicts, 69% preferred the AI change. A critic panel's preference is not a merge decision and not a correctness proof.

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 34 sAI pull requests vs merged human pull requests, judged blind

AI pull requests vs merged human pull requests, judged blind

Blind critics preferred the AI change on 9 of 12 tasks at the latest attempt and 6 of 12 at the first. Same-family bias disclosed.

Transcript
  1. Blind review · 12 real merged pull requests. AI change vs the merged human change. Critic models review both, unlabelled, once in each order.
  2. Latest attempt: the panel preferred the AI change on 9 of 12 tasks. First attempt: 6 of 12. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.
  3. The first attempt is the cleaner estimate: 50%, with a 95% interval from 25% to 75%. Chart: First attempt vs latest attempt (n = 5–12 each). Caveat: n = 12 tasks: intervals are wide.
  4. Vote by vote: 9 of 12 tasks went unanimously to the AI change, 2 unanimously to the human one. Chart: Blind panel votes per task (latest attempt) (n = 6–8 each). Caveat: 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.
  5. Cross-family check: OpenAI critics preferred the AI change in 10 of 12 verdicts. Small n, wide intervals. Chart: Does the judge’s model family matter? (n = 2–40 each). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
  6. Open benchmarks: intervals, sources and every failure kept.

Key numbers

50% (6/12)

Tasks where the panel preferred the AI change (first scored attempt)

95% CI 25%–75% · n = 12

60% (12/20)

All scored pairs where the panel preferred the AI change

95% CI 39%–78% · n = 20

80% (4/5)

Public OSS tasks, latest attempt

95% CI 38%–96% · n = 5

69% (91/132)

Single critic verdicts that preferred the AI change

95% CI 61%–76% · n = 132

$979 + $332

Notional spend: worker runs + critic panel

n = 20

4 of 20

Pairs flagged for position bias

n = 20

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

First scored attempt
Latest attempt
Public OSS tasks, latest
Private tasks, latest

Every interval overlaps every other: this chart does not order these rows.

4 rows. Highest Public OSS tasks, latest 80% (95% interval 38%–96%, n 5). Lowest First scored attempt 50% (95% interval 25%–75%, n 12). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 5–12 per row

Share of tasks where the blind panel preferred the AI change

Later attempts had the earlier attempts’ lessons, review replays and, on some tasks, operator answers. They are not independent first tries.

Source: Blind review panel: AI worker change vs merged human change

Share card (PNG)
  • Prefers AI change
  • Prefers human change
  • Tie
h3js/h3#1533
private-go-a (private Go)
open-telemetry/opentelemetry-go#8706
fastify/session#348
private-go-d (private Go)
redis/go-redis#3914
private-go-g (private Go)
private-go-f (private Go)
private-go-e (private Go)
aio-libs/aiohttp#13122
private-go-c (private Go)
private-go-b (private Go)

One square per verdict; counts at the right are exact and in legend order.

12 rows, 3 series: Prefers AI change, Prefers human change, Tie. Prefers AI change: highest h3js/h3#1533 8 (n 8). Lowest private-go-b (private Go) 0 (n 6). Prefers human change: highest private-go-c (private Go) 6 (n 6). Lowest private-go-e (private Go) 0 (n 6).

Notesn 6–8 per row

Each critic judges both orders without labels

The human change is the one the maintainers merged upstream. Private tasks are from one private Go service and carry neutral labels.

Source: Blind review panel: AI worker change vs merged human change

Share card (PNG)
  • AI change
  • Merged human change (square)
Sorted by gap, largest first.
Tests
Communication
Overall
Completeness
Conventions
Docs
Production readiness
Compatibility
Maintainability
Correctness
Root cause
Edge cases
Simplicity
Security

Gap labels, Merged human change vs AI change: Merged human change is x points higher (+) or lower (−) than AI change, calculated from the two values shown.

14 rows, 2 series: AI change, Merged human change. AI change: highest Communication 4.59 (n 12). Lowest Edge cases 3.95 (n 12). Merged human change: highest Security 4.14 (n 12). Lowest Overall 3.2 (n 12).

Notesn = 12 per row

Mean 1-5 score per dimension, latest attempt of 12 tasks

Unweighted mean over tasks of each pair’s mean critic score. Critics were not told which change came from a person.

Source: Blind review panel: AI worker change vs merged human change

Share card (PNG)
GPT 5.5 (OpenAI)
Codex (GPT) (OpenAI)
Claude Opus 5 (Anthropic)
Claude Sonnet (Anthropic)
Claude Fable 5 (Anthropic)

Every interval overlaps every other: this chart does not order these rows.

5 rows. Highest GPT 5.5 (OpenAI) 100% (95% interval 34%–100%, n 2). Lowest Claude Fable 5 (Anthropic) 65% (95% interval 50%–78%, n 40). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 2–40 per row

Share of single verdicts preferring the AI change, per critic model, all scored pairs

Whiskers are 95% Wilson intervals over single verdicts (verdicts within one pair are not independent). The worker model is Anthropic; OpenAI critics joined later and judged fewer pairs.

Source: Blind review panel: AI worker change vs merged human change

Share card (PNG)

Tables

Every scored pair

TaskLanguageKindAttemptPanel decisionVotes AI-humanCriticsOperator answersRun cost (notional)
aio-libs/aiohttp#13122Pythonfeature3Human preferred4-4Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet5$78.28
redis/go-redis#3914Goambiguous1Human preferred1-5Claude Opus 5, Claude Fable 5, Claude Sonnet0$21.05
redis/go-redis#3914Goambiguous2AI preferred8-0Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet1$31.50
h3js/h3#1533TypeScriptrefactor2AI preferred8-0Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet0$21.77
open-telemetry/opentelemetry-go#8706Gobug1Human preferred3-3Claude Opus 5, Claude Fable 5, Claude Sonnet2$35.17
open-telemetry/opentelemetry-go#8706Gobug2AI preferred8-0Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet3$80.48
private-go-a (private Go)Gobug1AI preferred6-0Claude Opus 5, Claude Fable 5, Claude Sonnet1$72.29
private-go-a (private Go)Gobug3AI preferred6-0Claude Opus 5, Claude Fable 5, Claude Sonnet3$54.51
private-go-a (private Go)Gobug7AI preferred8-0Claude Opus 5, Claude Fable 5, GPT 5.5, Claude Sonnet0$47.65
private-go-b (private Go)Gobug1Human preferred1-5Claude Opus 5, Claude Fable 5, Claude Sonnet0$74.97
private-go-b (private Go)Gobug2Human preferred0-6Claude Opus 5, Claude Fable 5, Claude Sonnet2$74.76
private-go-b (private Go)Gobug3Human preferred0-6Claude Opus 5, Claude Fable 5, Claude Sonnet4$70.28
private-go-c (private Go)Gobug1Human preferred0-6Claude Opus 5, Claude Fable 5, Claude Sonnet3$73.13
private-go-d (private Go)Gorefactor1AI preferred6-0Claude Opus 5, Claude Fable 5, Claude Sonnet3$43.39
private-go-e (private Go)Gofeature1Human preferred0-6Claude Opus 5, Claude Fable 5, Claude Sonnet2$22.93
private-go-e (private Go)Gofeature2AI preferred6-0Claude Opus 5, Claude Fable 5, Claude Sonnet1$32.75
private-go-e (private Go)Gofeature3AI preferred6-0Claude Opus 5, Claude Fable 5, Claude Sonnet3$26.81
private-go-f (private Go)Gofeature1AI preferred6-0Claude Opus 5, Claude Fable 5, Claude Sonnet2$53.93
private-go-g (private Go)Gofeature2AI preferred6-0Claude Opus 5, Claude Fable 5, Claude Sonnet1$41.35
fastify/session#348JavaScriptsecurity1AI preferred8-0Claude Opus 5, Codex (GPT), Claude Fable 5, Claude Sonnet0$21.53

Method

  1. Each task is a real merged pull request: the issue as the ticket, the repository at the base commit, the merged change as the human reference.
  2. The Agent worker solves the ticket in a sandbox without seeing the reference. Its pull request and the merged one form a pair.
  3. Critic models (Claude Opus 5, Claude Fable 5, Claude Sonnet, and on some pairs Codex or GPT 5.5) review both changes without labels, once in each order.
  4. Decision rule 2: the AI change wins on a strict majority of verdicts and no panel veto (serious issues flagged by at least two critics).
  5. Every scored attempt is kept. A re-scoring of the same pair under the current rule replaces the earlier scoring.

Caveats

  • n = 12 tasks: intervals are wide.
  • Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
  • Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.
  • 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.
  • 4 pairs were flagged for position bias (a critic flipped when the order swapped).
  • The human change was merged, reviewed and shipped. A critic preference says nothing about long-term maintenance cost.

Sources

  • Blind review panel: AI worker change vs merged human change

    Our recorded runs ·

    Each pair is judged by 3 or 4 critic models without labels, in both orders. 5 public OSS tasks and 7 private tasks. Private task names are replaced by neutral labels.

    Raw data: blind-review/attempts.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “AI pull requests vs merged human pull requests, judged blind”, updated October 5, 2026, https://agent.sasid.ai/benchmarks/blind-review-head-to-head.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

Models and comparisons in this study

More studies

All benchmarks
  • Prompt Caching
  • Cache Reuse

Does a new Claude Code session reuse the prompt cache of an earlier one?

30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.

0of 2 (95% interval 0% to 66%) · Later sessions with at least 50% of turn-1 input cached, A: new folder each time · n = 2

4 chartsUpdated October 7, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.