{"i":2,"study":{"slug":"blind-review-head-to-head","title":"AI pull requests vs merged human pull requests, judged blind","seoTitle":"AI vs human pull requests: a blind multi-model review","description":"A blind panel of Claude and GPT critics preferred Agent's change over the merged human change on 9 of 12 real tasks. Votes, scores, caveats.","question":"When critics cannot see which change came from a person, do they prefer the AI worker’s pull request or the one the maintainers merged?","answer":"On the latest attempt per task, the blind panel preferred the AI change on 9 of 12 tasks (75%, 95% interval 47% to 91%). On the first scored attempt it was 6 of 12 (25% to 75%). Later attempts learned from earlier ones, so the first-attempt figure is the cleaner estimate. Public OSS tasks: 4 of 5; private Go tasks: 5 of 7. Across 132 single verdicts, 69% preferred the AI change. A critic panel's preference is not a merge decision and not a correctness proof.","date":"2026-09-28","updated":"2026-10-05","tags":["code-review","ai-vs-human","llm-as-judge","pull-requests"],"caveats":["n = 12 tasks: intervals are wide.","Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.","Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.","7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.","4 pairs were flagged for position bias (a critic flipped when the order swapped).","The human change was merged, reviewed and shipped. A critic preference says nothing about long-term maintenance cost."],"sourceIds":["agent-blind-review"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["ai-preferred-latest","Tasks where the panel preferred the AI change (latest attempt)",0.75,"rate","75% (9/12)",12,[0.4677,0.9111],"Latest attempts come after earlier attempts on the same task; see caveats."],["ai-preferred-first","Tasks where the panel preferred the AI change (first scored attempt)",0.5,"rate","50% (6/12)",12,[0.2538,0.7462],"\u0001"],["ai-preferred-all-pairs","All scored pairs where the panel preferred the AI change",0.6,"rate","60% (12/20)",20,[0.3866,0.7812],"\u0001"],["ai-preferred-public","Public OSS tasks, latest attempt",0.8,"rate","80% (4/5)",5,[0.3755,0.9638],"\u0001"],["verdicts-ai","Single critic verdicts that preferred the AI change",0.6894,"rate","69% (91/132)",132,[0.606,0.762],"\u0001"],["review-spend","Notional spend: worker runs + critic panel",1310.5,"usd","$979 + $332",20,"\u0001","List-price estimates of subscription calls over all scored pairs."],["position-bias-flags","Pairs flagged for position bias",4,"count","4 of 20",20,"\u0001","A critic model flipped its preference when the two changes swapped places."]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["blind-review-first-vs-latest","First attempt vs latest attempt","Share of tasks where the blind panel preferred the AI change","dot-range","rate","Tasks preferred",[{"name":"AI preferred","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["First scored attempt",0.5,0.2538,0.7462,12,"\u0001"],["Latest attempt",0.75,0.4677,0.9111,12,true],["Public OSS tasks, latest",0.8,0.3755,0.9638,5,"\u0001"],["Private tasks, latest",0.7143,0.3589,0.9178,7,"\u0001"]]}}],"Later attempts had the earlier attempts’ lessons, review replays and, on some tasks, operator answers. They are not independent first tries.",["agent-blind-review"]],["blind-review-votes-by-task","Blind panel votes per task (latest attempt)","Each critic judges both orders without labels","stacked-bar","count","Verdicts",{"$k":["name","points"],"$r":[["Prefers AI change",{"$k":["label","value","n","highlight"],"$r":[["h3js/h3#1533",8,8,true],["private-go-a (private Go)",8,8,true],["open-telemetry/opentelemetry-go#8706",8,8,true],["fastify/session#348",8,8,true],["private-go-d (private Go)",6,6,true],["redis/go-redis#3914",8,8,true],["private-go-g (private Go)",6,6,true],["private-go-f (private Go)",6,6,true],["private-go-e (private Go)",6,6,true],["aio-libs/aiohttp#13122",4,8,false],["private-go-c (private Go)",0,6,false],["private-go-b (private Go)",0,6,false]]}],["Prefers human change",{"$k":["label","value","n"],"$r":[["h3js/h3#1533",0,8],["private-go-a (private Go)",0,8],["open-telemetry/opentelemetry-go#8706",0,8],["fastify/session#348",0,8],["private-go-d (private Go)",0,6],["redis/go-redis#3914",0,8],["private-go-g (private Go)",0,6],["private-go-f (private Go)",0,6],["private-go-e (private Go)",0,6],["aio-libs/aiohttp#13122",4,8],["private-go-c (private Go)",6,6],["private-go-b (private Go)",6,6]]}],["Tie",{"$k":["label","value","n"],"$r":[["h3js/h3#1533",0,8],["private-go-a (private Go)",0,8],["open-telemetry/opentelemetry-go#8706",0,8],["fastify/session#348",0,8],["private-go-d (private Go)",0,6],["redis/go-redis#3914",0,8],["private-go-g (private Go)",0,6],["private-go-f (private Go)",0,6],["private-go-e (private Go)",0,6],["aio-libs/aiohttp#13122",0,8],["private-go-c (private Go)",0,6],["private-go-b (private Go)",0,6]]}]]},"The human change is the one the maintainers merged upstream. Private tasks are from one private Go service and carry neutral labels.",["agent-blind-review"]],["blind-review-dimension-scores","What the critics scored higher","Mean 1-5 score per dimension, latest attempt of 12 tasks","grouped-bar","score","Mean score (1-5)",[{"name":"AI change","points":{"$k":["label","value","n","highlight"],"$r":[["Correctness",4.4,12,true],["Root cause",4.46,12,true],["Edge cases",3.95,12,true],["Tests",4.51,12,true],["Completeness",4.22,12,true],["Compatibility",4.39,12,true],["Security",4.38,12,true],["Maintainability",4.19,12,true],["Simplicity",4.19,12,true],["Conventions",4.36,12,true],["Docs",4.04,12,true],["Communication",4.59,12,true],["Production readiness",4.03,12,true],["Overall",4.14,12,true]]}},{"name":"Merged human change","points":{"$k":["label","value","n"],"$r":[["Correctness",3.86,12],["Root cause",3.97,12],["Edge cases",3.49,12],["Tests",3.29,12],["Completeness",3.34,12],["Compatibility",3.74,12],["Security",4.14,12],["Maintainability",3.64,12],["Simplicity",3.85,12],["Conventions",3.52,12],["Docs",3.2,12],["Communication",3.58,12],["Production readiness",3.21,12],["Overall",3.2,12]]}}],"Unweighted mean over tasks of each pair’s mean critic score. Critics were not told which change came from a person.",["agent-blind-review"]],["blind-review-critic-agreement","Does the judge’s model family matter?","Share of single verdicts preferring the AI change, per critic model, all scored pairs","dot-range","rate","Prefers AI change",[{"name":"Critic model","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT 5.5 (OpenAI)",1,0.3424,1,2],["Codex (GPT) (OpenAI)",0.8,0.4902,0.9433,10],["Claude Opus 5 (Anthropic)",0.7,0.5457,0.8193,40],["Claude Sonnet (Anthropic)",0.675,0.5202,0.7992,40],["Claude Fable 5 (Anthropic)",0.65,0.4951,0.7787,40]]}}],"Whiskers are 95% Wilson intervals over single verdicts (verdicts within one pair are not independent). The worker model is Anthropic; OpenAI critics joined later and judged fewer pairs.",["agent-blind-review"]]]}}}