PR trials
A PR trial replays a pull request you already merged. Agent builds it again blind on your computer, then compares the result, checks, time and estimated cost with the original.
What a trial doesLink to What a trial does
- You pick a merged GitHub.com pull request and give the original requirement.
- The runner on your computer rebuilds the repository's history up to before the change. The model does not see the merged change.
- Your own Claude Code or Codex builds a candidate within the turn, time and cost caps you set.
- When the candidate is frozen, the checks run without a model on separate copies of the base, the candidate and the reference.
- You compare the two diffs and record a verdict: Candidate preferred, Reference preferred, Comparable or Inconclusive.
What you needLink to What you need
- Agent runner 0.5 or newer on macOS, with a working sandbox;
- a merged pull request on GitHub.com;
- checks that run offline with the tools you already have.
StatusesLink to Statuses
Ready to run, Queued, Running, Verifying, Completed, Failed and Cancelled.
CostLink to Cost
The trial uses your existing CLI login on your computer. Agent does not give model credits for it. Costs shown are estimates, not bills.
Checked against the product on 2026-10-05.
Was this helpful?
Related articles
- The local runnerThe runner is a small program on your own computer. It starts your own Claude Code or Codex, under your own sign-in, for work that you asked for.
- The runner's sandboxOn macOS, work on your runner runs in a sandbox. It can write only its own task folder and reach only the model vendor and allowed package registries.
- Evidence and checksWorkers attach proof to their work: tests and results, live calls, screenshots and waivers. Agent never shows a waived check as passed.