Your codebase. Your proof.

Test Agent on a PR you already shipped.

Same starting point. A blind implementation. An answer you can inspect.

Give Agent the original task. Keep your shipped fix hidden. Watch the work, compare the evidence, and decide what deserves your next task.

Prepare: Start from the requirement and the original code. Hold the shipped patch back.
Inside a PR trialRecorded benchmark
Start from the requirement and the original code. Hold the shipped patch back.
The original task

Keep route middleware running.

Route middleware is skipped when ~getMiddleware is overridden. Restore it without duplicating middleware already included by the override.

  • Keep route middleware in the request flow.
  • Handle overrides that already include it.
  • Preserve the original fast dispatch path.
h3js/h3Historical base 4a9617e
Original requirementWhat needed to change
Fresh implementationWork from the base
Reference held backCompare after the run
What this recorded run produced
2,775sealed full-suite passed · 47 skipped
Run time14m 04sMeasured
Changed files4Recorded diff
h3js/h3 · recorded 23 Sep 2026
In your trial, the reference stays sealed until the candidate is locked.Your own run uses your runner and CLI login.
Blind from the startOriginal requirement. Original base.
Every result inspectableCode, checks, activity and cost provenance.
Your judgment decidesCompare both. Record the rationale.

The everyday workflow

Send the task.
Follow the work.

A message becomes a plan, implementation updates and a pull request you can review. The result returns to the conversation where the task began.

  1. 01
    Send a clear request

    The worker turns the task into a plan and asks about open questions.

  2. 02
    See the work take shape

    Implementation updates, changes and checks appear together.

  3. 03
    Review the pull request

    Follow approval and completion back to the original conversation.

Scripted example with a fictional worker and PR. Connected messaging routes depend on setup.

Example · scripted

You · WhatsApp

The payments list hides the drafts from the new invoice import. Can Ada take NW-142?

The protocol

A small experiment with a clear record.

The point is a trustworthy customer conversation: what was asked, what ran, what passed, how much human effort it took, and what a person decided.

  1. 01

    Bring one real PR

    Choose a merged GitHub pull request and describe the requirement that existed before its solution. The candidate starts from the requirement, never from the reference patch.

  2. 02

    Run it on your runner

    Pick the local runner, CLI profile, and model explicitly. Reproducible checks run as argv inside that runner’s sandbox.

  3. 03

    Inspect the evidence

    See the candidate and reference results, integrity hash, three-way checks, runner events, time, and cost provenance in one record.

  4. 04

    Make the customer call

    Record your verdict and total human effort. Missing evidence stays inconclusive until a person has inspected the result.

What you can inspect

Numbers keep their source.

Reported baselines, measured elapsed time, estimated CLI cost, verified charges, and unknown values stay distinct. A comparison signal never replaces the customer verdict.

Three-way gates

Base, candidate, and reference run the same explicit checks so failures are visible in context.

Runner events

Follow the local task status and event history while the work is queued, running, and verified.

Human review

Record the verdict, rationale, and review minutes after inspecting what the trial actually returned.

Bring the next real task.

Start with work you already understand. See what Agent can deliver, inspect the tradeoffs, and carry your next real task into the command center.

Eligible merged GitHub.com PRs · Agent runner 0.5+ on macOS · self-contained offline checks. Model cost is estimated; vendor charges stay separate.