Three-way gates
Base, candidate, and reference run the same explicit checks so failures are visible in context.
Your codebase. Your proof.
Same starting point. A blind implementation. An answer you can inspect.
Give Agent the original task. Keep your shipped fix hidden. Watch the work, compare the evidence, and decide what deserves your next task.
Sign in to run yoursRoute middleware is skipped when ~getMiddleware is overridden. Restore it without duplicating middleware already included by the override.
The everyday workflow
A message becomes a plan, implementation updates and a pull request you can review. The result returns to the conversation where the task began.
The worker turns the task into a plan and asks about open questions.
Implementation updates, changes and checks appear together.
Follow approval and completion back to the original conversation.
Scripted example with a fictional worker and PR. Connected messaging routes depend on setup.
You · WhatsApp
The payments list hides the drafts from the new invoice import. Can Ada take NW-142?The protocol
The point is a trustworthy customer conversation: what was asked, what ran, what passed, how much human effort it took, and what a person decided.
Choose a merged GitHub pull request and describe the requirement that existed before its solution. The candidate starts from the requirement, never from the reference patch.
Pick the local runner, CLI profile, and model explicitly. Reproducible checks run as argv inside that runner’s sandbox.
See the candidate and reference results, integrity hash, three-way checks, runner events, time, and cost provenance in one record.
Record your verdict and total human effort. Missing evidence stays inconclusive until a person has inspected the result.
What you can inspect
Reported baselines, measured elapsed time, estimated CLI cost, verified charges, and unknown values stay distinct. A comparison signal never replaces the customer verdict.
Base, candidate, and reference run the same explicit checks so failures are visible in context.
Follow the local task status and event history while the work is queued, running, and verified.
Record the verdict, rationale, and review minutes after inspecting what the trial actually returned.
Start with work you already understand. See what Agent can deliver, inspect the tradeoffs, and carry your next real task into the command center.
Eligible merged GitHub.com PRs · Agent runner 0.5+ on macOS · self-contained offline checks. Model cost is estimated; vendor charges stay separate.