• Calibration
  • Coding Agents
  • Failures
  • Fastify
  • H3
  • Uvicorn

Coding calibration: what broke on three real pull requests

On three real upstream issues, does the AI worker deliver a verified change, and what stops it when it does not?

Published · Updated · 4 charts · Download the data or a carousel

1 of 3

n = 3

Verified deliveries, latest build

The answer

Across four platform builds, verified delivery went from 0 of 3 to 1 of 3; the latest build delivered fastify/session cleanly, but h3 regressed to an empty patch when a formatting failure was lost across a context fold, and uvicorn passed its full suite yet stayed unverified for two gate-classification reasons. The latest slice cost $11.06 for the three tasks (notional). The first calibration run on 2026-10-02 produced a passing patch for fastify/session but stopped at its $5 cap before delivery. One attempt per cell finds defects; it does not measure a rate.

Key numbers

$11.06

Notional cost, latest build, all 3 tasks

n = 3

26 → 19

Guardrail refusals, first vs latest slice

n = 3

$4.89, stopped at cap

First calibration run (capped, fastify/session)

n = 1

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

  • Functional pass (offline gates)
  • Verified delivery
  • Pull request opened
Baseline (capped)
Fix wave 1 (capped)
Uncapped, build f0ac3a8a
Uncapped, build 236c0d3f

One square per item of n; 3 strips per row, one per class.

4 rows, 3 series: Functional pass (offline gates), Verified delivery, Pull request opened. Functional pass (offline gates): highest Baseline (capped) 2 (n 3). Lowest Uncapped, build 236c0d3f 1 (n 3). Verified delivery: highest Uncapped, build 236c0d3f 1 (n 3). Lowest Uncapped, build f0ac3a8a 0 (n 3).

Notesn = 3 per row

Tasks per slice: functional pass, verified delivery, pull request opened

One attempt per task per slice. The two capped slices stopped at 20 minutes or $5; the uncapped slices had no ceiling. Each slice is a different platform build, so a change is not a matched improvement.

Source: Coding calibration: fastify/session, h3, uvicorn

Share card (PNG)

Baseline (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Fix wave 1 (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build f0ac3a8a

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build 236c0d3f

fastify/session
h3js/h3
Kludex/uvicorn

One panel per series, all on the same axis.

3 rows, 4 series: Baseline (capped), Fix wave 1 (capped), Uncapped, build f0ac3a8a, Uncapped, build 236c0d3f. Baseline (capped): highest Kludex/uvicorn $4.88. Lowest fastify/session $3.43. Fix wave 1 (capped): highest fastify/session $4.85. Lowest h3js/h3 $3.29.

Notes

List-price estimates of subscription calls. Failed and capped attempts count.

Source: Coding calibration: fastify/session, h3, uvicorn

Share card (PNG)

Baseline (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Fix wave 1 (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build f0ac3a8a

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build 236c0d3f

fastify/session
h3js/h3
Kludex/uvicorn

One panel per series, all on the same axis.

3 rows, 4 series: Baseline (capped), Fix wave 1 (capped), Uncapped, build f0ac3a8a, Uncapped, build 236c0d3f. Baseline (capped): slowest Kludex/uvicorn 17.9 min. Fastest fastify/session 9 min. Fix wave 1 (capped): slowest Kludex/uvicorn 20.3 min. Fastest h3js/h3 13.5 min.

Notes

Onboarding included. uvicorn needed more than the old 20 minute cap once the caps were removed.

Source: Coding calibration: fastify/session, h3, uvicorn

Share card (PNG)
Baseline (capped)
Fix wave 1 (capped)
Uncapped, build f0ac3a8a
Uncapped, build 236c0d3f

Hover or focus a bar for its ratio to Uncapped, build 236c0d3f (the highlighted row): a ratio of the two values shown, not a measurement.

4 rows. Highest Fix wave 1 (capped) 28. Lowest Uncapped, build 236c0d3f 19.

Notes

Tool calls the platform turned back (plan schema, delivery contract, outgoing text, loops)

A refusal is a guardrail working, not always a failure: some catch real problems, some were platform defects that the next build fixed.

Source: Coding calibration: fastify/session, h3, uvicorn

Share card (PNG)

Tables

Every attempt, failures included

TaskSliceFunctional gatesVerified deliveryHow it endedMinutesCost (notional)CallsWhy not verified
fastify/sessionBaseline (capped)verifiednoEscalated to a person9 min$3.4343run-not-completed
fastify/sessionFix wave 1 (capped)verifiednoHit the $5 cap15.5 min$4.8573integrity-caveat:cost-capped, run-not-completed
fastify/sessionUncapped, build f0ac3a8afailednoEscalated to a person7.2 min$2.4429account-route-mismatch, run-not-completed, acceptance-failed, cross-test-command-failed:reference:default, cross-test-failed:reference:default, empty-patch
fastify/sessionUncapped, build 236c0d3fverifiedyesDelivered (in review)11.2 min$3.5352—
h3js/h3Baseline (capped)verifiednoEscalated to a person11.8 min$4.0053run-not-completed
h3js/h3Fix wave 1 (capped)verifiednoEscalated to a person13.5 min$3.2946run-not-completed
h3js/h3Uncapped, build f0ac3a8averifiednoDelivered (in review)12.5 min$3.8148account-route-mismatch
h3js/h3Uncapped, build 236c0d3ffailednoEscalated to a person9.2 min$2.9635run-not-completed, acceptance-failed, cross-test-command-failed:reference:default, cross-test-failed:reference:default, empty-patch
Kludex/uvicornBaseline (capped)unverifiednoHit the $5 cap17.9 min$4.8867integrity-caveat:cost-capped, run-not-completed, acceptance-missing, base-acceptance-not-run, base-passToPass-not-passed, candidate-acceptance-not-run, cross-tests-not-run:candidate:locked, cross-tests-not-run:candidate:ws17, cross-tests-not-run:reference:locked, cross-tests-not-run:reference:ws17, step-fail:install, step-skipped:build, step-skipped:lint, step-skipped:test, step-skipped:typecheck
Kludex/uvicornFix wave 1 (capped)unverifiednoHit the 20 min cap20.3 min$4.6368run-not-completed, acceptance-missing, base-acceptance-not-run, base-passToPass-not-passed, candidate-acceptance-not-run, cross-tests-not-run:candidate:locked, cross-tests-not-run:candidate:ws17, cross-tests-not-run:reference:locked, cross-tests-not-run:reference:ws17, step-fail:install, step-skipped:build, step-skipped:lint, step-skipped:test, step-skipped:typecheck
Kludex/uvicornUncapped, build f0ac3a8aunverifiednoDelivered (in review)22.7 min$4.7473account-route-mismatch, acceptance-missing, base-acceptance-not-run, base-passToPass-not-passed, candidate-acceptance-not-run, cross-tests-not-run:candidate:ws17, cross-tests-not-run:reference:ws17
Kludex/uvicornUncapped, build 236c0d3funverifiednoDelivered (in review)19.8 min$4.5762base-did-not-fail, cross-test-command-failed:candidate:ws17

Method

  1. Tasks: fastify/session #348, h3js/h3 #1533 and Kludex/uvicorn #3036, each frozen at the base commit with the merged change as reference.
  2. Fixed model claude-sonnet-5-5, routing off, balanced mode, real cold onboarding, no review replay, no operator answers.
  3. Worker and gates run offline with frozen dependencies. Gates run install, lint, typecheck, build and the full suite at base, candidate and reference, plus cross-tests.
  4. "Functional pass" means the frozen patch passed the gates. "Verified delivery" also needs a completed, clean, unassisted run.
  5. One attempt per task per slice. Every stop and failure is kept; nothing is rerun or replaced.

Caveats

  • One attempt per cell: these are defect-finding runs, not rates.
  • Each slice changes the platform build, and the last two also remove the caps. No two slices are a matched comparison.
  • Costs are list-price estimates for subscription calls, not invoices.
  • The f0ac3a8a h3 run was verified only after its account-route label was corrected; it counts as unverified here, as declared.

Sources

  • Coding calibration: fastify/session, h3, uvicorn

    Our recorded runs ·

    Three real upstream tasks, one attempt per task per platform slice, offline gates against the merged reference.

    Raw data: calibration/slices.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Coding calibration: what broke on three real pull requests”, updated October 5, 2026, https://agent.sasid.ai/benchmarks/coding-calibration.

Models and comparisons in this study

More studies

All benchmarks
Live story
  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.