{"i":15,"study":{"slug":"coding-calibration","title":"Coding calibration: what broke on three real pull requests","seoTitle":"AI coding agent calibration on fastify, h3 and uvicorn tasks","description":"One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.","question":"On three real upstream issues, does the AI worker deliver a verified change, and what stops it when it does not?","answer":"Across four platform builds, verified delivery went from 0 of 3 to 1 of 3; the latest build delivered fastify/session cleanly, but h3 regressed to an empty patch when a formatting failure was lost across a context fold, and uvicorn passed its full suite yet stayed unverified for two gate-classification reasons. The latest slice cost $11.06 for the three tasks (notional). The first calibration run on 2026-10-02 produced a passing patch for fastify/session but stopped at its $5 cap before delivery. One attempt per cell finds defects; it does not measure a rate.","date":"2026-10-02","updated":"2026-10-05","tags":["calibration","coding-agents","failures","fastify","h3","uvicorn"],"caveats":["One attempt per cell: these are defect-finding runs, not rates.","Each slice changes the platform build, and the last two also remove the caps. No two slices are a matched comparison.","Costs are list-price estimates for subscription calls, not invoices.","The f0ac3a8a h3 run was verified only after its account-route label was corrected; it counts as unverified here, as declared."],"sourceIds":["agent-coding-calibration"],"stats":{"$k":["id","label","value","unit","display","n"],"$r":[["verified-latest","Verified deliveries, latest build",1,"count","1 of 3",3],["cost-latest","Notional cost, latest build, all 3 tasks",11.06,"usd","$11.06",3],["refusals-trend","Guardrail refusals, first vs latest slice",19,"count","26 → 19",3],["first-run-cost","First calibration run (capped, fastify/session)",4.89,"usd","$4.89, stopped at cap",1]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["calibration-outcomes-by-slice","Three real tasks, four platform builds","Tasks per slice: functional pass, verified delivery, pull request opened","grouped-bar","count","Tasks (of 3)",{"$k":["name","points"],"$r":[["Functional pass (offline gates)",{"$k":["label","value","n"],"$r":[["Baseline (capped)",2,3],["Fix wave 1 (capped)",2,3],["Uncapped, build f0ac3a8a",1,3],["Uncapped, build 236c0d3f",1,3]]}],["Verified delivery",{"$k":["label","value","n","highlight"],"$r":[["Baseline (capped)",0,3,false],["Fix wave 1 (capped)",0,3,false],["Uncapped, build f0ac3a8a",0,3,false],["Uncapped, build 236c0d3f",1,3,true]]}],["Pull request opened",{"$k":["label","value","n"],"$r":[["Baseline (capped)",1,3],["Fix wave 1 (capped)",2,3],["Uncapped, build f0ac3a8a",2,3],["Uncapped, build 236c0d3f",2,3]]}]]},"One attempt per task per slice. The two capped slices stopped at 20 minutes or $5; the uncapped slices had no ceiling. Each slice is a different platform build, so a change is not a matched improvement.",["agent-coding-calibration"]],["calibration-cost-by-task","Notional model cost per task, by slice","\u0001","grouped-bar","usd","USD (notional)",{"$k":["name","points"],"$r":[["Baseline (capped)",{"$k":["label","value","highlight"],"$r":[["fastify/session",3.43,false],["h3js/h3",4,false],["Kludex/uvicorn",4.88,false]]}],["Fix wave 1 (capped)",{"$k":["label","value","highlight"],"$r":[["fastify/session",4.85,false],["h3js/h3",3.29,false],["Kludex/uvicorn",4.63,false]]}],["Uncapped, build f0ac3a8a",{"$k":["label","value","highlight"],"$r":[["fastify/session",2.44,false],["h3js/h3",3.81,false],["Kludex/uvicorn",4.74,false]]}],["Uncapped, build 236c0d3f",{"$k":["label","value","highlight"],"$r":[["fastify/session",3.53,true],["h3js/h3",2.96,true],["Kludex/uvicorn",4.57,true]]}]]},"List-price estimates of subscription calls. Failed and capped attempts count.",["agent-coding-calibration"]],["calibration-minutes-by-task","Wall time per task, by slice","\u0001","grouped-bar","minutes","Minutes",{"$k":["name","points"],"$r":[["Baseline (capped)",{"$k":["label","value","highlight"],"$r":[["fastify/session",9,false],["h3js/h3",11.8,false],["Kludex/uvicorn",17.9,false]]}],["Fix wave 1 (capped)",{"$k":["label","value","highlight"],"$r":[["fastify/session",15.5,false],["h3js/h3",13.5,false],["Kludex/uvicorn",20.3,false]]}],["Uncapped, build f0ac3a8a",{"$k":["label","value","highlight"],"$r":[["fastify/session",7.2,false],["h3js/h3",12.5,false],["Kludex/uvicorn",22.7,false]]}],["Uncapped, build 236c0d3f",{"$k":["label","value","highlight"],"$r":[["fastify/session",11.2,true],["h3js/h3",9.2,true],["Kludex/uvicorn",19.8,true]]}]]},"Onboarding included. uvicorn needed more than the old 20 minute cap once the caps were removed.",["agent-coding-calibration"]],["calibration-guardrail-refusals","Guardrail refusals per slice","Tool calls the platform turned back (plan schema, delivery contract, outgoing text, loops)","bar","count","Refusals (3 tasks)",[{"name":"Refusals","points":{"$k":["label","value","highlight"],"$r":[["Baseline (capped)",26,false],["Fix wave 1 (capped)",28,false],["Uncapped, build f0ac3a8a",25,false],["Uncapped, build 236c0d3f",19,true]]}}],"A refusal is a guardrail working, not always a failure: some catch real problems, some were platform defects that the next build fixed.",["agent-coding-calibration"]]]}}}