{"i":159,"comparison":{"slug":"agent-harness-vs-kimi-k2-5","a":"agent-harness","b":"kimi-k2-5","title":"Agent vs Kimi K2.5","seoTitle":"Agent vs Kimi K2.5: measured benchmarks","description":"Agent vs Kimi K2.5: 2 measured metrics from one study (Resolved rate on the same 33 SWE-bench Verified instances; more), with sample sizes and intervals.","verdict":"Agent and Kimi K2.5 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; Kimi K2.5 ran under a different, simpler harness, so this compares systems, not models.","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.7576,0.697,"rate","76% (25/33)","70% (23/33)","tie","The 95% intervals overlap (Agent 59% to 87%; Kimi K2.5 53% to 83%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.5898,0.8717],[0.5266,0.8262],"\u0001"],["Model calls per instance",49.5,56.7,"calls","49.5","56.7","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"full pipeline on Claude Sonnet 5.5","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",3.706,0.256,"usd","$3.71","$0.26","unclear","No interval or range was recorded for either side, so the gap ($3.71 vs $0.26, 14x) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",25,23,"full pipeline on Claude Sonnet 5.5 · notional","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001",true]]}}}