{"$k":["slug","entity","aEffort","bEffort","title","seoTitle","description","verdict","rows"],"$r":[["claude-sonnet-5-5-low-vs-medium","claude-sonnet-5-5","low","medium","Claude Sonnet 5.5: low vs medium effort","Claude Sonnet 5.5: low vs medium effort, measured","Claude Sonnet 5.5 at low vs medium effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at low effort and Claude Sonnet 5.5 at medium effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at low effort 81% to 100%; Claude Sonnet 5.5 at medium effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",5.82,7.63,"seconds","5.82 s","7.63 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at low effort 2.78 s to 20.0 s; Claude Sonnet 5.5 at medium effort 2.71 s to 24.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","range","minmax",[2.78,19.96],[2.71,24.01],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",667,770,"tokens","667","770","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01219,0.01352,"usd","$0.012","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.012 vs $0.014) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.004317,0.005946,"usd","$0.0043","$0.0059","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012191,0.01352,"usd","$0.012","$0.014","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort medium","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-low-vs-high","claude-sonnet-5-5","low","high","Claude Sonnet 5.5: low vs high effort","Claude Sonnet 5.5: low vs high effort, measured","Claude Sonnet 5.5 at low vs high effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at low effort and Claude Sonnet 5.5 at high effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at low effort 81% to 100%; Claude Sonnet 5.5 at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",5.82,8.81,"seconds","5.82 s","8.81 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at low effort 2.78 s to 20.0 s; Claude Sonnet 5.5 at high effort 2.93 s to 35.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","range","minmax",[2.78,19.96],[2.93,35.81],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",667,1192,"tokens","667","1,192","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01219,0.01671,"usd","$0.012","$0.017","unclear","No interval or range was recorded for either side, so the gap ($0.012 vs $0.017) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.004317,0.009369,"usd","$0.0043","$0.0094","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012191,0.016705,"usd","$0.012","$0.017","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-low-vs-default","claude-sonnet-5-5","low","default","Claude Sonnet 5.5: low vs default effort","Claude Sonnet 5.5: low vs default effort, measured","Claude Sonnet 5.5 at low vs default effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at low effort and Claude Sonnet 5.5 at default effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at low effort 81% to 100%; Claude Sonnet 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",5.82,7.97,"seconds","5.82 s","7.97 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at low effort 2.78 s to 20.0 s; Claude Sonnet 5.5 at default effort 2.26 s to 21.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[2.78,19.96],[2.26,21.61],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",667,1054,"tokens","667","1,054","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01219,0.01398,"usd","$0.012","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.012 vs $0.014) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.004317,0.006299,"usd","$0.0043","$0.0063","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012191,0.013978,"usd","$0.012","$0.014","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-medium-vs-high","claude-sonnet-5-5","medium","high","Claude Sonnet 5.5: medium vs high effort","Claude Sonnet 5.5: medium vs high effort, measured","Claude Sonnet 5.5 at medium vs high effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at medium effort and Claude Sonnet 5.5 at high effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at medium effort 81% to 100%; Claude Sonnet 5.5 at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.63,8.81,"seconds","7.63 s","8.81 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at medium effort 2.71 s to 24.0 s; Claude Sonnet 5.5 at high effort 2.93 s to 35.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","range","minmax",[2.71,24.01],[2.93,35.81],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",770,1192,"tokens","770","1,192","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01352,0.01671,"usd","$0.014","$0.017","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.017) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005946,0.009369,"usd","$0.0059","$0.0094","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.01352,0.016705,"usd","$0.014","$0.017","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-medium-vs-default","claude-sonnet-5-5","medium","default","Claude Sonnet 5.5: medium vs default effort","Claude Sonnet 5.5: medium vs default effort, measured","Claude Sonnet 5.5 at medium vs default effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at medium effort and Claude Sonnet 5.5 at default effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at medium effort 81% to 100%; Claude Sonnet 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.63,7.97,"seconds","7.63 s","7.97 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at medium effort 2.71 s to 24.0 s; Claude Sonnet 5.5 at default effort 2.26 s to 21.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[2.71,24.01],[2.26,21.61],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",770,1054,"tokens","770","1,054","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01352,0.01398,"usd","$0.014","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.014 vs $0.014) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005946,0.006299,"usd","$0.0059","$0.0063","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.01352,0.013978,"usd","$0.014","$0.014","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-sonnet-5-5-high-vs-default","claude-sonnet-5-5","high","default","Claude Sonnet 5.5: high vs default effort","Claude Sonnet 5.5: high vs default effort, measured","Claude Sonnet 5.5 at high vs default effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Sonnet 5.5 at high effort and Claude Sonnet 5.5 at default effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Sonnet 5.5 at high effort 81% to 100%; Claude Sonnet 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",8.81,7.97,"seconds","8.81 s","7.97 s","unclear","The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at high effort 2.93 s to 35.8 s; Claude Sonnet 5.5 at default effort 2.26 s to 21.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[2.93,35.81],[2.26,21.61],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",1192,1054,"tokens","1,192","1,054","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.01671,0.01398,"usd","$0.017","$0.014","unclear","No interval or range was recorded for either side, so the gap ($0.017 vs $0.014) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.009369,0.006299,"usd","$0.0094","$0.0063","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.016705,0.013978,"usd","$0.017","$0.014","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-low-vs-medium","claude-opus-5-5","low","medium","Claude Opus 5.5: low vs medium effort","Claude Opus 5.5: low vs medium effort, measured","Claude Opus 5.5 at low vs medium effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at low effort and Claude Opus 5.5 at medium effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 81% to 100%; Claude Opus 5.5 at medium effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.5,9.72,"seconds","7.50 s","9.72 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 3.34 s to 15.8 s; Claude Opus 5.5 at medium effort 4.78 s to 31.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","range","minmax",[3.34,15.82],[4.78,31.36],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",594,853,"tokens","594","853","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02115,0.02947,"usd","$0.021","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.021 vs $0.029) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005031,0.01344,"usd","$0.0050","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.021152,0.029475,"usd","$0.021","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort medium","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-low-vs-high","claude-opus-5-5","low","high","Claude Opus 5.5: low vs high effort","Claude Opus 5.5: low vs high effort, measured","Claude Opus 5.5 at low vs high effort: 9 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at low effort and Claude Opus 5.5 at high effort share 9 measured metrics and 5 list-price calculations from 3 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 11 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 80% to 100%; Claude Opus 5.5 at high effort 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.83,2.71,"seconds","2.83 s","2.71 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 2.35 s to 6.62 s; Claude Opus 5.5 at high effort 2.45 s to 11.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","range","minmax",[2.35,6.62],[2.45,11.78],"\u0001"],["Time to first useful output",2.39,2.04,"seconds","2.39 s","2.04 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 1.45 s to 4.90 s; Claude Opus 5.5 at high effort 1.40 s to 9.94 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","range","minmax",[1.45,4.9],[1.4,9.94],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1463,1463,"tokens","1,463","1,463","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",618,619,"tokens","618","619","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",64,78,"tokens","64","78","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00688,0.00694,"usd","$0.0069","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort $0.0058 to $0.018; Claude Opus 5.5 at high effort $0.0059 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","range","minmax",[0.00582,0.01793],[0.00592,0.02708],true],["List-price cost per passing answer (calculation)",0.00829,0.01049,"usd","$0.0083","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0083 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 81% to 100%; Claude Opus 5.5 at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.5,10.11,"seconds","7.50 s","10.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 3.34 s to 15.8 s; Claude Opus 5.5 at high effort 3.63 s to 63.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","range","minmax",[3.34,15.82],[3.63,63],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",594,1052,"tokens","594","1,052","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02115,0.03368,"usd","$0.021","$0.034","unclear","No interval or range was recorded for either side, so the gap ($0.021 vs $0.034) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005031,0.018034,"usd","$0.0050","$0.018","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.021152,0.033677,"usd","$0.021","$0.034","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-low-vs-default","claude-opus-5-5","low","default","Claude Opus 5.5: low vs default effort","Claude Opus 5.5: low vs default effort, measured","Claude Opus 5.5 at low vs default effort: 9 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at low effort and Claude Opus 5.5 at default effort share 9 measured metrics and 5 list-price calculations from 3 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 11 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 80% to 100%; Claude Opus 5.5 at default effort 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.83,2.75,"seconds","2.83 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 2.35 s to 6.62 s; Claude Opus 5.5 at default effort 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.35,6.62],[2.47,8.91],"\u0001"],["Time to first useful output",2.39,1.92,"seconds","2.39 s","1.92 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 1.45 s to 4.90 s; Claude Opus 5.5 at default effort 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[1.45,4.9],[1.56,7.23],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1463,1401,"tokens","1,463","1,401","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",618,680,"tokens","618","680","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",64,64,"tokens","64","64","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00688,0.00688,"usd","$0.0069","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort $0.0058 to $0.018; Claude Opus 5.5 at default effort $0.0059 to $0.022); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00582,0.01793],[0.00592,0.02226],true],["List-price cost per passing answer (calculation)",0.00829,0.01009,"usd","$0.0083","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.0083 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · effort low · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at low effort 81% to 100%; Claude Opus 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",7.5,9.18,"seconds","7.50 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at low effort 3.34 s to 15.8 s; Claude Opus 5.5 at default effort 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[3.34,15.82],[4.24,27.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",594,945,"tokens","594","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02115,0.02893,"usd","$0.021","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.021 vs $0.029) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort low · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.005031,0.013104,"usd","$0.0050","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.021152,0.028925,"usd","$0.021","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort low","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-medium-vs-high","claude-opus-5-5","medium","high","Claude Opus 5.5: medium vs high effort","Claude Opus 5.5: medium vs high effort, measured","Claude Opus 5.5 at medium vs high effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at medium effort and Claude Opus 5.5 at high effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at medium effort 81% to 100%; Claude Opus 5.5 at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",9.72,10.11,"seconds","9.72 s","10.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at medium effort 4.78 s to 31.4 s; Claude Opus 5.5 at high effort 3.63 s to 63.0 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","range","minmax",[4.78,31.36],[3.63,63],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",853,1052,"tokens","853","1,052","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02947,0.03368,"usd","$0.029","$0.034","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.034) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.01344,0.018034,"usd","$0.013","$0.018","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.029475,0.033677,"usd","$0.029","$0.034","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code · effort high","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-medium-vs-default","claude-opus-5-5","medium","default","Claude Opus 5.5: medium vs default effort","Claude Opus 5.5: medium vs default effort, measured","Claude Opus 5.5 at medium vs default effort: 3 measured metrics from one study, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at medium effort and Claude Opus 5.5 at default effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at medium effort 81% to 100%; Claude Opus 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",9.72,9.18,"seconds","9.72 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at medium effort 4.78 s to 31.4 s; Claude Opus 5.5 at default effort 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[4.78,31.36],[4.24,27.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",853,945,"tokens","853","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02947,0.02893,"usd","$0.029","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.029 vs $0.029) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort medium · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.01344,0.013104,"usd","$0.013","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.029475,0.028925,"usd","$0.029","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort medium","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}],["claude-opus-5-5-high-vs-default","claude-opus-5-5","high","default","Claude Opus 5.5: high vs default effort","Claude Opus 5.5: high vs default effort, measured","Claude Opus 5.5 at high vs default effort: 14 measured metrics from 3 studies, with sample sizes, intervals and every failure counted.","Claude Opus 5.5 at high effort and Claude Opus 5.5 at default effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. \"default\" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 4 ties and 22 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Opus 5.5 at high effort 80% to 100%; Claude Opus 5.5 at default effort 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",2.71,2.75,"seconds","2.71 s","2.75 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 2.45 s to 11.8 s; Claude Opus 5.5 at default effort 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[2.45,11.78],[2.47,8.91],"\u0001"],["Time to first useful output",2.04,1.92,"seconds","2.04 s","1.92 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 1.40 s to 9.94 s; Claude Opus 5.5 at default effort 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[1.4,9.94],[1.56,7.23],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",1463,1401,"tokens","1,463","1,401","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",619,680,"tokens","619","680","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",78,64,"tokens","78","64","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00694,0.00688,"usd","$0.0069","$0.0069","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort $0.0059 to $0.027; Claude Opus 5.5 at default effort $0.0059 to $0.022); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","range","minmax",[0.00592,0.02708],[0.00592,0.02226],true],["List-price cost per passing answer (calculation)",0.01049,0.01009,"usd","$0.010","$0.010","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.010) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · effort high · five short validated tasks","Claude Code · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Opus 5.5 at high effort 86% to 100%; Claude Opus 5.5 at default effort 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (24/24)","tie","The 95% intervals overlap (Claude Opus 5.5 at high effort 86% to 100%; Claude Opus 5.5 at default effort 86% to 100%), so this sample cannot separate them.","hard-model-head-to-head",24,"hard-h2h-pass-rate",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","ci95","ci95",[0.862,1],[0.862,1],"\u0001"],["Total time per call on hard tasks (separate batches)",11.03,9.18,"seconds","11.0 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 3.63 s to 63.0 s; Claude Opus 5.5 at default effort 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-total-latency",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[3.63,63],[4.24,27.21],"\u0001"],["Time to first useful output on hard tasks",7.13,6.78,"seconds","7.13 s","6.78 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 2.15 s to 56.2 s; Claude Opus 5.5 at default effort 2.39 s to 21.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",24,"hard-h2h-first-useful-latency",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","range","minmax",[2.15,56.23],[2.39,21.77],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1052,945,"tokens","1,052","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",24,"hard-h2h-output-tokens",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.03337,0.02824,"usd","$0.033","$0.028","unclear","No interval or range was recorded for either side, so the gap ($0.033 vs $0.028) is not tested against run-to-run variation.","hard-model-head-to-head",24,"hard-h2h-cost-per-pass",24,24,"Claude Code · effort high · eight hard validated tasks","Claude Code · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (Claude Opus 5.5 at high effort 81% to 100%; Claude Opus 5.5 at default effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",10.11,9.18,"seconds","10.1 s","9.18 s","unclear","The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 3.63 s to 63.0 s; Claude Opus 5.5 at default effort 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","range","minmax",[3.63,63],[4.24,27.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",1052,945,"tokens","1,052","945","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.03368,0.02893,"usd","$0.034","$0.029","unclear","No interval or range was recorded for either side, so the gap ($0.034 vs $0.029) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Claude Code · effort high · eight hard validated tasks, effort ladder","Claude Code · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",54.43,54.79,"percent","54.4%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-share",24,24,"Claude Code · effort high","Claude Code","range","minmax",[36.14,96.23],[29.92,95.6],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.017969,0.012528,"usd","$0.018","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.00799,0.008003,"usd","$0.0080","$0.0080","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.007407,0.00771,"usd","$0.0074","$0.0077","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-cost-per-call",24,24,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.018034,0.013104,"usd","$0.018","$0.013","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.033677,0.028925,"usd","$0.034","$0.029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Claude Code · effort high","Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",54.43,54.79,"percent","54.4%","54.8%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",24,"thinking-bill-short-vs-hard",24,24,"Claude Code · effort high","Claude Code","range","minmax",[36.14,96.23],[29.92,95.6],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",43.59,0,"percent","43.6%","0%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code · effort high","Claude Code","range","minmax",[0,93.33],[0,93.33],true]]}],["gpt-6-1-sol-codex-cli-low-vs-medium","gpt-6-1-sol-codex-cli","low","medium","GPT-6.1 Sol (Codex CLI): low vs medium effort","GPT-6.1 Sol (Codex CLI): low vs medium effort, measured","GPT-6.1 Sol (Codex CLI) at low vs medium effort: 9 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (Codex CLI) at low effort and GPT-6.1 Sol (Codex CLI) at medium effort share 9 measured metrics and 5 list-price calculations from 3 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 11 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 10 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation","n"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (10/10)","100% (15/15)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at low effort 72% to 100%; GPT-6.1 Sol (Codex CLI) at medium effort 80% to 100%), so this sample cannot separate them.","model-head-to-head","h2h-pass-rate",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.7225,1],[0.7961,1],"\u0001","\u0001"],["Total time per call",6.26,5.65,"seconds","6.26 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 4.65 s to 10.5 s; GPT-6.1 Sol (Codex CLI) at medium effort 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-total-latency",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[4.65,10.47],[4.1,25.46],"\u0001","\u0001"],["Time to first useful output",5.14,5.05,"seconds","5.14 s","5.05 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 4.02 s to 8.50 s; GPT-6.1 Sol (Codex CLI) at medium effort 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-first-useful-latency",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[4.02,8.5],[3.36,17.82],"\u0001","\u0001"],["Input tokens per call: what the CLI sends (Cache read)",8064,5180,"tokens","8,064","5,180","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head","h2h-input-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",4059,6943,"tokens","4,059","6,943","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head","h2h-input-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",42,42,"tokens","42","42","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head","h2h-output-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00769,0.01018,"usd","$0.0077","$0.010","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort $0.0074 to $0.026; GPT-6.1 Sol (Codex CLI) at medium effort $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-list-price-per-call",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.00742,0.02649],[0.0054,0.02686],true,"\u0001"],["List-price cost per passing answer (calculation)",0.00998,0.01564,"usd","$0.010","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.016) is not tested against run-to-run variation.","model-head-to-head","h2h-cost-per-pass",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true,"\u0001"],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at low effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%), so this sample cannot separate them.","effort-ladder","effort-ladder-pass-rate",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001",16],["Total time per call by effort on hard tasks",13.62,13.11,"seconds","13.6 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 7.94 s to 44.3 s; GPT-6.1 Sol (Codex CLI) at medium effort 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder","effort-ladder-total-latency",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","range","minmax",[7.94,44.29],[8.54,61.6],"\u0001",16],["Output tokens per call by effort on hard tasks (Output tokens)",284,335,"tokens","284","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder","effort-ladder-output-tokens",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001",16],["List-price cost per strict pass by effort (calculation)",0.01284,0.02564,"usd","$0.013","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.013 vs $0.026) is not tested against run-to-run variation.","effort-ladder","effort-ladder-cost-per-pass",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort medium · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true,16],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.001223,0.002273,"usd","$0.0012","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","thinking-bill-by-effort",16,16,"Codex CLI · effort low","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true,16],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012837,0.025637,"usd","$0.013","$0.026","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","thinking-bill-by-effort",16,16,"Codex CLI · effort low","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true,16]]}],["gpt-6-1-sol-codex-cli-low-vs-high","gpt-6-1-sol-codex-cli","low","high","GPT-6.1 Sol (Codex CLI): low vs high effort","GPT-6.1 Sol (Codex CLI): low vs high effort, measured","GPT-6.1 Sol (Codex CLI) at low vs high effort: 14 measured metrics from 3 studies, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (Codex CLI) at low effort and GPT-6.1 Sol (Codex CLI) at high effort share 14 measured metrics and 5 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 16 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation","n"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (10/10)","100% (15/15)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at low effort 72% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 80% to 100%), so this sample cannot separate them.","model-head-to-head","h2h-pass-rate",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","ci95","ci95",[0.7225,1],[0.7961,1],"\u0001","\u0001"],["Total time per call",6.26,5.6,"seconds","6.26 s","5.60 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 4.65 s to 10.5 s; GPT-6.1 Sol (Codex CLI) at high effort 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-total-latency",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[4.65,10.47],[4.05,19.52],"\u0001","\u0001"],["Time to first useful output",5.14,5.32,"seconds","5.14 s","5.32 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 4.02 s to 8.50 s; GPT-6.1 Sol (Codex CLI) at high effort 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-first-useful-latency",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[4.02,8.5],[3.64,16.37],"\u0001","\u0001"],["Input tokens per call: what the CLI sends (Cache read)",8064,6716,"tokens","8,064","6,716","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head","h2h-input-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",4059,5406,"tokens","4,059","5,406","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head","h2h-input-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",42,42,"tokens","42","42","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head","h2h-output-tokens",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00769,0.01047,"usd","$0.0077","$0.010","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort $0.0074 to $0.026; GPT-6.1 Sol (Codex CLI) at high effort $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head","h2h-list-price-per-call",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[0.00742,0.02649],[0.0066,0.02812],true,"\u0001"],["List-price cost per passing answer (calculation)",0.00998,0.01322,"usd","$0.010","$0.013","unclear","No interval or range was recorded for either side, so the gap ($0.010 vs $0.013) is not tested against run-to-run variation.","model-head-to-head","h2h-cost-per-pass",10,15,"Codex CLI · effort low · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true,"\u0001"],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at low effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder","effort-ladder-pass-rate",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001",16],["Total time per call by effort on hard tasks",13.62,18.12,"seconds","13.6 s","18.1 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 7.94 s to 44.3 s; GPT-6.1 Sol (Codex CLI) at high effort 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder","effort-ladder-total-latency",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","range","minmax",[7.94,44.29],[11.67,92.21],"\u0001",16],["Output tokens per call by effort on hard tasks (Output tokens)",284,436,"tokens","284","436","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder","effort-ladder-output-tokens",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001",16],["List-price cost per strict pass by effort (calculation)",0.01284,0.01514,"usd","$0.013","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.013 vs $0.015) is not tested against run-to-run variation.","effort-ladder","effort-ladder-cost-per-pass",16,16,"Codex CLI · effort low · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true,16],["CLI vs API: time for a one-line answer (Total time)",4.18,4.19,"seconds","4.18 s","4.19 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 3.86 s to 4.53 s; GPT-6.1 Sol (Codex CLI) at high effort 3.81 s to 4.69 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens","cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort low · fixed exact reply, 5 runs","Codex CLI · effort high · fixed exact reply, 5 runs","range","minmax",[3.86,4.53],[3.81,4.69],"\u0001",5],["CLI vs API: time for a one-line answer (First useful output)",3.75,3.79,"seconds","3.75 s","3.79 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at low effort 3.44 s to 4.10 s; GPT-6.1 Sol (Codex CLI) at high effort 3.37 s to 4.30 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens","cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort low · fixed exact reply, 5 runs","Codex CLI · effort high · fixed exact reply, 5 runs","range","minmax",[3.44,4.1],[3.37,4.3],"\u0001",5],["CLI vs API: time for a small coding task (Total time)",14.15,17.85,"seconds","14.2 s","17.9 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) at low effort 13.0 s to 14.4 s; GPT-6.1 Sol (Codex CLI) at high effort 17.7 s to 22.4 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens","cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort low · small coding task, 3 runs","Codex CLI · effort high · small coding task, 3 runs","range","minmax",[13.02,14.41],[17.68,22.42],"\u0001",3],["CLI vs API: time for a small coding task (First useful output)",13.6,17.27,"seconds","13.6 s","17.3 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) at low effort 12.5 s to 13.8 s; GPT-6.1 Sol (Codex CLI) at high effort 17.1 s to 21.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens","cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort low · small coding task, 3 runs","Codex CLI · effort high · small coding task, 3 runs","range","minmax",[12.52,13.83],[17.13,21.86],"\u0001",3],["Hidden prompt: input tokens for the same one-line request",19551,19555,"tokens","19,551","19,555","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens","cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort low · short fixed tasks","Codex CLI · effort high · short fixed tasks","\u0001","\u0001","\u0001","\u0001","\u0001",5],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.001223,0.004114,"usd","$0.0012","$0.0041","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","thinking-bill-by-effort",16,16,"Codex CLI · effort low","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true,16],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.012837,0.015137,"usd","$0.013","$0.015","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","thinking-bill-by-effort",16,16,"Codex CLI · effort low","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true,16]]}],["gpt-6-1-sol-codex-cli-medium-vs-high","gpt-6-1-sol-codex-cli","medium","high","GPT-6.1 Sol (Codex CLI): medium vs high effort","GPT-6.1 Sol (Codex CLI): medium vs high effort, measured","GPT-6.1 Sol (Codex CLI) at medium vs high effort: 14 measured metrics from 3 studies, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (Codex CLI) at medium effort and GPT-6.1 Sol (Codex CLI) at high effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 5 ties and 21 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 80% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",5.65,5.6,"seconds","5.65 s","5.60 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 4.10 s to 25.5 s; GPT-6.1 Sol (Codex CLI) at high effort 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[4.1,25.46],[4.05,19.52],"\u0001"],["Time to first useful output",5.05,5.32,"seconds","5.05 s","5.32 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 3.36 s to 17.8 s; GPT-6.1 Sol (Codex CLI) at high effort 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[3.36,17.82],[3.64,16.37],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",5180,6716,"tokens","5,180","6,716","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",6943,5406,"tokens","6,943","5,406","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",42,42,"tokens","42","42","tie","Same value. More or fewer is not better by itself for this metric.","model-head-to-head",15,"h2h-output-tokens",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.01018,0.01047,"usd","$0.010","$0.010","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort $0.0054 to $0.027; GPT-6.1 Sol (Codex CLI) at high effort $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","range","minmax",[0.0054,0.02686],[0.0066,0.02812],true],["List-price cost per passing answer (calculation)",0.01564,0.01322,"usd","$0.016","$0.013","unclear","No interval or range was recorded for either side, so the gap ($0.016 vs $0.013) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Codex CLI · effort medium · five short validated tasks","Codex CLI · effort high · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head",16,"hard-h2h-pass-rate",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head",16,"hard-h2h-pass-rate",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",13.11,18.12,"seconds","13.1 s","18.1 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 8.54 s to 61.6 s; GPT-6.1 Sol (Codex CLI) at high effort 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",16,"hard-h2h-total-latency",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","range","minmax",[8.54,61.6],[11.67,92.21],"\u0001"],["Time to first useful output on hard tasks",10.23,12.69,"seconds","10.2 s","12.7 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 6.09 s to 40.4 s; GPT-6.1 Sol (Codex CLI) at high effort 8.93 s to 75.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head",16,"hard-h2h-first-useful-latency",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","range","minmax",[6.09,40.41],[8.93,75.91],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",335,436,"tokens","335","436","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head",16,"hard-h2h-output-tokens",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.02564,0.01514,"usd","$0.026","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.026 vs $0.015) is not tested against run-to-run variation.","hard-model-head-to-head",16,"hard-h2h-cost-per-pass",16,16,"Codex CLI · effort medium · eight hard validated tasks","Codex CLI · effort high · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Strict pass rate by effort on eight hard tasks",1,1,"rate","100% (16/16)","100% (16/16)","tie","The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.","effort-ladder",16,"effort-ladder-pass-rate",16,16,"Codex CLI · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","ci95","ci95",[0.8064,1],[0.8064,1],"\u0001"],["Total time per call by effort on hard tasks",13.11,18.12,"seconds","13.1 s","18.1 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 8.54 s to 61.6 s; GPT-6.1 Sol (Codex CLI) at high effort 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","effort-ladder",16,"effort-ladder-total-latency",16,16,"Codex CLI · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","range","minmax",[8.54,61.6],[11.67,92.21],"\u0001"],["Output tokens per call by effort on hard tasks (Output tokens)",335,436,"tokens","335","436","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","effort-ladder",16,"effort-ladder-output-tokens",16,16,"Codex CLI · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass by effort (calculation)",0.02564,0.01514,"usd","$0.026","$0.015","unclear","No interval or range was recorded for either side, so the gap ($0.026 vs $0.015) is not tested against run-to-run variation.","effort-ladder",16,"effort-ladder-cost-per-pass",16,16,"Codex CLI · effort medium · eight hard validated tasks, effort ladder","Codex CLI · effort high · eight hard validated tasks, effort ladder","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",46.33,57.01,"percent","46.3%","57%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-share",16,16,"Codex CLI · effort medium","Codex CLI · effort high","range","minmax",[11.42,86.85],[29.19,90.8],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.002273,0.004114,"usd","$0.0023","$0.0041","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-cost-per-call",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.003025,0.002905,"usd","$0.0030","$0.0029","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-cost-per-call",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.020339,0.008117,"usd","$0.020","$0.0081","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-cost-per-call",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)",0.002273,0.004114,"usd","$0.0023","$0.0041","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)",0.025637,0.015137,"usd","$0.026","$0.015","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-by-effort",16,16,"Codex CLI · effort medium","Codex CLI · effort high","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",46.33,57.01,"percent","46.3%","57%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",16,"thinking-bill-short-vs-hard",16,16,"Codex CLI · effort medium","Codex CLI · effort high","range","minmax",[11.42,86.85],[29.19,90.8],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",41.05,58.06,"percent","41%","58.1%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Codex CLI · effort medium","Codex CLI · effort high","range","minmax",[0,71.43],[0,75.76],true]]}],["gpt-6-1-sol-openai-api-low-vs-high","gpt-6-1-sol-openai-api","low","high","GPT-6.1 Sol (OpenAI API): low vs high effort","GPT-6.1 Sol (OpenAI API): low vs high effort, measured","GPT-6.1 Sol (OpenAI API) at low vs high effort: 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","GPT-6.1 Sol (OpenAI API) at low effort and GPT-6.1 Sol (OpenAI API) at high effort share 5 measured metrics from one study. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 4 unclear; each row says why. Some rows rest on small samples (n = 3 at the smallest).",{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",1.02,1.52,"seconds","1.02 s","1.52 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.96 s to 1.87 s; GPT-6.1 Sol (OpenAI API) at high effort 1.35 s to 2.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort low · fixed exact reply, 5 runs","OpenAI API · effort high · fixed exact reply, 5 runs","range","minmax",[0.96,1.87],[1.35,2.23]],["CLI vs API: time for a one-line answer (First useful output)",0.87,1.34,"seconds","0.87 s","1.34 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.84 s to 1.74 s; GPT-6.1 Sol (OpenAI API) at high effort 1.26 s to 2.12 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort low · fixed exact reply, 5 runs","OpenAI API · effort high · fixed exact reply, 5 runs","range","minmax",[0.84,1.74],[1.26,2.12]],["CLI vs API: time for a small coding task (Total time)",6,9.56,"seconds","6.00 s","9.56 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) at low effort 5.44 s to 6.20 s; GPT-6.1 Sol (OpenAI API) at high effort 9.44 s to 10.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort low · small coding task, 3 runs","OpenAI API · effort high · small coding task, 3 runs","range","minmax",[5.44,6.2],[9.44,10.94]],["CLI vs API: time for a small coding task (First useful output)",1.05,5.31,"seconds","1.05 s","5.31 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) at low effort 0.97 s to 1.40 s; GPT-6.1 Sol (OpenAI API) at high effort 4.99 s to 6.42 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort low · small coding task, 3 runs","OpenAI API · effort high · small coding task, 3 runs","range","minmax",[0.97,1.4],[4.99,6.42]],["Hidden prompt: input tokens for the same one-line request",17,17,"tokens","17","17","tie","Same value. More or fewer is not better by itself for this metric.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"OpenAI API · effort low · short fixed tasks","OpenAI API · effort high · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}]]}