{"i":6,"study":{"slug":"effort-ladder","title":"Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks","seoTitle":"Effort ladder: does more AI effort buy quality?","description":"176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.","question":"On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate for Sonnet, Opus and GPT-6.1 Sol, and what does it cost in time, tokens and list price per pass?","answer":"Not on this set. Every one of the 11 configurations (Sonnet at low, medium, high and default, Opus at low, medium, high and default and GPT-6.1 Sol through Codex CLI at low, medium and high) passed all 16 calls strictly (95% interval 81% to 100% each), so pass rate does not separate any effort level. The hard set has a ceiling for these models: with 16 calls per cell it cannot rule out a difference of up to about 19 points. What more effort did change is output and time. Median total time per call by effort: Sonnet: low 5.8 s, medium 7.6 s, high 8.8 s, default 8.0 s (median output tokens 667 / 770 / 1,192 / 1,054); Opus: low 7.5 s, medium 9.7 s, high 10.1 s, default 9.2 s (median output tokens 594 / 853 / 1,052 / 945); GPT-6.1 Sol (Codex CLI): low 13.6 s, medium 13.1 s, high 18.1 s (median output tokens 284 / 335 / 436). Within each model, the fastest and slowest calls of every effort overlap, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on subscriptions), cost per strict pass went from $0.0122 at low to $0.0167 at high for Sonnet, $0.0212 at low to $0.0337 at high for Opus and $0.0128 at low to $0.0151 at high for GPT-6.1 Sol (Codex CLI).","date":"2026-10-06","updated":"2026-10-06","tags":["effort","reasoning-effort","hard-tasks","claude-sonnet","claude-opus","gpt-6-1-sol","claude-code","codex-cli","latency","tokens"],"caveats":["Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.","Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.","Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.","Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.","List-price costs are calculations; the calls used flat subscriptions."],"sourceIds":["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"],"stats":{"$k":["id","label","value","unit","display","n","ci"],"$r":[["effort-ladder-pass-new","New effort-ladder calls that passed strictly",1,"rate","100% (96/96)",96,[0.9615,1]],["effort-ladder-pass-all","All effort-ladder calls that passed strictly, reference cells included",1,"rate","100% (176/176)",176,[0.9786,1]],["effort-ladder-cells","Configurations on the ladder (new + reference)",11,"count","11 (6 new, 5 reference)",176,"\u0001"],["effort-ladder-format-misses","Format misses and wrong answers on the ladder",0,"count","0 format misses, 0 wrong answers",176,"\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","whisker","sourceIds","xLabel"],"$r":[["effort-ladder-pass-rate","Strict pass rate by effort on eight hard tasks","Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals","dot-range","rate","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Sonnet 5.5 · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (low) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (medium) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 (high) · Claude Code",1,0.8064,1,16],["Claude Opus 5.5 · Claude Code",1,0.8064,1,16],["GPT-6.1 Sol (low) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16]]}}],"Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.","ci95",["agent-effort-ladder"],"\u0001"],["effort-ladder-total-latency","Total time per call by effort on hard tasks","Median per configuration; whiskers = fastest and slowest call","dot-range","seconds","Seconds",[{"name":"Total time per call by effort on hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",5.82,2.78,19.96,16],["Claude Sonnet 5.5 (medium) · Claude Code",7.63,2.71,24.01,16],["Claude Sonnet 5.5 (high) · Claude Code",8.81,2.93,35.81,16],["Claude Sonnet 5.5 · Claude Code",7.97,2.26,21.61,16],["Claude Opus 5.5 (low) · Claude Code",7.5,3.34,15.82,16],["Claude Opus 5.5 (medium) · Claude Code",9.72,4.78,31.36,16],["Claude Opus 5.5 (high) · Claude Code",10.11,3.63,63,16],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,16],["GPT-6.1 Sol (low) · Codex CLI",13.62,7.94,44.29,16],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16]]}}],"Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.","minmax",["agent-effort-ladder"],"\u0001"],["effort-ladder-time-by-effort","Median total time per call, by effort","One line per model and route; default = the effort flag was not passed","line","seconds","Seconds (median)",{"$k":["name","points"],"$r":[["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["low",5.82,16],["medium",7.63,16],["high",8.81,16],["default",7.97,16]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["low",7.5,16],["medium",9.72,16],["high",10.11,16],["default",9.18,16]]}],["GPT-6.1 Sol · Codex CLI",{"$k":["label","value","n"],"$r":[["low",13.62,16],["medium",13.11,16],["high",18.12,16]]}]]},"Medians only; the per-call ranges are in the total-time chart and they overlap. \"default\" is placed last because its level is not known: the CLI chose it.","\u0001",["agent-effort-ladder"],"Effort"],["effort-ladder-output-tokens","Output tokens per call by effort on hard tasks","Median per configuration; reasoning tokens as the CLI reports them","grouped-bar","tokens","Tokens",[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",667,16],["Claude Sonnet 5.5 (medium) · Claude Code",770,16],["Claude Sonnet 5.5 (high) · Claude Code",1192,16],["Claude Sonnet 5.5 · Claude Code",1054,16],["Claude Opus 5.5 (low) · Claude Code",594,16],["Claude Opus 5.5 (medium) · Claude Code",853,16],["Claude Opus 5.5 (high) · Claude Code",1052,16],["Claude Opus 5.5 · Claude Code",945,16],["GPT-6.1 Sol (low) · Codex CLI",284,16],["GPT-6.1 Sol (medium) · Codex CLI",335,16],["GPT-6.1 Sol (high) · Codex CLI",436,16]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",273,16],["Claude Sonnet 5.5 (medium) · Claude Code",422,16],["Claude Sonnet 5.5 (high) · Claude Code",745,16],["Claude Sonnet 5.5 · Claude Code",668,16],["Claude Opus 5.5 (low) · Claude Code",87,16],["Claude Opus 5.5 (medium) · Claude Code",518,16],["Claude Opus 5.5 (high) · Claude Code",614,16],["Claude Opus 5.5 · Claude Code",538,16],["GPT-6.1 Sol (low) · Codex CLI",63,16],["GPT-6.1 Sol (medium) · Codex CLI",150,16],["GPT-6.1 Sol (high) · Codex CLI",225,16]]}}],"Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean \"not reported\". Their content is never captured. More tokens is not better or worse by itself.","\u0001",["agent-effort-ladder"],"\u0001"],["effort-ladder-cost-per-pass","List-price cost per strict pass by effort (calculation)","All calls in a configuration divided by its strict passes","bar","usd","USD per strict pass",[{"name":"Cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.01219,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.01352,16],["Claude Sonnet 5.5 (high) · Claude Code",0.01671,16],["Claude Sonnet 5.5 · Claude Code",0.01398,16],["Claude Opus 5.5 (low) · Claude Code",0.02115,16],["Claude Opus 5.5 (medium) · Claude Code",0.02947,16],["Claude Opus 5.5 (high) · Claude Code",0.03368,16],["Claude Opus 5.5 · Claude Code",0.02893,16],["GPT-6.1 Sol (low) · Codex CLI",0.01284,16],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16]]}}],"Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.","\u0001",["agent-effort-ladder","calc-repricing","price-anthropic","price-openai"],"\u0001"]]},"related":["hard-model-head-to-head","caching-consistency"]}}