{"i":4,"study":{"slug":"hard-model-head-to-head","title":"Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks","seoTitle":"Hard tasks: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol","description":"152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.","question":"On 8 hard tasks with deterministic validators, does pass rate separate the Claude Code models and GPT-6.1 Sol through the Codex CLI, and what do speed, tokens and cost per pass add?","answer":"139 of 152 calls that reached a model passed strictly (91%). 6 of 7 configurations passed every call: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, Claude Opus 5.5 (high) · Claude Code, Claude Fable 5.1 · Claude Code (24/24 each, 95% interval 86% to 100%) and GPT-6.1 Sol (medium) · Codex CLI, GPT-6.1 Sol (high) · Codex CLI (16/16 each, 95% interval 81% to 100%), so the hard set still has a ceiling for these models and pass rate does not separate them. Claude Haiku 4.5 · Claude Code passed 11/24 strictly (46%, 95% interval 28% to 65%). 5 more replies had the right answer in the wrong format (for example inside a code fence), so 16/24 on a lenient reading (47% to 82%); 8 replies were wrong. It passed none of the tasks “Predict JavaScript event-loop output order”, “Solve a multi-constraint room schedule” and “Write a SQLite reporting query (fan-out, ties, boundaries)”. Median total time per call was Sonnet 7.7 s, Opus 9.2 s, Opus (high) 11.0 s, GPT-6.1 Sol (medium) 13.1 s, Fable 16.1 s, GPT-6.1 Sol (high) 18.1 s, Haiku 39.0 s. The fastest and slowest single calls of every configuration overlap with every other, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on a subscription), the lowest cost per strict pass was Claude Sonnet 5.5 · Claude Code at $0.0143; the quality-vs-cost frontier is Claude Sonnet 5.5 · Claude Code. 30 earlier attempts were blocked before any model call (Codex CLI: the CLI reported no signed-in account); they are reported, not scored, and that route ran in a later batch.","date":"2026-10-06","updated":"2026-10-06","tags":["head-to-head","hard-tasks","claude-haiku","claude-sonnet","claude-opus","claude-fable","gpt-6-1-sol","codex-cli","format-misses","latency"],"caveats":["6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.","Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.","Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.","The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.","Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.","CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.","Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.","List-price costs are calculations; the calls used a flat subscription."],"sourceIds":["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["hard-h2h-pass-all","Calls that passed strictly (hard set)",0.9145,"rate","91% (139/152)",152,[0.8592,0.9493],"\u0001"],["hard-h2h-correct-all","Calls with a correct answer, format misses included (lenient reading)",0.9474,"rate","95% (144/152)",152,[0.8996,0.9731],"\u0001"],["hard-h2h-format-misses","Non-passes that were format misses, not wrong answers",5,"count","5 of 13 non-passes (8 wrong answers)",13,"\u0001","\u0001"],["hard-h2h-perfect-configs","Configurations that passed every call",6,"count","6 of 7 (4 at 24/24, 2 at 16/16)",7,"\u0001","\u0001"],["hard-h2h-fastest-perfect","Lowest observed median time among configurations that passed every call (separate batches)",7.75,"seconds","Claude Sonnet 5.5 · Claude Code: 7.7 s",24,"\u0001","The counted Claude and Codex batches ran hours apart on one host and network. Host load was not controlled; this does not isolate model speed."],["hard-h2h-cheapest-per-pass","Lowest list-price cost per strict pass (calculation)",0.01435,"usd","Claude Sonnet 5.5 · Claude Code: $0.0143",24,"\u0001","\u0001"],["hard-h2h-blocked","Attempts blocked before any model call (not scored)",30,"count","30 (Codex CLI; 0 model calls)",182,"\u0001","\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds","xLabel"],"$r":[["hard-h2h-pass-rate","Pass rate on eight hard tasks","Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","dot-range","rate","Passed",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.4583,0.2789,0.6493,24]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 · Claude Code",1,0.862,1,24],["Claude Opus 5.5 (high) · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (medium) · Codex CLI",1,0.8064,1,16],["Claude Fable 5.1 · Claude Code",1,0.862,1,24],["GPT-6.1 Sol (high) · Codex CLI",1,0.8064,1,16],["Claude Haiku 4.5 · Claude Code",0.6667,0.4671,0.8203,24]]}}],"Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.",["agent-provider-h2h-hard"],"\u0001"],["hard-h2h-outcomes","What happened on every call","Counts per configuration: strict passes, format misses and wrong answers","stacked-bar","count","Calls",{"$k":["name","points"],"$r":[["Strict pass",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",24,24],["Claude Opus 5.5 · Claude Code",24,24],["Claude Opus 5.5 (high) · Claude Code",24,24],["GPT-6.1 Sol (medium) · Codex CLI",16,16],["Claude Fable 5.1 · Claude Code",24,24],["GPT-6.1 Sol (high) · Codex CLI",16,16],["Claude Haiku 4.5 · Claude Code",11,24]]}],["Format miss (correct answer, wrong format)",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0,24],["Claude Opus 5.5 · Claude Code",0,24],["Claude Opus 5.5 (high) · Claude Code",0,24],["GPT-6.1 Sol (medium) · Codex CLI",0,16],["Claude Fable 5.1 · Claude Code",0,24],["GPT-6.1 Sol (high) · Codex CLI",0,16],["Claude Haiku 4.5 · Claude Code",5,24]]}],["Wrong answer",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0,24],["Claude Opus 5.5 · Claude Code",0,24],["Claude Opus 5.5 (high) · Claude Code",0,24],["GPT-6.1 Sol (medium) · Codex CLI",0,16],["Claude Fable 5.1 · Claude Code",0,24],["GPT-6.1 Sol (high) · Codex CLI",0,16],["Claude Haiku 4.5 · Claude Code",8,24]]}]]},"A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.",["agent-provider-h2h-hard"],"\u0001"],["hard-h2h-total-latency","Total time per call on hard tasks (separate batches)","Median per configuration; whiskers = fastest and slowest call","dot-range","seconds","Seconds",[{"name":"Total time per call on hard tasks (separate batches)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",7.75,2.26,34.79,24,true],["Claude Opus 5.5 · Claude Code",9.18,4.24,27.21,24,true],["Claude Opus 5.5 (high) · Claude Code",11.03,3.63,63,24,true],["GPT-6.1 Sol (medium) · Codex CLI",13.11,8.54,61.6,16,true],["Claude Fable 5.1 · Claude Code",16.13,4.46,90,24,true],["GPT-6.1 Sol (high) · Codex CLI",18.12,11.67,92.21,16,true],["Claude Haiku 4.5 · Claude Code",39.01,15.27,75.13,24,false]]}}],"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.",["agent-provider-h2h-hard"],"\u0001"],["hard-h2h-first-useful-latency","Time to first useful output on hard tasks","Median per configuration; whiskers = fastest and slowest call","dot-range","seconds","Seconds",[{"name":"Time to first useful output on hard tasks","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",5.95,0.86,30.57,24,true],["Claude Opus 5.5 · Claude Code",6.78,2.39,21.77,24,true],["Claude Opus 5.5 (high) · Claude Code",7.13,2.15,56.23,24,true],["GPT-6.1 Sol (medium) · Codex CLI",10.23,6.09,40.41,16,true],["Claude Fable 5.1 · Claude Code",11.63,2,85.33,24,true],["GPT-6.1 Sol (high) · Codex CLI",12.69,8.93,75.91,16,true],["Claude Haiku 4.5 · Claude Code",35.54,12.88,70.31,24,false]]}}],"One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.",["agent-provider-h2h-hard"],"\u0001"],["hard-h2h-output-tokens","Output tokens per call on hard tasks","Median per configuration; reasoning tokens as the CLI reports them","grouped-bar","tokens","Tokens",[{"name":"Output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",1050,24],["Claude Opus 5.5 · Claude Code",945,24],["Claude Opus 5.5 (high) · Claude Code",1052,24],["GPT-6.1 Sol (medium) · Codex CLI",335,16],["Claude Fable 5.1 · Claude Code",1366,24],["GPT-6.1 Sol (high) · Codex CLI",436,16],["Claude Haiku 4.5 · Claude Code",5064,24]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",585,24],["Claude Opus 5.5 · Claude Code",529,24],["Claude Opus 5.5 (high) · Claude Code",614,24],["GPT-6.1 Sol (medium) · Codex CLI",150,16],["Claude Fable 5.1 · Claude Code",889,24],["GPT-6.1 Sol (high) · Codex CLI",225,16],["Claude Haiku 4.5 · Claude Code",4556,24]]}}],"Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.",["agent-provider-h2h-hard"],"\u0001"],["hard-h2h-cost-per-pass","List-price cost per strict pass on hard tasks (calculation)","All calls in a configuration, failures and format misses included, divided by its strict passes","bar","usd","USD per strict pass",[{"name":"Cost per strict pass","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.01435,24,true],["GPT-6.1 Sol (high) · Codex CLI",0.01514,16,false],["GPT-6.1 Sol (medium) · Codex CLI",0.02564,16,false],["Claude Opus 5.5 · Claude Code",0.02824,24,false],["Claude Opus 5.5 (high) · Claude Code",0.03337,24,false],["Claude Haiku 4.5 · Claude Code",0.0672,24,false],["Claude Fable 5.1 · Claude Code",0.09331,24,false]]}}],"Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.",["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],"\u0001"],["hard-h2h-frontier","Quality vs cost frontier on hard tasks","Strict pass rate against list-price cost per strict pass","scatter","rate","Strict pass rate",[{"name":"Claude Code","points":{"$k":["label","x","value","n","highlight"],"$r":[["Claude Sonnet 5.5 · Claude Code",0.01435,1,24,true],["Claude Opus 5.5 · Claude Code",0.02824,1,24,false],["Claude Opus 5.5 (high) · Claude Code",0.03337,1,24,false],["Claude Fable 5.1 · Claude Code",0.09331,1,24,false],["Claude Haiku 4.5 · Claude Code",0.0672,0.4583,24,false]]}},{"name":"Codex CLI","points":[{"label":"GPT-6.1 Sol (medium) · Codex CLI","x":0.02564,"value":1,"n":16,"highlight":false},{"label":"GPT-6.1 Sol (high) · Codex CLI","x":0.01514,"value":1,"n":16,"highlight":false}]}],"Upper-left is better. Highlighted points are on the frontier: no other configuration passes at least as often for at most the same cost per pass. Frontier: Claude Sonnet 5.5 · Claude Code. Costs are calculations from tokens. Pass rates with their 95% intervals are in the pass-rate chart.",["agent-provider-h2h-hard","calc-repricing","price-anthropic","price-openai"],"USD per strict pass (list-price calculation)"]]},"related":["model-head-to-head"]}}