{"i":25,"study":{"slug":"harder-tasks-head-to-head","title":"GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks","seoTitle":"GPT-6.1 Sol vs Claude Opus 5.5 on harder tasks","description":"56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.","question":"On a task set built so that Claude Sonnet 5.5 did not pass it every time, does pass rate separate GPT-6.1 Sol (Codex CLI) from Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 (Claude Code)?","answer":"22 of 56 counted calls passed strictly (39%, 95% Wilson interval 28% to 52%) on 4 tasks. The Sonnet pilot did not pass these twice. GPT-6.1 Sol (medium) passed 11/16 strictly (69%, 95% interval 44% to 86%). Opus 5.5 passed 5/12 strictly (42%, 95% interval 19% to 68%). Sonnet 5.5 passed 6/16 strictly (38%, 95% interval 18% to 61%). Haiku 4.5 passed 0/12 strictly (0%, 95% interval 0% to 24%). GPT-6.1 Sol (medium) is ahead of Haiku 4.5 (the 95% intervals do not overlap). The other 5 of 6 pairs overlap, so this set cannot rank them. The lenient reading counts format misses. Opus 5.5 is also ahead of Haiku 4.5 on the lenient reading. Opus 5.5 6/12 (n = 12, 95% Wilson interval 25.4% to 74.6%); Haiku 4.5 0/12 (n = 12, 95% Wilson interval 0.0% to 24.2%). The intervals miss by 1.1 points (calculation), so this is fragile. Haiku 4.5 passed none (a floor for that configuration on this set). No configuration passed every call across the full set. Some per-task cells still hit a ceiling; see the task chart. Of 34 non-passes, 1 was a format miss with the right answer. Another 21 were wrong answers. 12 gave no answer; 10 hit the 300 s timeout. 11 of 56 calls tried a tool although tools were off. Per configuration: GPT-6.1 Sol (medium) 0 of 16, Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12. None of them passed. The Codex runner also tells the model not to call tools; the Claude Code runner does not. Median total time per completed call (wrong answers and format misses included; timeouts and tool-call parse errors excluded; ranges are not intervals): GPT-6.1 Sol (medium) 120.2 s (n = 13, range 46.2 s to 273.5 s). Opus 5.5 80.3 s (n = 9, range 3.8 s to 279.5 s). Sonnet 5.5 70.4 s (n = 12, range 4.3 s to 210.1 s). Haiku 4.5 109.0 s (n = 10, range 25.7 s to 223.9 s). Every configuration’s fastest-to-slowest range overlaps every other, so the medians describe this run and are not a tested ranking. Cost is a list-price calculation; the calls ran on subscriptions. The lowest recorded lower bound was GPT-6.1 Sol (medium) · Codex CLI: $0.083 (a lower bound: 3 timed-out calls report no tokens; $0.100 if each had cost a median call, an assumption). Selection effect: the study picked tasks with mixed or failed Sonnet pilot results. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls. The counted calls are new calls. Selection can produce this pattern; the run does not establish its cause.","date":"2026-10-07","updated":"2026-10-07","tags":["head-to-head","harder-tasks","gpt-6-1-sol","codex-cli","claude-sonnet","claude-opus","claude-haiku","reasoning","selection-effect","format-misses","latency"],"caveats":["Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.","Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.","Small samples: GPT-6.1 Sol (medium) n = 16, Opus 5.5 n = 12, Sonnet 5.5 n = 16, Haiku 4.5 n = 12. Per-task cells have 3 to 4 calls. The intervals are wide, and the calls on one task are not independent, so the intervals are likely narrower than the truth.","Each row pairs a CLI with a model on one subscription. The CLIs use different subscriptions, system prompts and start-up steps. A Claude-vs-GPT row compares route + model pairs, not the models alone.","Both CLIs disabled tools. Every prompt asks for the answer only. The runners differ in one way: the Codex runner adds a developer instruction not to call tools, and the Claude Code runner adds none. 11 of 40 Claude Code calls tried a tool anyway. By model: Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12.\n\nNone of them passed. None of the 16 Codex CLI calls did. We did not test the instruction, so we cannot say how much of the gap it explains. With tools on, the Claude Code models might have run code and passed on these tasks. This study does not test that either.","The runner stopped 10 of 56 calls at the 300 s timeout and counted them as no answer. Counts: GPT-6.1 Sol (medium) 3, Opus 5.5 1, Sonnet 5.5 4 and Haiku 4.5 2. A longer limit might turn some into answers, passes or fails.\n\nTiming medians cover completed calls only. They include wrong answers and format misses. They omit only timeouts and tool-call parse errors in this run.","Claude Code ran with an output-token cap setting of 16,000. Reported totals exceeded 16,000 in 11 of 40 Claude Code calls. The setting did not bound reported totals. We cannot tell whether it cut any reply. The tool-call parse errors came at 16,746 and 16,859 reported output tokens. The Codex CLI calls had no such cap.","The kept tasks are mostly long exact-search or computation tasks. In the pilot Claude Sonnet 5.5 passed 10 of 10 code, SQL, spec, numeric and simulation candidates 2 of 2. This set says little about everyday coding.","Claude Haiku 4.5 · Claude Code passed none: a floor on this set, not a general rating.","Strict format rules decide part of the result: a reply with extra text fails. The chart shows the lenient reading next to the strict score.","Arena servers shared the Mac during part of the run. Host load was not controlled. CLI timings include start-up and each CLI’s system prompt. These times do not isolate model speed.","The protocol first said to drop tasks with validator defects. An amendment instead repaired the JavaScript validators and re-scored stored pilot replies. None of those code-writing tasks entered the counted set.","No configuration passed every call across the full set. Some per-task cells hit a ceiling: Sonnet and Opus on the nonogram, and Sol on Skyscrapers. Their perfect cells do not prove equal ability.","Wilson intervals treat calls as independent. The same four tasks repeat, so these are descriptive call-level intervals, not population intervals for coding tasks. The study ran no task-level paired significance test.","Default effort means the effort flag was not passed; the CLI chose. Median reasoning tokens per completed call, where the CLI reports them (n and ranges are in the token chart): GPT-6.1 Sol (medium) 4,971, Opus 5.5 8,352, Sonnet 5.5 6,557, Haiku 4.5 12,483. They are part of the output tokens.","List-price costs are calculations; the calls used flat subscriptions. Opus 5.5 figures are provisional: its cache-read price is under re-check. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. Those cells show a lower bound on cost per pass."],"sourceIds":["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["harder-h2h-pass-all","Counted calls that passed strictly (harder set)",0.3929,"rate","39% (22/56)",56,[0.2758,0.5237],"\u0001"],["harder-h2h-correct-all","Counted calls with a correct answer, format misses included (lenient reading)",0.4107,"rate","41% (23/56)",56,[0.2917,0.5412],"\u0001"],["harder-h2h-pass-sol","GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks",0.6875,"rate","69% (11/16)",16,[0.444,0.8584],"\u0001"],["harder-h2h-pass-opus","Opus 5.5 (Claude Code): strict pass rate on the harder tasks",0.4167,"rate","42% (5/12)",12,[0.1933,0.6805],"\u0001"],["harder-h2h-pass-sonnet","Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks",0.375,"rate","38% (6/16)",16,[0.1848,0.6136],"\u0001"],["harder-h2h-pass-haiku","Haiku 4.5 (Claude Code): strict pass rate on the harder tasks",0,"rate","0% (0/12)",12,[0,0.2425],"\u0001"],["harder-h2h-format-misses","Non-passes that were format misses, not wrong answers",1,"count","1 of 34 non-passes (21 wrong answers, 12 no answer)",34,"\u0001","\u0001"],["harder-h2h-tool-attempts","Counted calls that tried a tool although tools were off (a behaviour, not a quality score)",0.1964,"rate","20% (11/56)",56,[0.1134,0.3184],"Per configuration: GPT-6.1 Sol (medium) 0 of 16, Opus 5.5 5 of 12, Sonnet 5.5 5 of 16 and Haiku 4.5 1 of 12. None of these calls passed. The flag comes from the reply text: tool-call markup, or a CLI message that a tool call could not be parsed. The Codex runner also tells the model not to call tools; the Claude Code runner does not."],["harder-h2h-timeouts","Counted calls that ran past the 300 s limit and gave no answer",0.1786,"rate","18% (10/56)",56,[0.1,0.2984],"Per configuration: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12, Sonnet 5.5 4 of 16 and Haiku 4.5 2 of 12. A timeout counts as a non-pass."],["harder-h2h-pilot-sonnet","Pilot (not scored): strict passes of Claude Sonnet 5.5 on all candidate tasks",0.8125,"rate","81% (26/32)",32,[0.6469,0.9111],"Two calls per candidate. Its rates have a selection effect and repeated-task dependence. The pilot decided which tasks were kept; it is not part of any counted cell."],["harder-h2h-pilot-sonnet-kept","Pilot (not scored): strict passes of Claude Sonnet 5.5 on the tasks that were kept",0.25,"rate","25% (2/8)",8,[0.0715,0.5907],"The same four tasks the counted calls use. The counted Sonnet rate is in harder-h2h-pass-sonnet. Selection against Sonnet predicts a lower pilot rate than a fresh rate."],["harder-h2h-tasks-kept","Candidate tasks kept for the counted set",4,"count","4 of 16 (Sonnet passed 12 candidates 2 of 2, including 10 of 10 code, SQL, spec, numeric and simulation tasks)",16,"\u0001","\u0001"],["harder-h2h-cheapest-per-pass","Lowest recorded cost lower bound per strict pass (calculation)",0.08293,"usd","GPT-6.1 Sol (medium) · Codex CLI: $0.0829 (lower bound)",16,"\u0001","All passing configurations have unpriced timeout calls. Unknown costs can change their order. This is a recorded lower bound, not a cost ranking."]]},"charts":{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","whisker","series","note","sourceIds","xLabel"],"$r":[["harder-h2h-pass-rate","Pass rate on 4 harder tasks","Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","dot-range","rate","higher","Passed","ci95",[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.5,0.2538,0.7462,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}}],"Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row.\n\nCounted calls are new calls.",["agent-harder-tasks"],"\u0001"],["harder-h2h-outcomes","What happened on every call","Counts per configuration: strict passes, format misses, wrong answers and calls with no answer","stacked-bar","count","\u0001","Calls","\u0001",{"$k":["name","points"],"$r":[["Strict pass",{"$k":["label","value","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",11,16],["Claude Opus 5.5 · Claude Code",5,12],["Claude Sonnet 5.5 · Claude Code",6,16],["Claude Haiku 4.5 · Claude Code",0,12]]}],["Format miss (correct answer, wrong format)",{"$k":["label","value","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0,16],["Claude Opus 5.5 · Claude Code",1,12],["Claude Sonnet 5.5 · Claude Code",0,16],["Claude Haiku 4.5 · Claude Code",0,12]]}],["Wrong answer",{"$k":["label","value","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",2,16],["Claude Opus 5.5 · Claude Code",3,12],["Claude Sonnet 5.5 · Claude Code",6,16],["Claude Haiku 4.5 · Claude Code",10,12]]}],["No answer (timeout or error)",{"$k":["label","value","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",3,16],["Claude Opus 5.5 · Claude Code",3,12],["Claude Sonnet 5.5 · Claude Code",4,16],["Claude Haiku 4.5 · Claude Code",2,12]]}]]},"A format miss fails strictly but has an extracted answer that passes the same validator. Extra working or grids can cause this outcome. It is not a pass. A call with no answer is a timeout or an error; it counts as a non-pass.",["agent-harder-tasks"],"\u0001"],["harder-h2h-tool-attempts","Calls that tried a tool although tools were off","Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse","dot-range","rate","none","Calls with a tool attempt","ci95",[{"name":"Tool attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0,0,0.1936,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.3125,0.1416,0.556,16],["Claude Haiku 4.5 · Claude Code",0.0833,0.0149,0.3539,12]]}}],"Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator.\n\nIt is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.",["agent-harder-tasks"],"\u0001"],["harder-h2h-pass-by-task","Strict pass rate by task","One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide","grouped-bar","rate","higher","Strict pass rate","ci95",{"$k":["name","points"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",0.75,0.3006,0.9544,4],["Sudoku, 22 givens",0.25,0.0456,0.6994,4],["6x6 Skyscrapers",1,0.5101,1,4],["Seeded shuffle output",0.75,0.3006,0.9544,4]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",1,0.4385,1,3],["Sudoku, 22 givens",0,0,0.5615,3],["6x6 Skyscrapers",0.3333,0.0615,0.7923,3],["Seeded shuffle output",0.3333,0.0615,0.7923,3]]}],["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",1,0.5101,1,4],["Sudoku, 22 givens",0,0,0.4899,4],["6x6 Skyscrapers",0,0,0.4899,4],["Seeded shuffle output",0.5,0.15,0.85,4]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10x10 nonogram",0,0,0.5615,3],["Sudoku, 22 givens",0,0,0.5615,3],["6x6 Skyscrapers",0,0,0.5615,3],["Seeded shuffle output",0,0,0.5615,3]]}]]},"Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.",["agent-harder-tasks"],"\u0001"],["harder-h2h-total-latency","Total time per call on harder tasks","Median per configuration; whiskers = fastest and slowest call","dot-range","seconds","\u0001","Seconds","minmax",[{"name":"Total time per call","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",120.24,46.24,273.46,13,false],["Claude Opus 5.5 · Claude Code",80.34,3.82,279.5,9,false],["Claude Sonnet 5.5 · Claude Code",70.43,4.32,210.08,12,false],["Claude Haiku 4.5 · Claude Code",108.98,25.73,223.95,10,false]]}}],"Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session.\n\nArena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.",["agent-harder-tasks"],"\u0001"],["harder-h2h-output-tokens","Output tokens per call on harder tasks","Median per configuration; reasoning tokens as the CLI reports them","grouped-bar","tokens","none","Tokens","minmax",[{"name":"Output tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",4994,2099,13413,13],["Claude Opus 5.5 · Claude Code",8420,279,40044,9],["Claude Sonnet 5.5 · Claude Code",9287,407,27921,12],["Claude Haiku 4.5 · Claude Code",12508,2965,26532,10]]}},{"name":"Reasoning tokens","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",4971,2070,13372,13],["Claude Opus 5.5 · Claude Code",8352,21,9897,9],["Claude Sonnet 5.5 · Claude Code",6557,63,27902,12],["Claude Haiku 4.5 · Claude Code",12483,2924,26510,10]]}}],"Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured.\n\nClaude Code used an output-token cap setting of 16,000. Some reported totals exceeded it.\n\nCodex CLI had no cap. More tokens is not better or worse by itself.",["agent-harder-tasks"],"\u0001"],["harder-h2h-cost-per-pass","List-price cost per strict pass on harder tasks (calculation)","All calls in a configuration, failures and format misses included, divided by its strict passes","bar","usd","\u0001","USD per strict pass","\u0001",[{"name":"Cost per strict pass","points":{"$k":["label","value","n","highlight"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.08293,16,true],["Claude Sonnet 5.5 · Claude Code",0.23843,16,false],["Claude Opus 5.5 · Claude Code",0.59333,12,false]]}}],"Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. These cells show a lower bound.\n\nAssume each unpriced call cost its cell’s median priced call. This sensitivity calculation gives GPT-6.1 Sol (medium) $0.100, Opus 5.5 $0.633 and Sonnet 5.5 $0.303. Opus 5.5 figures are provisional: its cache-read price is under re-check.\n\nHighlights mark the observed frontier of these lower-bound costs. Unknown timeout costs can change it; this is not a cost ranking. Claude Haiku 4.5 · Claude Code had no strict pass, so it has no cost per pass.",["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"],"\u0001"],["harder-h2h-frontier","Observed quality vs cost frontier (calculation)","Strict pass rate against list-price cost per strict pass","scatter","rate","higher","Strict pass rate","ci95",[{"name":"Codex CLI","points":[{"label":"GPT-6.1 Sol (medium) · Codex CLI","value":0.6875,"lo":0.444,"hi":0.8584,"n":16,"x":0.08293,"highlight":true}]},{"name":"Claude Code","points":[{"label":"Claude Opus 5.5 · Claude Code","value":0.4167,"lo":0.1933,"hi":0.6805,"n":12,"x":0.59333,"highlight":false},{"label":"Claude Sonnet 5.5 · Claude Code","value":0.375,"lo":0.1848,"hi":0.6136,"n":16,"x":0.23843,"highlight":false}]}],"Upper-left has a higher observed pass rate and lower recorded cost per pass. Highlights mark the observed frontier of lower-bound costs. Unknown timeout costs can change it. This is not a tested ranking. Frontier: GPT-6.1 Sol (medium) · Codex CLI.\n\nCosts are calculations from tokens; calls that timed out are not priced (GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16), so those cells are lower bounds. The 95% Wilson intervals are listed below; the plot shows point estimates.\n\nStrict rates: GPT-6.1 Sol (medium) 11/16 (n = 16, 95% interval 44.4% to 85.8%); Opus 5.5 5/12 (n = 12, 95% interval 19.3% to 68.0%); Sonnet 5.5 6/16 (n = 16, 95% interval 18.5% to 61.4%).",["agent-harder-tasks","calc-repricing","price-anthropic","price-openai"],"USD per strict pass (list-price calculation)"]]},"related":["hard-model-head-to-head","effort-ladder"]}}