{"i":12,"comparison":{"slug":"claude-fable-5-1-vs-gpt-6-1-sol-codex-cli","a":"claude-fable-5-1","b":"gpt-6-1-sol-codex-cli","title":"Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI)","seoTitle":"Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI): benchmarks","description":"Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI): 12 measured metrics from 3 studies (Pass rate on five validated tasks; more), with sample sizes and intervals.","verdict":"Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 4 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Pass rate on five validated tasks",1,1,"rate","100% (15/15)","100% (15/15)","tie","The 95% intervals overlap (Claude Fable 5.1 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.","model-head-to-head",15,"h2h-pass-rate",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","ci95","ci95",[0.7961,1],[0.7961,1],"\u0001"],["Total time per call",1.94,5.65,"seconds","1.94 s","5.65 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 1.41 s to 9.83 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-total-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[1.41,9.83],[4.1,25.46],"\u0001"],["Time to first useful output",1.2,5.05,"seconds","1.20 s","5.05 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 0.95 s to 7.90 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-first-useful-latency",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.95,7.9],[3.36,17.82],"\u0001"],["Input tokens per call: what the CLI sends (Cache read)",2760,5180,"tokens","2,760","5,180","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Input tokens per call: what the CLI sends (Other input)",473,6943,"tokens","473","6,943","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-input-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Output tokens per call (Output tokens)",64,42,"tokens","64","42","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","model-head-to-head",15,"h2h-output-tokens",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per call (calculation)",0.00987,0.01018,"usd","$0.0099","$0.010","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 $0.0049 to $0.058; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.","model-head-to-head",15,"h2h-list-price-per-call",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","range","minmax",[0.0049,0.05843],[0.0054,0.02686],true],["List-price cost per passing answer (calculation)",0.02054,0.01564,"usd","$0.021","$0.016","unclear","No interval or range was recorded for either side, so the gap ($0.021 vs $0.016) is not tested against run-to-run variation.","model-head-to-head",15,"h2h-cost-per-pass",15,15,"Claude Code · five short validated tasks","Codex CLI · effort medium · five short validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Pass rate on eight hard tasks (Strict pass)",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Pass rate on eight hard tasks (Lenient (format misses counted))",1,1,"rate","100% (24/24)","100% (16/16)","tie","The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.","hard-model-head-to-head","\u0001","hard-h2h-pass-rate",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","ci95","ci95",[0.862,1],[0.8064,1],"\u0001"],["Total time per call on hard tasks (separate batches)",16.13,13.11,"seconds","16.1 s","13.1 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 4.46 s to 90.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-total-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[4.46,90],[8.54,61.6],"\u0001"],["Time to first useful output on hard tasks",11.63,10.23,"seconds","11.6 s","10.2 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.00 s to 85.3 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","hard-model-head-to-head","\u0001","hard-h2h-first-useful-latency",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","range","minmax",[2,85.33],[6.09,40.41],"\u0001"],["Output tokens per call on hard tasks (Output tokens)",1366,335,"tokens","1,366","335","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","hard-model-head-to-head","\u0001","hard-h2h-output-tokens",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["List-price cost per strict pass on hard tasks (calculation)",0.09331,0.02564,"usd","$0.093","$0.026","unclear","No interval or range was recorded for either side, so the gap ($0.093 vs $0.026, 3.6x) is not tested against run-to-run variation.","hard-model-head-to-head","\u0001","hard-h2h-cost-per-pass",24,16,"Claude Code · eight hard validated tasks","Codex CLI · effort medium · eight hard validated tasks","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share of output tokens per call on hard tasks (calculation)",64.24,46.33,"percent","64.2%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-share",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[23.44,97.19],[11.42,86.85],true],["List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))",0.053696,0.002273,"usd","$0.054","$0.0023","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))",0.018702,0.003025,"usd","$0.019","$0.0030","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))",0.02091,0.020339,"usd","$0.021","$0.020","unclear","More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-cost-per-call",24,16,"Claude Code","Codex CLI · effort medium","\u0001","\u0001","\u0001","\u0001",true],["Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)",64.24,46.33,"percent","64.2%","46.3%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill","\u0001","thinking-bill-short-vs-hard",24,16,"Claude Code","Codex CLI · effort medium","range","minmax",[23.44,97.19],[11.42,86.85],true],["Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)",0,41.05,"percent","0%","41%","unclear","More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.","thinking-token-bill",15,"thinking-bill-short-vs-hard",15,15,"Claude Code","Codex CLI · effort medium","range","minmax",[0,74.01],[0,71.43],true],["Time to first text: a 250-line answer, six models",4.43,3.52,"seconds","4.43 s","3.52 s","unclear","The run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.27 s to 4.64 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[2.27,4.64],[2.75,4.42],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",122.6,79.6,"tokens","123","80","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[120.9,131.4],[71.6,80.5],true],["Output speed in characters per second after the first text (calculation)",273,323,"count","273","323","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Claude Code","Codex CLI · effort low","range","minmax",[270,293],[291,327],true]]}}}