{"i":9,"comparison":{"slug":"gpt-6-1-sol-codex-cli-vs-gpt-6-1-sol-openai-api","a":"gpt-6-1-sol-codex-cli","b":"gpt-6-1-sol-openai-api","title":"GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API)","seoTitle":"GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API)","description":"GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API): 8 measured metrics from one study, with sample sizes, intervals and every failure counted.","verdict":"GPT-6.1 Sol (Codex CLI) and GPT-6.1 Sol (OpenAI API) share 8 measured metrics from one study. GPT-6.1 Sol (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 1.52 s vs 4.19 s; CLI vs API: time for a one-line answer (First useful output), 1.34 s vs 3.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 6 unclear; each row says why. Every row ran the two sides through different routes (for example Codex CLI vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",4.19,1.52,"seconds","4.19 s","1.52 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","OpenAI API · effort high · fixed exact reply, 5 runs","range","minmax",[3.81,4.69],[1.35,2.23]],["CLI vs API: time for a one-line answer (First useful output)",3.79,1.34,"seconds","3.79 s","1.34 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","OpenAI API · effort high · fixed exact reply, 5 runs","range","minmax",[3.37,4.3],[1.26,2.12]],["CLI vs API: time for a small coding task (Total time)",17.85,9.56,"seconds","17.9 s","9.56 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6.1 Sol (OpenAI API) 9.44 s to 10.9 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","OpenAI API · effort high · small coding task, 3 runs","range","minmax",[17.68,22.42],[9.44,10.94]],["CLI vs API: time for a small coding task (First useful output)",17.27,5.31,"seconds","17.3 s","5.31 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6.1 Sol (OpenAI API) 4.99 s to 6.42 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","OpenAI API · effort high · small coding task, 3 runs","range","minmax",[17.13,21.86],[4.99,6.42]],["Hidden prompt: input tokens for the same one-line request",19555,17,"tokens","19,555","17","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort high · short fixed tasks","OpenAI API · effort high · short fixed tasks","\u0001","\u0001","\u0001","\u0001"],["Repairing a scheduler: Claude Code vs Codex vs API (Total time)",61.16,17.32,"seconds","61.2 s","17.3 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 59.9 s to 69.5 s; GPT-6.1 Sol (OpenAI API) 16.3 s to 18.6 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs","OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[59.9,69.51],[16.28,18.61]],["Repairing a scheduler: Claude Code vs Codex vs API (First useful output)",15.56,7.46,"seconds","15.6 s","7.46 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 13.7 s to 23.0 s; GPT-6.1 Sol (OpenAI API) 6.68 s to 9.05 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"scheduler-repair-claude-vs-codex",3,3,"Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs","OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs","range","minmax",[13.65,23.04],[6.68,9.05]],["Output tokens to repair the scheduler (Output tokens)",1181,1313,"tokens","1,181","1,313","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",3,"scheduler-repair-output-tokens",3,3,"Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs","OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs","\u0001","\u0001","\u0001","\u0001"]]}}}