{"i":59,"comparison":{"slug":"gpt-6-1-sol-codex-cli-vs-gpt-6-luna-openai-api","a":"gpt-6-1-sol-codex-cli","b":"gpt-6-luna-openai-api","title":"GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (OpenAI API)","seoTitle":"GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (OpenAI API)","description":"GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (OpenAI API): 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","verdict":"GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (OpenAI API) share 5 measured metrics from one study. GPT-6 Luna (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 0.97 s vs 4.19 s; CLI vs API: time for a one-line answer (First useful output), 0.82 s vs 3.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example Codex CLI vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",4.19,0.97,"seconds","4.19 s","0.97 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6 Luna (OpenAI API) 0.65 s to 1.50 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[3.81,4.69],[0.65,1.5]],["CLI vs API: time for a one-line answer (First useful output)",3.79,0.82,"seconds","3.79 s","0.82 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6 Luna (OpenAI API) 0.51 s to 1.37 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[3.37,4.3],[0.51,1.37]],["CLI vs API: time for a small coding task (Total time)",17.85,4.01,"seconds","17.9 s","4.01 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6 Luna (OpenAI API) 3.83 s to 4.35 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[17.68,22.42],[3.83,4.35]],["CLI vs API: time for a small coding task (First useful output)",17.27,0.67,"seconds","17.3 s","0.67 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6 Luna (OpenAI API) 0.62 s to 0.81 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[17.13,21.86],[0.62,0.81]],["Hidden prompt: input tokens for the same one-line request",19555,17,"tokens","19,555","17","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort high · short fixed tasks","OpenAI API · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}}}