{"i":62,"comparison":{"slug":"gpt-6-luna-codex-cli-vs-gpt-6-luna-openai-api","a":"gpt-6-luna-codex-cli","b":"gpt-6-luna-openai-api","title":"GPT-6 Luna (Codex CLI) vs GPT-6 Luna (OpenAI API)","seoTitle":"GPT-6 Luna (Codex CLI) vs GPT-6 Luna (OpenAI API)","description":"GPT-6 Luna (Codex CLI) vs GPT-6 Luna (OpenAI API): 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","verdict":"GPT-6 Luna (Codex CLI) and GPT-6 Luna (OpenAI API) share 5 measured metrics from one study. GPT-6 Luna (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 0.97 s vs 3.19 s; CLI vs API: time for a one-line answer (First useful output), 0.82 s vs 2.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example Codex CLI vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",3.19,0.97,"seconds","3.19 s","0.97 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6 Luna (Codex CLI) 2.88 s to 3.83 s; GPT-6 Luna (OpenAI API) 0.65 s to 1.50 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort none · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[2.88,3.83],[0.65,1.5]],["CLI vs API: time for a one-line answer (First useful output)",2.79,0.82,"seconds","2.79 s","0.82 s","b","The run ranges (fastest to slowest) do not overlap (GPT-6 Luna (Codex CLI) 2.46 s to 3.42 s; GPT-6 Luna (OpenAI API) 0.51 s to 1.37 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort none · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[2.46,3.42],[0.51,1.37]],["CLI vs API: time for a small coding task (Total time)",9.23,4.01,"seconds","9.23 s","4.01 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6 Luna (Codex CLI) 8.99 s to 11.7 s; GPT-6 Luna (OpenAI API) 3.83 s to 4.35 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort none · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[8.99,11.68],[3.83,4.35]],["CLI vs API: time for a small coding task (First useful output)",8.68,0.67,"seconds","8.68 s","0.67 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6 Luna (Codex CLI) 8.27 s to 11.0 s; GPT-6 Luna (OpenAI API) 0.62 s to 0.81 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort none · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[8.27,11.01],[0.62,0.81]],["Hidden prompt: input tokens for the same one-line request",18859,17,"tokens","18,859","17","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort none · short fixed tasks","OpenAI API · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}}}