{"i":60,"comparison":{"slug":"gpt-6-1-sol-openai-api-vs-gpt-6-luna-codex-cli","a":"gpt-6-1-sol-openai-api","b":"gpt-6-luna-codex-cli","title":"GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI)","seoTitle":"GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI)","description":"GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI): 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","verdict":"GPT-6.1 Sol (OpenAI API) and GPT-6 Luna (Codex CLI) share 5 measured metrics from one study. GPT-6.1 Sol (OpenAI API) leads on 2 rows: CLI vs API: time for a one-line answer (Total time), 1.52 s vs 3.19 s; CLI vs API: time for a one-line answer (First useful output), 1.34 s vs 2.79 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example OpenAI API vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",1.52,3.19,"seconds","1.52 s","3.19 s","a","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s; GPT-6 Luna (Codex CLI) 2.88 s to 3.83 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort high · fixed exact reply, 5 runs","Codex CLI · effort none · fixed exact reply, 5 runs","range","minmax",[1.35,2.23],[2.88,3.83]],["CLI vs API: time for a one-line answer (First useful output)",1.34,2.79,"seconds","1.34 s","2.79 s","a","The run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s; GPT-6 Luna (Codex CLI) 2.46 s to 3.42 s). A range is not a confidence interval. Samples are small (5 runs per side).","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort high · fixed exact reply, 5 runs","Codex CLI · effort none · fixed exact reply, 5 runs","range","minmax",[1.26,2.12],[2.46,3.42]],["CLI vs API: time for a small coding task (Total time)",9.56,9.23,"seconds","9.56 s","9.23 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) 9.44 s to 10.9 s; GPT-6 Luna (Codex CLI) 8.99 s to 11.7 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort high · small coding task, 3 runs","Codex CLI · effort none · small coding task, 3 runs","range","minmax",[9.44,10.94],[8.99,11.68]],["CLI vs API: time for a small coding task (First useful output)",5.31,8.68,"seconds","5.31 s","8.68 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 4.99 s to 6.42 s; GPT-6 Luna (Codex CLI) 8.27 s to 11.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort high · small coding task, 3 runs","Codex CLI · effort none · small coding task, 3 runs","range","minmax",[4.99,6.42],[8.27,11.01]],["Hidden prompt: input tokens for the same one-line request",17,18859,"tokens","17","18,859","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"OpenAI API · effort high · short fixed tasks","Codex CLI · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}}}