{"i":57,"comparison":{"slug":"gpt-6-1-sol-codex-cli-vs-gpt-6-luna-codex-cli","a":"gpt-6-1-sol-codex-cli","b":"gpt-6-luna-codex-cli","title":"GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (Codex CLI)","seoTitle":"GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (Codex CLI)","description":"GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (Codex CLI): 6 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","verdict":"GPT-6.1 Sol (Codex CLI) and GPT-6 Luna (Codex CLI) share 6 measured metrics and 2 list-price calculations from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 8 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["CLI vs API: time for a one-line answer (Total time)",4.19,3.19,"seconds","4.19 s","3.19 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.81 s to 4.69 s; GPT-6 Luna (Codex CLI) 2.88 s to 3.83 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","Codex CLI · effort none · fixed exact reply, 5 runs","range","minmax",[3.81,4.69],[2.88,3.83],"\u0001"],["CLI vs API: time for a one-line answer (First useful output)",3.79,2.79,"seconds","3.79 s","2.79 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 3.37 s to 4.30 s; GPT-6 Luna (Codex CLI) 2.46 s to 3.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"Codex CLI · effort high · fixed exact reply, 5 runs","Codex CLI · effort none · fixed exact reply, 5 runs","range","minmax",[3.37,4.3],[2.46,3.42],"\u0001"],["CLI vs API: time for a small coding task (Total time)",17.85,9.23,"seconds","17.9 s","9.23 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.7 s to 22.4 s; GPT-6 Luna (Codex CLI) 8.99 s to 11.7 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","Codex CLI · effort none · small coding task, 3 runs","range","minmax",[17.68,22.42],[8.99,11.68],"\u0001"],["CLI vs API: time for a small coding task (First useful output)",17.27,8.68,"seconds","17.3 s","8.68 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (Codex CLI) 17.1 s to 21.9 s; GPT-6 Luna (Codex CLI) 8.27 s to 11.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"Codex CLI · effort high · small coding task, 3 runs","Codex CLI · effort none · small coding task, 3 runs","range","minmax",[17.13,21.86],[8.27,11.01],"\u0001"],["Hidden prompt: input tokens for the same one-line request",19555,18859,"tokens","19,555","18,859","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"Codex CLI · effort high · short fixed tasks","Codex CLI · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001","\u0001"],["Time to first text: a 250-line answer, six models",3.52,3.3,"seconds","3.52 s","3.30 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","llm-speed-anatomy",4,"speed-anatomy-first-text",4,4,"Codex CLI · effort low","Codex CLI · effort low","range","minmax",[2.75,4.42],[3.19,3.47],"\u0001"],["Output speed after the first text: visible tokens per second (calculation)",79.6,129.1,"tokens","80","129","unclear","More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-output-speed",4,4,"Codex CLI · effort low","Codex CLI · effort low","range","minmax",[71.6,80.5],[55.5,259.1],true],["Output speed in characters per second after the first text (calculation)",323,524,"count","323","524","unclear","More or fewer count is not better or worse by itself; this row describes behaviour, not a winner.","llm-speed-anatomy",4,"speed-anatomy-chars-per-second",4,4,"Codex CLI · effort low","Codex CLI · effort low","range","minmax",[291,327],[225,1052],true]]}}}