{"i":61,"comparison":{"slug":"gpt-6-1-sol-openai-api-vs-gpt-6-luna-openai-api","a":"gpt-6-1-sol-openai-api","b":"gpt-6-luna-openai-api","title":"GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (OpenAI API)","seoTitle":"GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (OpenAI API)","description":"GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (OpenAI API): 5 measured metrics from one study, with sample sizes, intervals and every failure counted.","verdict":"GPT-6.1 Sol (OpenAI API) and GPT-6 Luna (OpenAI API) share 5 measured metrics from one study. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 4 unclear; each row says why. Some rows rest on small samples (n = 3 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["CLI vs API: time for a one-line answer (Total time)",1.52,0.97,"seconds","1.52 s","0.97 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) 1.35 s to 2.23 s; GPT-6 Luna (OpenAI API) 0.65 s to 1.50 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort high · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[1.35,2.23],[0.65,1.5]],["CLI vs API: time for a one-line answer (First useful output)",1.34,0.82,"seconds","1.34 s","0.82 s","unclear","The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (OpenAI API) 1.26 s to 2.12 s; GPT-6 Luna (OpenAI API) 0.51 s to 1.37 s); the medians alone do not show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",5,"cli-vs-api-exact-reply-latency",5,5,"OpenAI API · effort high · fixed exact reply, 5 runs","OpenAI API · effort none · fixed exact reply, 5 runs","range","minmax",[1.26,2.12],[0.51,1.37]],["CLI vs API: time for a small coding task (Total time)",9.56,4.01,"seconds","9.56 s","4.01 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 9.44 s to 10.9 s; GPT-6 Luna (OpenAI API) 3.83 s to 4.35 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort high · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[9.44,10.94],[3.83,4.35]],["CLI vs API: time for a small coding task (First useful output)",5.31,0.67,"seconds","5.31 s","0.67 s","unclear","Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (GPT-6.1 Sol (OpenAI API) 4.99 s to 6.42 s; GPT-6 Luna (OpenAI API) 0.62 s to 0.81 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.","cli-model-latency-tokens",3,"cli-vs-api-small-coding-latency",3,3,"OpenAI API · effort high · small coding task, 3 runs","OpenAI API · effort none · small coding task, 3 runs","range","minmax",[4.99,6.42],[0.62,0.81]],["Hidden prompt: input tokens for the same one-line request",17,17,"tokens","17","17","tie","Same value. More or fewer is not better by itself for this metric.","cli-model-latency-tokens",5,"cli-vs-api-prompt-overhead",5,5,"OpenAI API · effort high · short fixed tasks","OpenAI API · effort none · short fixed tasks","\u0001","\u0001","\u0001","\u0001"]]}}}