Model · OpenAI

GPT-6 Luna (Codex CLI)

OpenAI’s GPT-6 Luna model run through the Codex CLI.

34 values from 3 studies (4 are list-price calculations) · Updated

At a glance

The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.

Qualitypass rates, accuracy and scores

63% (10/16)

95% CI 39%–82% · n = 16

Strict pass rate: single call vs agent loop on eight hard tasks

Codex CLI · single call · Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks

Speedtime per call or decision

5.16s

range 3.6 s–11.3 s · n = 16

Total time per attempt: single call vs agent loop

Codex CLI · single call · Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks

CostUS dollars per call, pass or decision

$0.0012

n = 16 · list-price calculation

List-price cost per strict pass: single call vs agent loop (calculation)

Codex CLI · single call · Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks

GPT-6 Luna price per 1M tokens

The 2026-10-06 OpenRouter snapshot reported 2 standard-tier providers at $0.10 per 1M input tokens and $0.50 per 1M output tokens. These are third-party-reported prices.

Third-party-reported · snapshot 2026-10-06

Third-party-reported

$0.10

Input, per 1M tokens

Third-party-reported

$0.50

Output, per 1M tokens

Third-party-reported

$0.01

Cache read, per 1M tokens

No current vendor context limit is proved for this exact model here.

No vendor list price for this model name in the price table, so the tiles show the provider snapshot.

Providers in the dated snapshot

2 standard-tier providers · snapshot 2026-10-06 · USD per 1M tokens

Same snapshot price at all 2 providers
  • OpenAIfirst-party
  • Azure

Third-party-reported, not measured by Agent · Standard tier only: regional, flex, fast and priority tiers are left out · One row per provider: its cheapest standard endpoint

GPT-6 Luna: 2 standard-tier providers, all at the same price ($0.10 input, $0.50 output per 1M tokens). Third-party-reported, snapshot 2026-10-06.

Source: OpenRouter public API: models and provider endpoints (snapshot) (). The diamond marks the first-party provider in this snapshot. Current availability is not checked.

Where it sits

Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.

CLI vs API: time for a one-line answer (Total time)n = 5 · range 2.9 s–3.8 s · Codex CLI · effort none · fixed exact reply, 5 runs
3.19 s
CLI vs API: time for a one-line answer (First useful output)n = 5 · range 2.5 s–3.4 s · Codex CLI · effort none · fixed exact reply, 5 runs
2.79 s
CLI vs API: time for a small coding task (Total time)n = 3 · range 9 s–11.7 s · Codex CLI · effort none · small coding task, 3 runs
9.23 s
CLI vs API: time for a small coding task (First useful output)n = 3 · range 8.3 s–11 s · Codex CLI · effort none · small coding task, 3 runs
8.68 s
Hidden prompt: input tokens for the same one-line requestn = 5 · Codex CLI · effort none · short fixed tasks
18,859

Lines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chart

GPT-6 Luna (Codex CLI) in Claude Code CLI vs Codex CLI vs the API: latency and tokens: 5 values, first CLI vs API: time for a one-line answer (Total time) 3.19 s.

Strict pass rate: single call vs agent loop on eight hard tasksn = 16 · 95% CI 39%–82% · Codex CLI · single call
63% (10/16)
Strict pass rate: single call vs agent loop on eight hard tasksn = 14 · 95% CI 60%–96% · Codex CLI · agent loop
86% (12/14)
Strict passes per task: single call vs agent loop: Interval merge fixn = 2 · 95% CI 34%–100% · Codex CLI · single call
100% (2/2)
Strict passes per task: single call vs agent loop: DST day-length fixn = 2 · 95% CI 34%–100% · Codex CLI · single call
100% (2/2)
Strict passes per task: single call vs agent loop: CSV parsern = 2 · 95% CI 34%–100% · Codex CLI · single call
100% (2/2)
Strict passes per task: single call vs agent loop: Event-loop ordern = 2 · 95% CI 0%–66% · Codex CLI · single call
0% (0/2)

Whiskers: 95% Wilson intervalLines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation

GPT-6 Luna (Codex CLI) in Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks: 26 values, first Strict pass rate: single call vs agent loop on eight hard tasks 63% (10/16).

Time to first text: a 250-line answer, six modelsn = 4 · range 3.2 s–3.5 s · Codex CLI · effort low
3.30 s
Output speed after the first text: visible tokens per second (calculation)n = 4 · range 56–259 · Codex CLI · effort low
129Calculation
Output speed in characters per second after the first text (calculation)n = 4 · range 225–1,052 · Codex CLI · effort low
524Calculation

Lines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation

GPT-6 Luna (Codex CLI) in Where the seconds go: first text, output speed and prompt size for 6 LLMs: 3 values, first Time to first text: a 250-line answer, six models 3.30 s.

Compare GPT-6 Luna (Codex CLI)

Each bar counts the rows of one comparison: a side ahead only where its interval or range is apart, otherwise a tie or unclear.

Watch

Live story · 33 sClaude Code CLI vs Codex CLI vs the API: a latency race

Claude Code CLI vs Codex CLI vs the API: a latency race

For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.

Transcript
  1. Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
  2. A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
  3. It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  4. A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
  5. Open benchmarks: intervals, sources and every failure kept.

Write-ups that use this data

All posts

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.