Model · OpenAI
GPT-6.1 Sol (OpenAI API)
OpenAI’s GPT-6.1 Sol model called directly through the OpenAI API, without a CLI.
13 values from 1 study · Updated
At a glance
The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.
Qualitypass rates, accuracy and scores
Not measured in any study yet.
Speedtime per call or decision
1.52s
range 1.4 s–2.2 s · n = 5
CLI vs API: time for a one-line answer (Total time)
OpenAI API · effort high · fixed exact reply, 5 runs · Claude Code CLI vs Codex CLI vs the API: latency and tokens
CostUS dollars per call, pass or decision
Not measured in any study yet.
GPT-6.1 Sol price per 1M tokens
The 2026-10-03 vendor-price record lists GPT-6.1 Sol at $2.00 per 1M input tokens and $10.00 per 1M output tokens. Cache reads are recorded at $0.10 per 1M tokens.
Recorded vendor list price
$2.00
Input, per 1M tokens
Recorded vendor list price
$10.00
Output, per 1M tokens
Recorded vendor list price
$0.10
Cache read, per 1M tokens
No current vendor context limit is proved for this exact model here.
Recorded study price source: OpenAI list prices (). Token prices as listed by the vendor on 2026-10-03.
No provider prices for this model name in the OpenRouter snapshot.
Estimate your monthly costPrices of every model at every provider
Where it sits
Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| CLI vs API: time for a one-line answer (Total time) | 1.02 s | 5 | 1 s–1.9 s (range) | OpenAI API · effort low · fixed exact reply, 5 runs |
| CLI vs API: time for a one-line answer (Total time) | 1.52 s | 5 | 1.4 s–2.2 s (range) | OpenAI API · effort high · fixed exact reply, 5 runs |
| CLI vs API: time for a one-line answer (First useful output) | 0.87 s | 5 | 0.8 s–1.7 s (range) | OpenAI API · effort low · fixed exact reply, 5 runs |
| CLI vs API: time for a one-line answer (First useful output) | 1.34 s | 5 | 1.3 s–2.1 s (range) | OpenAI API · effort high · fixed exact reply, 5 runs |
| CLI vs API: time for a small coding task (Total time) | 6.00 s | 3 | 5.4 s–6.2 s (range) | OpenAI API · effort low · small coding task, 3 runs |
| CLI vs API: time for a small coding task (Total time) | 9.56 s | 3 | 9.4 s–10.9 s (range) | OpenAI API · effort high · small coding task, 3 runs |
| CLI vs API: time for a small coding task (First useful output) | 1.05 s | 3 | 1 s–1.4 s (range) | OpenAI API · effort low · small coding task, 3 runs |
| CLI vs API: time for a small coding task (First useful output) | 5.31 s | 3 | 5 s–6.4 s (range) | OpenAI API · effort high · small coding task, 3 runs |
| Hidden prompt: input tokens for the same one-line request | 17 | 5 | — | OpenAI API · effort low · short fixed tasks |
| Hidden prompt: input tokens for the same one-line request | 17 | 5 | — | OpenAI API · effort high · short fixed tasks |
| Repairing a scheduler: Claude Code vs Codex vs API (Total time) | 17.3 s | 3 | 16.3 s–18.6 s (range) | OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs |
| Repairing a scheduler: Claude Code vs Codex vs API (First useful output) | 7.46 s | 3 | 6.7 s–9.1 s (range) | OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs |
| Output tokens to repair the scheduler (Output tokens) | 1,313 | 3 | — | OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs |
Lines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chart
GPT-6.1 Sol (OpenAI API) in Claude Code CLI vs Codex CLI vs the API: latency and tokens: 13 values, first CLI vs API: time for a one-line answer (Total time) 1.02 s.
Compare GPT-6.1 Sol (OpenAI API)
Each bar counts the rows of one comparison: a side ahead only where its interval or range is apart, otherwise a tie or unclear.
vs model
GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API)
GPT-6.1 Sol (OpenAI API) ahead on 2 · 6 unclear
vs model
GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI)
GPT-6.1 Sol (OpenAI API) ahead on 2 · 3 unclear
vs model
GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (OpenAI API)
1 tie · 4 unclear
vs model
Claude Sonnet 5.5 vs GPT-6.1 Sol (OpenAI API)
3 unclear
Watch
Claude Code CLI vs Codex CLI vs the API: a latency race
For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.
Transcript
- Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
- A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
- It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
- A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
- Open benchmarks: intervals, sources and every failure kept.
Write-ups that use this data
All postsA latency budget for voice agents: which LLM steps fit in one turn?
136.5 ms for Jev, 0.82 s for a small-model API, 2.79 s to 3.79 s for Codex CLI: which steps fit a voice agent latency budget? A thought experiment.
A voice agent latency budget, with measured times: what fits in one turn?
Rules and Jev 1.13 fit every budget we assumed; a Claude router through a CLI fits none. 14 measured steps vs 300, 800 and 1,500 ms. A thought experiment.
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.