3 measured metrics · 2 studies
Gemini 3 FlashvsGPT 5 mini
No row separates them: 1 tie, 2 unclear.
The verdict
Gemini 3 Flash and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Gemini 3 Flash | GPT 5 mini | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 82% (27/33)effort high · public mini-SWE-agent v2 run, same instances | 64% (21/33)public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 66%–91% vs 47%–78% | Tie | The 95% intervals overlap (Gemini 3 Flash 66% to 91%; GPT 5 mini 47% to 78%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.44effort high · public mini-SWE-agent v2 run, same instances | $0.080public mini-SWE-agent v2 run, same instances | 27 / 21 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.44 vs $0.080, 5.5x) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
| Model calls per instance | 54.2effort high · public mini-SWE-agent v2 run, same instances | 20.8public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
Marks: 95% intervals (Wilson for rates)n is shown per side on every row
3 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Recorded cost per resolved instance: Agent vs the public panel, 5.5x (Gemini 3 Flash larger).
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Gemini 3 Flash
- GPT 5 mini
- 95% interval
- where the two overlap
These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.
Agent on SWE-bench Verified vs 11 public models
- Resolved rate on the same 33 SWE-bench Verified instances82% (27/33)n 3364% (21/33)n 33TieResolved rate on the same 33 SWE-bench Verified instances: Gemini 3 Flash 82% (27/33) (n 33, 95% interval 66%–91%); GPT 5 mini 64% (21/33) (n 33, 95% interval 47%–78%). Tie.
Calculation: at these rates, about 90 runs per side would separate them.
- Model calls per instance54.2n 3320.8n 33UnclearModel calls per instance: Gemini 3 Flash 54.2 (n 33); GPT 5 mini 20.8 (n 33). Unclear.
What if every call ran on Opus? Repricing real agent tokens
- Recorded cost per resolved instance: Agent vs the public panel$0.44n 27$0.080n 21UnclearRecorded cost per resolved instance: Agent vs the public panel: Gemini 3 Flash $0.44 (n 27); GPT 5 mini $0.080 (n 21). Unclear.
| Metric | Gemini 3 Flash | GPT 5 mini | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 82% (27/33)effort high · public mini-SWE-agent v2 run, same instances | 64% (21/33)public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 66%–91% vs 47%–78% | Tie | The 95% intervals overlap (Gemini 3 Flash 66% to 91%; GPT 5 mini 47% to 78%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 54.2effort high · public mini-SWE-agent v2 run, same instances | 20.8public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.44effort high · public mini-SWE-agent v2 run, same instances | $0.080public mini-SWE-agent v2 run, same instances | 27 / 21 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.44 vs $0.080, 5.5x) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
Marks: 95% intervals (Wilson for rates)n is shown per side on every row
3 rows from 2 studies. No row separates them: 1 tie, 2 unclear.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Gemini 3 Flash
No row in this data puts Gemini 3 Flash ahead of GPT 5 mini. Pick on other grounds (price, access, the tasks you run), or measure your own workload.
When to pick GPT 5 mini
No row in this data puts GPT 5 mini ahead of Gemini 3 Flash. Pick on other grounds (price, access, the tasks you run), or measure your own workload.
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Resolved rate | 95% interval | n |
|---|---|---|---|
| Gemini 3 Flash (high) | 82% | 66%–91% | 33 |
| GPT 5 mini | 64% | 47%–78% | 33 |
2 rows. Highest Gemini 3 Flash (high) 82% (95% interval 66%–91%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 33 per row
Agent vs 11 public mini-SWE-agent v2 runs, one attempt each
Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
Model calls per instance
Mean over the same 33 instances
| Item | Mean calls | n |
|---|---|---|
| Gemini 3 Flash (high) | 54.2 | 33 |
| GPT 5 mini | 20.8 | 33 |
2 rows. Highest Gemini 3 Flash (high) 54.2 (n 33). Lowest GPT 5 mini 20.8 (n 33).
Notesn = 33 per row
A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review).
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
| Item | Cost per resolved instance | n |
|---|---|---|
| Gemini 3 Flash (high) | $0.44 | 27 |
| GPT 5 mini | $0.08 | 21 |
2 rows. Highest Gemini 3 Flash (high) $0.44 (n 27). Lowest GPT 5 mini $0.08 (n 21).
Notesn 21–27 per row
Same 33 SWE-bench Verified instances; all attempts in the numerator
Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Gemini 3 Flash or GPT 5 mini?
- Gemini 3 Flash and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
- How were Gemini 3 Flash and GPT 5 mini measured?
- They share 3 measured metrics from 2 public studies: Agent on SWE-bench Verified vs 11 public models; What if every call ran on Opus? Repricing real agent tokens. Every row names its configuration, its sample size and its interval or range.
- How do Gemini 3 Flash and GPT 5 mini compare on resolved rate on the same 33 SWE-bench Verified instances?
- Gemini 3 Flash: 82% (27/33) (effort high · public mini-SWE-agent v2 run, same instances; n = 33; 95% interval 66% to 91%). GPT 5 mini: 64% (21/33) (public mini-SWE-agent v2 run, same instances; n = 33; 95% interval 47% to 78%). The 95% intervals overlap (Gemini 3 Flash 66% to 91%; GPT 5 mini 47% to 78%), so this sample cannot separate them.
- How do Gemini 3 Flash and GPT 5 mini compare on recorded cost per resolved instance: Agent vs the public panel?
- Gemini 3 Flash: $0.44 (effort high · public mini-SWE-agent v2 run, same instances; n = 27). GPT 5 mini: $0.080 (public mini-SWE-agent v2 run, same instances; n = 21). No interval or range was recorded for either side, so the gap ($0.44 vs $0.080, 5.5x) is not tested against run-to-run variation.
The studies behind this page
Agent on SWE-bench Verified vs 11 public models
Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.
What if every call ran on Opus? Repricing real agent tokens
Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.