3 measured metrics · 2 studies
MiniMax M2.5vsKimi K2.5
No row separates them: 1 tie, 2 unclear.
The verdict
MiniMax M2.5 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | MiniMax M2.5 | Kimi K2.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 70% (23/33)effort high · public mini-SWE-agent v2 run, same instances | 70% (23/33)effort high · public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 53%–83% vs 53%–83% | Tie | The 95% intervals overlap (MiniMax M2.5 53% to 83%; Kimi K2.5 53% to 83%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.11effort high · public mini-SWE-agent v2 run, same instances | $0.26effort high · public mini-SWE-agent v2 run, same instances | 23 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.11 vs $0.26, 2.4x) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
| Model calls per instance | 58.4effort high · public mini-SWE-agent v2 run, same instances | 56.7effort high · public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
Marks: 95% intervals (Wilson for rates)n is shown per side on every row
3 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Recorded cost per resolved instance: Agent vs the public panel, 2.4x (Kimi K2.5 larger).
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- MiniMax M2.5
- Kimi K2.5
- 95% interval
- where the two overlap
Agent on SWE-bench Verified vs 11 public models
- Resolved rate on the same 33 SWE-bench Verified instances70% (23/33)n 3370% (23/33)n 33TieResolved rate on the same 33 SWE-bench Verified instances: MiniMax M2.5 70% (23/33) (n 33, 95% interval 53%–83%); Kimi K2.5 70% (23/33) (n 33, 95% interval 53%–83%). Tie.
- Model calls per instance58.4n 3356.7n 33UnclearModel calls per instance: MiniMax M2.5 58.4 (n 33); Kimi K2.5 56.7 (n 33). Unclear.
What if every call ran on Opus? Repricing real agent tokens
- Recorded cost per resolved instance: Agent vs the public panel$0.11n 23$0.26n 23UnclearRecorded cost per resolved instance: Agent vs the public panel: MiniMax M2.5 $0.11 (n 23); Kimi K2.5 $0.26 (n 23). Unclear.
| Metric | MiniMax M2.5 | Kimi K2.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 70% (23/33)effort high · public mini-SWE-agent v2 run, same instances | 70% (23/33)effort high · public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 53%–83% vs 53%–83% | Tie | The 95% intervals overlap (MiniMax M2.5 53% to 83%; Kimi K2.5 53% to 83%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 58.4effort high · public mini-SWE-agent v2 run, same instances | 56.7effort high · public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.11effort high · public mini-SWE-agent v2 run, same instances | $0.26effort high · public mini-SWE-agent v2 run, same instances | 23 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.11 vs $0.26, 2.4x) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
Marks: 95% intervals (Wilson for rates)n is shown per side on every row
3 rows from 2 studies. No row separates them: 1 tie, 2 unclear.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick MiniMax M2.5
No row in this data puts MiniMax M2.5 ahead of Kimi K2.5. Pick on other grounds (price, access, the tasks you run), or measure your own workload.
When to pick Kimi K2.5
No row in this data puts Kimi K2.5 ahead of MiniMax M2.5. Pick on other grounds (price, access, the tasks you run), or measure your own workload.
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Resolved rate | 95% interval | n |
|---|---|---|---|
| MiniMax M2.5 (high) | 70% | 53%–83% | 33 |
| Kimi K2.5 (high) | 70% | 53%–83% | 33 |
2 rows. All at 70%.
NotesWhiskers: 95% Wilson intervaln = 33 per row
Agent vs 11 public mini-SWE-agent v2 runs, one attempt each
Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
Model calls per instance
Mean over the same 33 instances
| Item | Mean calls | n |
|---|---|---|
| MiniMax M2.5 (high) | 58.4 | 33 |
| Kimi K2.5 (high) | 56.7 | 33 |
2 rows. Highest MiniMax M2.5 (high) 58.4 (n 33). Lowest Kimi K2.5 (high) 56.7 (n 33).
Notesn = 33 per row
A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review).
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
| Item | Cost per resolved instance | n |
|---|---|---|
| Kimi K2.5 (high) | $0.26 | 23 |
| MiniMax M2.5 (high) | $0.11 | 23 |
2 rows. Highest Kimi K2.5 (high) $0.26 (n 23). Lowest MiniMax M2.5 (high) $0.11 (n 23).
Notesn = 23 per row
Same 33 SWE-bench Verified instances; all attempts in the numerator
Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, MiniMax M2.5 or Kimi K2.5?
- MiniMax M2.5 and Kimi K2.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
- How were MiniMax M2.5 and Kimi K2.5 measured?
- They share 3 measured metrics from 2 public studies: Agent on SWE-bench Verified vs 11 public models; What if every call ran on Opus? Repricing real agent tokens. Every row names its configuration, its sample size and its interval or range.
- How do MiniMax M2.5 and Kimi K2.5 compare on resolved rate on the same 33 SWE-bench Verified instances?
- MiniMax M2.5: 70% (23/33) (effort high · public mini-SWE-agent v2 run, same instances; n = 33; 95% interval 53% to 83%). Kimi K2.5: 70% (23/33) (effort high · public mini-SWE-agent v2 run, same instances; n = 33; 95% interval 53% to 83%). The 95% intervals overlap (MiniMax M2.5 53% to 83%; Kimi K2.5 53% to 83%), so this sample cannot separate them.
- How do MiniMax M2.5 and Kimi K2.5 compare on recorded cost per resolved instance: Agent vs the public panel?
- MiniMax M2.5: $0.11 (effort high · public mini-SWE-agent v2 run, same instances; n = 23). Kimi K2.5: $0.26 (effort high · public mini-SWE-agent v2 run, same instances; n = 23). No interval or range was recorded for either side, so the gap ($0.11 vs $0.26, 2.4x) is not tested against run-to-run variation.
The studies behind this page
Agent on SWE-bench Verified vs 11 public models
Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.
What if every call ran on Opus? Repricing real agent tokens
Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.