MiniMax·released 2026-02-12✓ 2 sources
MiniMax-M2.5 benchmark scores: 8 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 63.67 on ARC-AGI (unverified)·0 of 8 independently reproduced·$0.3/$1.2 per M tokens
Consensus: LiteLLM · Cross-check: OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC-AGI | 63.67accuracy (%) | unverified· optimizedT2 | 2026-02-12 |
| Terminal Bench | 42.7accuracy (%) | unverified· optimizedT1 | 2026-02-12https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| CL-bench | 11.4accuracy (%) | unverifiedT2 | 2026-02-12 |
| PostTrainBench | 9.5accuracy (%) | unverifiedT2 | 2026-02-12 |
| CL-bench Life | 6.3accuracy (%) | unverifiedT2 | 2026-02-12 |
| APEX-Agents | 6.2accuracy (%) | unverifiedT2 | 2026-02-12 |
| ARC-AGI-2 | 4.86accuracy (%) | unverifiedT2 | 2026-02-12 |
| ProofBench | 4accuracy (%) | unverifiedT1 | 2026-02-12https://www.vals.ai/benchmarks/proof_bench |
Claims drawn from cited facts, not live model generation.
Developed by MiniMax, this model is a frontier-scale 228.7-billion-parameter system released in 2026. Its large scale offers substantial headroom for complex tasks, though it carries correspondingly high computational expense.
4 cited facts
This model is scored on 8 tracked benchmarks and holds no current top score on any of them. Because its results are unverified, treat them as neither independently confirmed nor disputed by outside evaluation.
3 cited facts
Based on an average gap of 45.37 points under the disclosed harness, the model trails the leader by a wide margin, placing it well behind the front-of-pack. It holds no top scores on any benchmark (sota leader count of 0), so it has no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 1 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models, litellm_prices, openrouter_models
MiniMax-M2.5 is an AI model developed by MiniMax, released 2026-02-12. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
MiniMax-M2.5 has recorded scores on 8 benchmarks, each shown with its evidence status.
MiniMax-M2.5 has recorded scores on 8 benchmarks — ARC-AGI, ARC-AGI-2, APEX-Agents, PostTrainBench, ProofBench, Terminal Bench, and 2 more. The full table above shows each score with its evidence status.
0 of 8 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
MiniMax-M2.5 has 8 tracked claims: 8 unverified.
Listed API pricing: $0.3 per million input tokens, $1.2 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.