MiniMax·released 2026-06-011 source
MiniMax-M3 benchmark scores: 9 benchmarks tracked. 33% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 87.88 on GPQA diamond (reproduced)·3 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GPQA diamond | 87.88accuracy (%) | reproduced· optimizedT1 | 2026-06-01Epoch AI |
| OTIS Mock AIME 2024-2025 | 71.08accuracy (%) | reproduced· optimizedT1 | 2026-06-01Epoch AI |
| Surface Evolver Bench | 53.12accuracy (%) | unverifiedT1 | 2026-06-01https://yhenon.github.io/surface-evolver-llm-eval/ |
| SimpleBench | 34.96accuracy (%) | unverifiedT1 | 2026-06-01SimpleBench Leaderboard |
| ProofBench | 18accuracy (%) | unverifiedT1 | 2026-06-01https://www.vals.ai/benchmarks/proof_bench |
| FrontierCode | 14.7accuracy (%) | unverifiedT1 | 2026-06-01https://cognition.com/frontiercode |
| Chess Puzzles | 9.51accuracy (%) | reproducedT1 | 2026-06-01Epoch AI |
| OSWorld 2.0 | 4.6accuracy (%) | unverifiedT1 | 2026-06-01https://osworld-v2.xlang.ai/ |
| CritPt | 3.71accuracy (%) | unverifiedT2 | 2026-06-01 |
Claims drawn from cited facts, not live model generation.
This model was developed by MiniMax and released in 2026, making it a recently introduced system whose scale is not specified here.
2 cited facts
This model is tracked on nine benchmarks. Of those, it currently tops none. The record is independently reproduced, so outside evaluation backs the scores up.
3 cited facts
With an average gap of 37.37 points behind the leader, this model sits well behind the front of the pack on its benchmarks. It holds the top score on none of the tracked SOTA leaderboards, so it currently has no category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
MiniMax-M3 is an AI model developed by MiniMax, released 2026-06-01. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
MiniMax-M3 has recorded scores on 9 benchmarks, each shown with its evidence status.
MiniMax-M3 has recorded scores on 9 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleBench, CritPt, FrontierCode, and 3 more. The full table above shows each score with its evidence status.
3 of 9 recorded scores (33%) are independently reproduced rather than self-reported by the lab.
MiniMax-M3 has 9 tracked claims: 3 independently reproduced, 6 unverified.