FrontierMath-2025-02-28-Private
Integrity rank #8 of 61 · 45 models scored · top score 68.42 · Muse Spark
unknown
Strongest on discrimination, weakest on freshness.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Muse Spark | 68.42 | reproduced· optimizedT1 | 2026-04-08 |
| 2 | Gemini 3 Pro | 65.96 | reproduced· optimizedT1 | 2025-11-18 |
| 3 | GLM-5.1 | 58.68 | reproduced· optimizedT1 | 2026-04-07 |
| 4 | Claude Sonnet 4.6 | 56.84 | reproduced· optimizedT1 | 2026-02-17 |
| 5 | GPT-5.1 | 54.45 | reproduced· optimizedT1 | 2025-11-13 |
| 6 | Kimi K2.5 | 48.95 | reproduced· optimizedT1 | 2026-02-02 |
| 7 | Qwen 3.6 Plus | 45.98 | reproduced· optimizedT1 | 2026-04-01 |
| 8 | Qwen 3.6 Max (Preview) | 40.53 | reproduced· optimizedT1 | 2026-04-20 |
| 9 | DeepSeek-V3.2 | 38.77 | reproduced· optimizedT1 | 2025-12-01 |
| 10 | Kimi K2 Thinking | 37.55 | reproduced· optimizedT1 | 2025-11-06 |
| 11 | Qwen 3.5 Plus (hosted 397B-A17B) | 36.9 | reproduced· optimizedT1 | 2026-02-16 |
| 12 | Grok 4 | 34.48 | reproduced· optimizedT1 | 2025-07-09 |
| 13 | o3 | 32.78 | reproduced· optimizedT1 | 2024-12-20 |
| 14 | GLM-5 | 28.83 | reproduced· optimizedT1 | 2026-02-11 |
| 15 | Qwen 3.6 Flash | 18.15 | reproduced· optimizedT1 | 2026-04-26 |
| 16 | o1 | 16.33 | reproduced· optimizedT1 | 2024-12-05 |
| 17 | Qwen3-235B-A22B-Thinking (Jul 2025) | 14.88 | reproduced· optimizedT1 | 2025-07-25 |
| 18 | Qwen 3.5 Flash (hosted 35B-A3B) | 10.89 | reproduced· optimizedT1 | 2026-02-25 |
| 19 | Claude Haiku 4.5 | 10.36 | reproduced· optimizedT1 | 2025-10-15 |
| 20 | Grok-3 mini | 10.28 | reproduced· optimizedT1 | 2025-02-19 |
| 21 | GPT-4.1 | 9.68 | reproduced· optimizedT1 | 2025-04-14 |
| 22 | Gemini 2.5 Flash (Jun 2025) | 8.5 | reproduced· optimizedT1 | 2025-06-17 |
| 23 | Claude Opus 4 | 7.86 | reproduced· optimizedT1 | 2025-05-22 |
| 24 | GPT-4.1 mini | 7.86 | reproduced· optimizedT1 | 2025-04-14 |
| 25 | Claude 3.7 Sonnet | 7.26 | reproduced· optimizedT1 | 2025-02-24 |
| 26 | Claude Sonnet 4 | 7.26 | reproduced· optimizedT1 | 2025-05-22 |
| 27 | GLM-4.6 | 6.7 | reproduced· optimizedT1 | 2025-09-30 |
| 28 | Grok 3 | 6.65 | reproduced· optimizedT1 | 2025-02-17 |
| 29 | GLM-4.7 | 4.28 | reproduced· optimizedT1 | 2025-12-22 |
| 30 | Claude 3.5 Sonnet (October 2024) | 3.63 | reproduced· optimizedT1 | 2024-10-22 |
| 31 | DeepSeek-V3 | 3.02 | reproduced· optimizedT1 | 2024-12-24 |
| 32 | o1-mini | 3.02 | reproduced· optimizedT1 | 2024-09-12 |
| 33 | Gemini 2.0 Flash (Feb 2025) | 3.02 | reproduced· optimizedT1 | 2024-12-11 |
| 34 | Claude 3.5 Sonnet | 1.81 | reproduced· optimizedT1 | 2024-06-20 |
| 35 | Qwen2.5-Max | 1.81 | reproduced· optimizedT1 | 2025-01-25 |
| 36 | GPT-4.1 nano | 1.81 | reproduced· optimizedT1 | 2025-04-14 |
| 37 | Grok-2 (Dec 2024) | 1.21 | reproduced· optimizedT1 | 2024-08-13 |
| 38 | Llama 4 Maverick | 1.21 | reproduced· optimizedT1 | 2025-04-05 |
| 39 | Mistral Medium 3 | 0.61 | reproduced· optimizedT1 | 2025-05-07 |
| 40 | Claude 3.5 Haiku | 0.6 | reproduced· optimizedT1 | 2024-10-22 |
| 41 | Mistral Large 2 (Nov 2024) | 0.6 | reproduced· optimizedT1 | 2024-07-24 |
| 42 | GPT-4o (Aug 2024) | 0.6 | reproduced· optimizedT1 | 2024-05-13 |
| 43 | GPT-4o (Nov 2024) | 0.6 | reproduced· optimizedT1 | 2024-05-13 |
| 44 | Llama 4 Scout | 0 | reproduced· optimizedT1 | 2025-04-05 |
| 45 | Gemini 1.5 Flash (Sep 2024) | 0 | reproduced· optimizedT1 | 2024-05-10 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
The integrity score of 99 reflects strong benchmark health as a discriminator of frontier models under a disclosed harness. With a rank of 6 out of 51 benchmarks in the universe, this benchmark maintains high standing in reliability and consistency.
3 cited facts
Space remains at the top of this benchmark, so it can still do its job — separating frontier models — instead of bunching them against a ceiling. The top score stands at 68.42, leaving a headroom of 31.58 points to the theoretical ceiling, which means there is still genuine room to separate the field at the top. Only a single model forms the saturation cluster, so today's ranking near the top is not close to noise and differences there reflect real capability gaps rather than a tight cluster.
5 cited facts
What it measures is undocumented, limiting clarity on its objectives. The construction details of this benchmark are unknown, hindering understanding of its test-building process. Key_assumptions for this benchmark remain unspecified, which could affect its applicability. Limitations of this benchmark are undocumented, making its potential biases or errors unassessable.
7 cited facts
FrontierMath-2025-02-28-Private: FrontierMath-2025-02-28-Private as reported in Epoch AI's Capabilities Index CSV.
Muse Spark leads FrontierMath-2025-02-28-Private at 68.42 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
45 models have recorded scores on FrontierMath-2025-02-28-Private, spanning a score spread of 68.42.
tensor.news grades FrontierMath-2025-02-28-Private A for integrity (score 99/100), ranking #8 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — FrontierMath-2025-02-28-Private still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on FrontierMath-2025-02-28-Private are reasonably apples-to-apples.
Every FrontierMath-2025-02-28-Private measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.