FrontierMath-Tier-4-2025-07-01-Private
Integrity rank #25 of 61 · 27 models scored · top score 31.25 · Gemini 3 Pro
unknown
Strongest on contamination resistance, weakest on discrimination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Gemini 3 Pro | 31.25 | reproducedT1 | 2025-11-18 |
| 2 | Muse Spark | 24.33 | reproducedT1 | 2026-04-08 |
| 3 | GLM-5.1 | 20.83 | reproducedT1 | 2026-04-07 |
| 4 | GPT-5.1 | 20.83 | reproducedT1 | 2025-11-13 |
| 5 | Qwen 3.6 Plus | 13.89 | reproducedT1 | 2026-04-01 |
| 6 | Claude Sonnet 4.6 | 13.83 | reproducedT1 | 2026-02-17 |
| 7 | Kimi K2.5 | 7 | reproducedT1 | 2026-02-02 |
| 8 | Gemini 2.5 Flash (Jun 2025) | 6.94 | reproducedT1 | 2025-06-17 |
| 9 | Claude Opus 4 | 6.94 | reproducedT1 | 2025-05-22 |
| 10 | Qwen 3.6 Max (Preview) | 6.94 | reproducedT1 | 2026-04-20 |
| 11 | GLM-4.6 | 3.55 | reproducedT1 | 2025-09-30 |
| 12 | DeepSeek-V3.2 | 3.5 | reproducedT1 | 2025-12-01 |
| 13 | GLM-5 | 3.5 | reproducedT1 | 2026-02-11 |
| 14 | Grok 4 | 3.47 | reproducedT1 | 2025-07-09 |
| 15 | Claude Haiku 4.5 | 3.47 | reproducedT1 | 2025-10-15 |
| 16 | Qwen 3.5 Plus (hosted 397B-A17B) | 3.47 | reproducedT1 | 2026-02-16 |
| 17 | o3 | 3.47 | reproducedT1 | 2024-12-20 |
| 18 | Grok 3 | 0 | reproducedT1 | 2025-02-17 |
| 19 | Claude 3.5 Sonnet | 0 | reproducedT1 | 2024-06-20 |
| 20 | Kimi K2 Thinking | 0 | reproducedT1 | 2025-11-06 |
| 21 | Claude 3.5 Sonnet (October 2024) | 0 | reproducedT1 | 2024-10-22 |
| 22 | Qwen 3.5 Flash (hosted 35B-A3B) | 0 | reproducedT1 | 2026-02-25 |
| 23 | Qwen 3.6 Flash | 0 | reproducedT1 | 2026-04-26 |
| 24 | Qwen3-235B-A22B-Thinking (Jul 2025) | 0 | reproducedT1 | 2025-07-25 |
| 25 | Claude Sonnet 4 | 0 | reproducedT1 | 2025-05-22 |
| 26 | GLM-4.7 | 0 | reproducedT1 | 2025-12-22 |
| 27 | GPT-4.1 | 0 | reproducedT1 | 2025-04-14 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark measures unknown aspects, as what it measures is not documented. The construction of this benchmark is not documented, so its design details are unavailable. Key assumptions underlying its operation are not documented, making them unknown. Because its limitations are not documented, the benchmark may mislead readers about its capabilities.
4 cited facts
FrontierMath-Tier-4-2025-07-01-Private: FrontierMath-Tier-4-2025-07-01-Private as reported in Epoch AI's Capabilities Index CSV.
Gemini 3 Pro leads FrontierMath-Tier-4-2025-07-01-Private at 31.25 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
27 models have recorded scores on FrontierMath-Tier-4-2025-07-01-Private, spanning a score spread of 31.25.
tensor.news grades FrontierMath-Tier-4-2025-07-01-Private A for integrity (score 93/100), ranking #25 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — FrontierMath-Tier-4-2025-07-01-Private still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on FrontierMath-Tier-4-2025-07-01-Private are reasonably apples-to-apples.
Every FrontierMath-Tier-4-2025-07-01-Private measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.