FrontierMath-Tier-4-v2-Private
Integrity rank #9 of 61 · 48 models scored · top score 87.8 · Claude Fable 5
unknown
Strongest on discrimination, weakest on saturation headroom.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 87.8 | reproducedT1 | 2026-06-09 |
| 2 | GPT-5.6 Sol | 82.93 | reproducedT1 | 2026-07-09 |
| 3 | GPT-5.5 Pro | 78.05 | reproducedT1 | 2026-04-23 |
| 4 | Claude Opus 5 | 73.17 | reproducedT1 | 2026-07-24 |
| 5 | GPT-5.5 | 72.5 | reproducedT1 | 2026-04-23 |
| 6 | GPT-5.6 Terra | 70.73 | reproducedT1 | 2026-07-09 |
| 7 | GPT-5.6 Luna | 60.98 | reproducedT1 | 2026-07-09 |
| 8 | GPT-5.4 Pro | 58.54 | reproducedT1 | 2026-03-05 |
| 9 | Claude Opus 4.8 | 56.1 | reproducedT1 | 2026-05-28 |
| 10 | GPT-5.4 | 49 | reproducedT1 | 2026-03-05 |
| 11 | Qwen 3.8 Max | 46.34 | reproducedT1 | 2026-07-19 |
| 12 | GPT-5.2 Pro | 46 | reproducedT1 | 2025-12-11 |
| 13 | Kimi K3 | 39.02 | reproducedT1 | 2026-07-16 |
| 14 | Gemini 3.7 Flash | 36.59 | reproducedT1 | 2026-08-13 |
| 15 | Qwen3.7-Max | 34.15 | reproducedT1 | 2026-05-19 |
| 16 | Grok 4.6 | 31.71 | reproducedT1 | 2026-08-12 |
| 17 | Claude Opus 4.7 | 31.71 | reproducedT1 | 2026-04-16 |
| 18 | GPT-5.2 | 31.7 | reproducedT1 | 2025-12-11 |
| 19 | GLM-5.2 | 29.27 | reproducedT1 | 2026-06-16 |
| 20 | Claude Sonnet 5 | 29.27 | reproducedT1 | 2026-06-30 |
| 21 | Gemini 3.1 Pro | 26.83 | reproducedT1 | 2026-02-19 |
| 22 | Gemini 3.5 Flash | 26.83 | reproducedT1 | 2026-05-19 |
| 23 | Claude Opus 4.6 | 26.83 | reproducedT1 | 2026-02-05 |
| 24 | Kimi K2.6 | 25.64 | reproducedT1 | 2026-04-20 |
| 25 | Grok 4.5 | 24.39 | reproducedT1 | 2026-07-08 |
| 26 | DeepSeek V4 Flash 0731 | 24.39 | reproducedT1 | 2026-07-31 |
| 27 | Gemini 3.6 Flash | 21.95 | reproducedT1 | 2026-07-21 |
| 28 | GPT-5 | 21.95 | reproducedT1 | 2025-08-07 |
| 29 | GPT-5 Pro | 19.51 | reproducedT1 | 2025-10-07 |
| 30 | Gemini 3 Flash | 17.07 | reproducedT1 | 2025-12-17 |
| 31 | Grok 4.20 | 17.07 | reproducedT1 | 2026-02-17 |
| 32 | Inkling-Small | 17.07 | reproducedT1 | 2026-07-30 |
| 33 | Grok 4.3 Beta | 14.63 | reproducedT1 | 2026-04-17 |
| 34 | Kimi K2.7 Code | 12.2 | reproducedT1 | 2026-06-12 |
| 35 | GPT-5 mini | 12.2 | reproducedT1 | 2025-08-07 |
| 36 | GPT-5.4 Nano | 12.2 | reproducedT1 | 2026-03-17 |
| 37 | GPT-5.4 Mini | 9.76 | reproducedT1 | 2026-03-17 |
| 38 | Inkling | 4.88 | reproducedT1 | 2026-07-15 |
| 39 | Claude Opus 4.5 | 4.88 | reproducedT1 | 2025-11-24 |
| 40 | o4-mini | 4.88 | reproducedT1 | 2025-04-16 |
| 41 | DeepSeek-V4-Pro | 2.44 | reproducedT1 | 2026-04-24 |
| 42 | Claude Opus 4.1 | 2.44 | reproducedT1 | 2025-08-05 |
| 43 | Claude Sonnet 4.5 | 2.44 | reproducedT1 | 2025-09-29 |
| 44 | GPT-5 nano | 2.44 | reproducedT1 | 2025-08-07 |
| 45 | GPT-5.5 Instant | 2.44 | reproducedT1 | 2026-05-05 |
| 46 | Gemini 2.5 Pro (Jun 2025) | 0 | reproducedT1 | 2025-06-05 |
| 47 | Gemini 3.5 Flash-Lite | 0 | reproducedT1 | 2026-07-21 |
| 48 | o3-mini | 0 | reproducedT1 | 2025-01-31 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 48 models, with the best and worst results separated by 87.8 points. This wide spread indicates that score differences between models are likely to reflect genuine capability gaps rather than noise.
3 cited facts
The benchmark receives an integrity score of 99, a grade of A, ranking 9th out of 61 benchmarks. Its weakest component is saturation, meaning the ceiling is crowded and the benchmark's ability to discriminate among the strongest models is compressed.
6 cited facts
The benchmark is not saturated. The top score is 87.8, leaving 12.2 points of headroom to the ceiling. Only one model sits within a couple of points of the top, so there is genuine room to separate the field.
4 cited facts
Task performance scores are reported under a consistent harness, so they are directly comparable across configurations. The test set for this benchmark is held out, meaning high scores warrant less scrutiny for contamination than a public test set would. Contamination history is unknown, and the consistent harness limits sensitivity to setup differences.
4 cited facts
This benchmark's measurement target is not documented, so its intended scope cannot be stated from the available facts. How the test was built is also undocumented, and no key assumptions are recorded. The sharpest caveat is that its limitations are undocumented, leaving no documented basis for judging where it is most likely to mislead.
5 cited facts
FrontierMath-Tier-4-v2-Private: FrontierMath-Tier-4-v2-Private as reported in Epoch AI's Capabilities Index CSV.
Claude Fable 5 leads FrontierMath-Tier-4-v2-Private at 87.8 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
48 models have recorded scores on FrontierMath-Tier-4-v2-Private, spanning a score spread of 87.8.
tensor.news grades FrontierMath-Tier-4-v2-Private A for integrity (score 99/100), ranking #9 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — FrontierMath-Tier-4-v2-Private still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on FrontierMath-Tier-4-v2-Private are reasonably apples-to-apples.
Every FrontierMath-Tier-4-v2-Private measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.