Benchmark explainer
What is FrontierMath Tier 4?
unknown
FrontierMath-Tier-4-v2-Private as reported in Epoch AI's Capabilities Index CSV.
How it's scored
- Metric
- normalized performance in Epoch CSV, rendered as accuracy (%)
- Score ceiling
- 100
- Construction
- unknown
- Human baseline
- unknown
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
98.2/100How far from ceiling / clustered at the top?
Contamination resistance
100/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
100/100How old is the benchmark?
Contamination history: unknown
Limitations: unknown
Who leads FrontierMath Tier 4
| Model | Score | Evidence |
|---|---|---|
| Claude Fable 5 | 87.8 | reproduced |
| GPT-5.6 Sol | 82.93 | reproduced |
| GPT-5.5 Pro | 78.05 | reproduced |
| Claude Opus 5 | 73.17 | reproduced |
| GPT-5.5 | 72.5 | reproduced |
Frequently asked questions
FrontierMath-Tier-4-v2-Private: FrontierMath-Tier-4-v2-Private as reported in Epoch AI's Capabilities Index CSV.
Claude Fable 5 leads FrontierMath-Tier-4-v2-Private at 87.8 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
48 models have recorded scores on FrontierMath-Tier-4-v2-Private, spanning a score spread of 87.8.
tensor.news grades FrontierMath-Tier-4-v2-Private A for integrity (score 99/100), ranking #9 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — FrontierMath-Tier-4-v2-Private still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on FrontierMath-Tier-4-v2-Private are reasonably apples-to-apples.
Every FrontierMath-Tier-4-v2-Private measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.