Skip to content

Benchmark explainer

What is FrontierMath Tier 4?

unknown

FrontierMath-Tier-4-v2-Private as reported in Epoch AI's Capabilities Index CSV.

AINTEGRITY 99 / 100tensor.news
consistent harnessheld-out test set

How it's scored

Metric
normalized performance in Epoch CSV, rendered as accuracy (%)
Score ceiling
100
Construction
unknown
Human baseline
unknown

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

98.2/100

How far from ceiling / clustered at the top?

Contamination resistance

100/100

Public vs held-out; training-leak risk.

Harness comparability

100/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

100/100

How old is the benchmark?

Contamination history: unknown

Limitations: unknown

Who leads FrontierMath Tier 4

ModelScoreEvidence
Claude Fable 587.8reproduced
GPT-5.6 Sol82.93reproduced
GPT-5.5 Pro78.05reproduced
Claude Opus 573.17reproduced
GPT-5.572.5reproduced

Frequently asked questions

FrontierMath-Tier-4-v2-Private: FrontierMath-Tier-4-v2-Private as reported in Epoch AI's Capabilities Index CSV.

Claude Fable 5 leads FrontierMath-Tier-4-v2-Private at 87.8 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.

48 models have recorded scores on FrontierMath-Tier-4-v2-Private, spanning a score spread of 87.8.

tensor.news grades FrontierMath-Tier-4-v2-Private A for integrity (score 99/100), ranking #9 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

No — FrontierMath-Tier-4-v2-Private still has headroom and continues to discriminate between models rather than bunching them at the ceiling.

Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.

Scores are reported under a consistent harness, so comparisons on FrontierMath-Tier-4-v2-Private are reasonably apples-to-apples.

Every FrontierMath-Tier-4-v2-Private measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.

Follow the record