Benchmark explainer
What is HLE?
Frontier expert knowledge/reasoning across 100+ subjects
Humanity's Last Exam — Frontier expert knowledge/reasoning across 100+ subjects; scored as Accuracy + calibration/confidence-error.
How it's scored
- Metric
- Accuracy + calibration/confidence-error
- Score ceiling
- 100
- Construction
- ~2500 Qs from ~1000 expert contributors, multi-stage review; ~10-14% multimodal
- Human baseline
- No aggregate (at/beyond individual expert frontier)
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
98.8/100How far from ceiling / clustered at the top?
Contamination resistance
100/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
100/100How old is the benchmark?
Contamination history: Public + private holdout to detect gaming; public leakage expected
Limitations: Short-answer/MC lucky hits; obscure recall; public portion contaminates; not agentic
Who leads HLE
| Model | Score | Evidence |
|---|---|---|
| Gemini 3.1 Pro | 43.74 | unverified |
| GPT-5.4 Pro | 41.51 | unverified |
| Muse Spark | 37.56 | unverified |
| Gemini 3 Pro | 34.37 | unverified |
| GPT-5.4 | 33.03 | unverified |
Frequently asked questions
HLE (Humanity's Last Exam): Humanity's Last Exam — Frontier expert knowledge/reasoning across 100+ subjects; scored as Accuracy + calibration/confidence-error.
Gemini 3.1 Pro leads HLE at 43.74. The full leaderboard above lists every recorded measurement, not just the headline number.
37 models have recorded scores on HLE, spanning a score spread of 43.74.
tensor.news grades HLE A for integrity (score 100/100), ranking #3 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — HLE still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on HLE are reasonably apples-to-apples.
Every HLE measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites agi.safe.ai.