Skip to content

Benchmark explainer

What is HLE?

Frontier expert knowledge/reasoning across 100+ subjects

Humanity's Last Exam — Frontier expert knowledge/reasoning across 100+ subjects; scored as Accuracy + calibration/confidence-error.

AINTEGRITY 100 / 100tensor.news
consistent harnesspublic test set

How it's scored

Metric
Accuracy + calibration/confidence-error
Score ceiling
100
Construction
~2500 Qs from ~1000 expert contributors, multi-stage review; ~10-14% multimodal
Human baseline
No aggregate (at/beyond individual expert frontier)

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

98.8/100

How far from ceiling / clustered at the top?

Contamination resistance

100/100

Public vs held-out; training-leak risk.

Harness comparability

100/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

100/100

How old is the benchmark?

Contamination history: Public + private holdout to detect gaming; public leakage expected

Limitations: Short-answer/MC lucky hits; obscure recall; public portion contaminates; not agentic

ModelScoreEvidence
Gemini 3.1 Pro43.74unverified
GPT-5.4 Pro41.51unverified
Muse Spark37.56unverified
Gemini 3 Pro34.37unverified
GPT-5.433.03unverified

Frequently asked questions

HLE (Humanity's Last Exam): Humanity's Last Exam — Frontier expert knowledge/reasoning across 100+ subjects; scored as Accuracy + calibration/confidence-error.

Gemini 3.1 Pro leads HLE at 43.74. The full leaderboard above lists every recorded measurement, not just the headline number.

37 models have recorded scores on HLE, spanning a score spread of 43.74.

tensor.news grades HLE A for integrity (score 100/100), ranking #3 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

No — HLE still has headroom and continues to discriminate between models rather than bunching them at the ceiling.

Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.

Scores are reported under a consistent harness, so comparisons on HLE are reasonably apples-to-apples.

Every HLE measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites agi.safe.ai.

Follow the record