Skip to content

Benchmark explainer

What is Chess Puzzles?

Spatial reasoning and planning via best-move selection

An Epoch AI internal benchmark of 100 programmatically generated chess positions (given as FEN); the model must output the single Stockfish-best move, scored as exact-match accuracy. Feeds quick Epoch Capabilities Index (ECI) estimates; the top frontier score was under 40% at its 2025-12 release.

AINTEGRITY 100 / 100tensor.news
consistent harnessheld-out test set

How it's scored

Metric
Single-best-move accuracy via exact string match of the move (e.g., 'b1c3')
Score ceiling
100
Construction
100 novel positions generated by Epoch: first ~6 full moves played randomly, then Stockfish plays at a randomly-drawn Elo; positions are kept only where Stockfish identifies one clear best move (roughly one kept per five simulated games).
Human baseline
unknown (Epoch reports no human baseline; a strong human or engine would score high, but no number is published)

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

98.5/100

How far from ceiling / clustered at the top?

Contamination resistance

100/100

Public vs held-out; training-leak risk.

Harness comparability

100/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

100/100

How old is the benchmark?

Contamination history: Positions were generated programmatically by Epoch and do not appear in any other source, so it is contamination-resistant by design (no known leakage as of its 2025-12 release)

Limitations: Only 100 items gives high variance; 'best move' is engine-defined (Stockfish-dependent); FEN/move-format parsing errors penalize models that reason correctly; brand-new with little external validation

Who leads Chess Puzzles

ModelScoreEvidence
GPT-5.5 Pro62.12reproduced
GPT-5.6 Sol62.12reproduced
GPT-5.4 Pro56.44reproduced
Gemini 3.1 Pro52.65reproduced
GPT-5.551.6reproduced

Frequently asked questions

Chess Puzzles (Epoch AI Chess Puzzles (ECI internal benchmark)): An Epoch AI internal benchmark of 100 programmatically generated chess positions (given as FEN); the model must output the single Stockfish-best move, scored as exact-match accuracy. Feeds quick Epoch Capabilities Index (ECI) estimates; the top frontier score was under 40% at its 2025-12 release.

GPT-5.5 Pro leads Chess Puzzles at 62.12 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.

81 models have recorded scores on Chess Puzzles, spanning a score spread of 62.12.

tensor.news grades Chess Puzzles A for integrity (score 100/100), ranking #2 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

No — Chess Puzzles still has headroom and continues to discriminate between models rather than bunching them at the ceiling.

Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.

Scores are reported under a consistent harness, so comparisons on Chess Puzzles are reasonably apples-to-apples.

Every Chess Puzzles measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/benchmarks/chess-puzzles.

Follow the record