Benchmark explainer
What is Chess Puzzles?
Spatial reasoning and planning via best-move selection
An Epoch AI internal benchmark of 100 programmatically generated chess positions (given as FEN); the model must output the single Stockfish-best move, scored as exact-match accuracy. Feeds quick Epoch Capabilities Index (ECI) estimates; the top frontier score was under 40% at its 2025-12 release.
How it's scored
- Metric
- Single-best-move accuracy via exact string match of the move (e.g., 'b1c3')
- Score ceiling
- 100
- Construction
- 100 novel positions generated by Epoch: first ~6 full moves played randomly, then Stockfish plays at a randomly-drawn Elo; positions are kept only where Stockfish identifies one clear best move (roughly one kept per five simulated games).
- Human baseline
- unknown (Epoch reports no human baseline; a strong human or engine would score high, but no number is published)
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
98.5/100How far from ceiling / clustered at the top?
Contamination resistance
100/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
100/100How old is the benchmark?
Contamination history: Positions were generated programmatically by Epoch and do not appear in any other source, so it is contamination-resistant by design (no known leakage as of its 2025-12 release)
Limitations: Only 100 items gives high variance; 'best move' is engine-defined (Stockfish-dependent); FEN/move-format parsing errors penalize models that reason correctly; brand-new with little external validation
Who leads Chess Puzzles
| Model | Score | Evidence |
|---|---|---|
| GPT-5.5 Pro | 62.12 | reproduced |
| GPT-5.6 Sol | 62.12 | reproduced |
| GPT-5.4 Pro | 56.44 | reproduced |
| Gemini 3.1 Pro | 52.65 | reproduced |
| GPT-5.5 | 51.6 | reproduced |
Frequently asked questions
Chess Puzzles (Epoch AI Chess Puzzles (ECI internal benchmark)): An Epoch AI internal benchmark of 100 programmatically generated chess positions (given as FEN); the model must output the single Stockfish-best move, scored as exact-match accuracy. Feeds quick Epoch Capabilities Index (ECI) estimates; the top frontier score was under 40% at its 2025-12 release.
GPT-5.5 Pro leads Chess Puzzles at 62.12 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
81 models have recorded scores on Chess Puzzles, spanning a score spread of 62.12.
tensor.news grades Chess Puzzles A for integrity (score 100/100), ranking #2 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — Chess Puzzles still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on Chess Puzzles are reasonably apples-to-apples.
Every Chess Puzzles measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/benchmarks/chess-puzzles.