Benchmark explainer
What is SimpleBench?
Basic human common-sense reasoning that resists memorized knowledge
A multiple-choice text benchmark of adversarial everyday-reasoning questions (spatio-temporal, social, and trick questions) on which unspecialized humans outperform frontier models. 213 questions total but only 10 are public; the rest are held out.
How it's scored
- Metric
- Multiple-choice accuracy (6 options), reported AVG@5 at temperature 0.7 / top-p 0.95
- Score ceiling
- 100
- Construction
- 213 hand-written adversarial multiple-choice questions (6 options each); only 10 released publicly in simple_bench_public.json, ~203 held private
- Human baseline
- 83.7% (small sample of 9 unspecialized humans); top model ~81.9% as of mid-2026, still below humans
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
98/100How far from ceiling / clustered at the top?
Contamination resistance
100/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
82.2/100How old is the benchmark?
Contamination history: Held-out design is the primary defense; only the 10-example sample public since 2024; no confirmed train-set leakage found
Limitations: Very small human sample (n=9); only 10 public items limits external auditability; 213 items is a small set with non-trivial per-question noise; trick-question framing may reward a specific answer style
Who leads SimpleBench
| Model | Score | Evidence |
|---|---|---|
| Claude Fable 5 | 78.28 | unverified |
| Claude Opus 5 | 76.72 | unverified |
| Gemini 3.1 Pro | 75.52 | unverified |
| GPT-5.5 Pro | 72.28 | unverified |
| Gemini 3.5 Flash | 72.04 | unverified |
Compare the top SimpleBench scorers
Frequently asked questions
SimpleBench: A multiple-choice text benchmark of adversarial everyday-reasoning questions (spatio-temporal, social, and trick questions) on which unspecialized humans outperform frontier models. 213 questions total but only 10 are public; the rest are held out.
Claude Fable 5 leads SimpleBench at 78.28. The full leaderboard above lists every recorded measurement, not just the headline number.
80 models have recorded scores on SimpleBench, spanning a score spread of 78.28.
tensor.news grades SimpleBench A for integrity (score 98/100), ranking #10 of 62 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — SimpleBench still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on SimpleBench are reasonably apples-to-apples.
Every SimpleBench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites simple-bench.com.