Skip to content

Benchmark explainer

What is SimpleBench?

Basic human common-sense reasoning that resists memorized knowledge

A multiple-choice text benchmark of adversarial everyday-reasoning questions (spatio-temporal, social, and trick questions) on which unspecialized humans outperform frontier models. 213 questions total but only 10 are public; the rest are held out.

AINTEGRITY 98 / 100tensor.news
consistent harnessheld-out test set

How it's scored

Metric
Multiple-choice accuracy (6 options), reported AVG@5 at temperature 0.7 / top-p 0.95
Score ceiling
100
Construction
213 hand-written adversarial multiple-choice questions (6 options each); only 10 released publicly in simple_bench_public.json, ~203 held private
Human baseline
83.7% (small sample of 9 unspecialized humans); top model ~81.9% as of mid-2026, still below humans

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

98/100

How far from ceiling / clustered at the top?

Contamination resistance

100/100

Public vs held-out; training-leak risk.

Harness comparability

100/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

82.2/100

How old is the benchmark?

Contamination history: Held-out design is the primary defense; only the 10-example sample public since 2024; no confirmed train-set leakage found

Limitations: Very small human sample (n=9); only 10 public items limits external auditability; 213 items is a small set with non-trivial per-question noise; trick-question framing may reward a specific answer style

Who leads SimpleBench

ModelScoreEvidence
Claude Fable 578.28unverified
Claude Opus 576.72unverified
Gemini 3.1 Pro75.52unverified
GPT-5.5 Pro72.28unverified
Gemini 3.5 Flash72.04unverified

Compare the top SimpleBench scorers

Frequently asked questions

SimpleBench: A multiple-choice text benchmark of adversarial everyday-reasoning questions (spatio-temporal, social, and trick questions) on which unspecialized humans outperform frontier models. 213 questions total but only 10 are public; the rest are held out.

Claude Fable 5 leads SimpleBench at 78.28. The full leaderboard above lists every recorded measurement, not just the headline number.

80 models have recorded scores on SimpleBench, spanning a score spread of 78.28.

tensor.news grades SimpleBench A for integrity (score 98/100), ranking #10 of 62 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

No — SimpleBench still has headroom and continues to discriminate between models rather than bunching them at the ceiling.

Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.

Scores are reported under a consistent harness, so comparisons on SimpleBench are reasonably apples-to-apples.

Every SimpleBench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites simple-bench.com.

Follow the record