Skip to content

Benchmark explainer

What is GSM8K?

Multi-step grade-school arithmetic word problems

Grade School Math 8K — Multi-step grade-school arithmetic word problems; scored as Final-answer exact match.

CINTEGRITY 68 / 100tensor.news
mixed harness — not comparablepublic test set

How it's scored

Metric
Final-answer exact match
Score ceiling
100
Construction
8.5k human-written problems (7.5k train/1319 test)
Human baseline
Not measured (middle-school level)

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

95.1/100

How far from ceiling / clustered at the top?

Contamination resistance

0/100

Public vs held-out; training-leak risk.

Harness comparability

50/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

23.5/100

How old is the benchmark?

Contamination history: Public since 2021; GSM1k clone showed contamination-indicative drops

Limitations: Saturated (>97%); ignores flawed reasoning; GSM1k showed overfitting; trivial w/ tools

Who leads GSM8K

ModelScoreEvidence
GPT-4 (Mar 2023)92self-reported
GPT-4o mini91.3self-reported
Qwen2.5-Coder-32B91.1self-reported
GPT-4 (Jun 2023)89.99unverified
Qwen2.5-Coder-14B88.7self-reported

Frequently asked questions

GSM8K (Grade School Math 8K): Grade School Math 8K — Multi-step grade-school arithmetic word problems; scored as Final-answer exact match.

GPT-4 (Mar 2023) leads GSM8K at 92 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.

56 models have recorded scores on GSM8K, spanning a score spread of 87.6.

tensor.news grades GSM8K C for integrity (score 68/100), ranking #51 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

No — GSM8K still has headroom and continues to discriminate between models rather than bunching them at the ceiling.

Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.

Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on GSM8K is not fully apples-to-apples.

Every GSM8K measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2110.14168.

Follow the record