Grade School Math 8K
Integrity rank #51 of 61 · 56 models scored · top score 92 · GPT-4 (Mar 2023)
Multi-step grade-school arithmetic word problems
Strongest on discrimination, weakest on contamination resistance. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-4 (Mar 2023) | 92 | self-reported· optimizedT1 | 2023-03-15 |
| 2 | GPT-4o mini | 91.3 | self-reported· optimizedT1 | 2024-07-18 |
| 3 | Qwen2.5-Coder-32B | 91.1 | self-reported· optimizedT1 | 2024-09-18 |
| 4 | GPT-4 (Jun 2023) | 89.99 | unverified· optimizedT1 | 2023-06-13 |
| 5 | Qwen2.5-Coder-14B | 88.7 | self-reported· optimizedT1 | 2024-09-18 |
| 6 | Claude Instant | 86.7 | unverified· optimizedT1 | 2023-08-09 |
| 7 | Qwen2.5-Coder (7B) | 86.7 | self-reported· optimizedT1 | 2024-09-18 |
| 8 | Gemma 2 9B | 84.9 | self-reported· optimizedT1 | 2024-06-24 |
| 9 | Mistral NeMo | 84.2 | self-reported· optimizedT1 | 2024-07-18 |
| 10 | Llama 3.1-8B | 82.4 | self-reported· optimizedT1 | 2024-07-23 |
| 11 | Gemini 1.5 Flash (May 2024) | 82.4 | self-reported· optimizedT1 | 2024-05-10 |
| 12 | Yi-34B | 76 | unverified· optimizedT1 | 2023-11-02 |
| 13 | Qwen2.5-Coder-3B | 75.7 | self-reported· optimizedT1 | 2024-09-18 |
| 14 | Mixtral 8x7B | 74.4 | unverified· optimizedT1 | 2023-12-11 |
| 15 | Llama 2-70B | 69.6 | unverified· optimizedT1 | 2023-07-18 |
| 16 | Stable Beluga 2 | 69.6 | self-reported· optimizedT1 | 2023-07-20 |
| 17 | DeepSeek-Coder-V2-Lite-Base | 67.1 | self-reported· optimizedT1 | 2024-06-13 |
| 18 | Qwen2.5-Coder (1.5B) | 65.8 | self-reported· optimizedT1 | 2024-09-18 |
| 19 | internlm-20b | 62.9 | unverified· optimizedT1 | 2023-09-18 |
| 20 | Qwen-14B | 61.3 | self-reported· optimizedT1 | 2023-09-24 |
| 21 | GPT-3.5 Turbo (Jun 2023) | 57.77 | unverified· optimizedT1 | 2023-06-13 |
| 22 | StarCoder 2 15B | 57.7 | self-reported· optimizedT1 | 2024-02-29 |
| 23 | LLaMA-65B | 54.4 | unverified· optimizedT1 | 2023-02-24 |
| 24 | Mistral 7B v0.1 | 54.4 | unverified· optimizedT1 | 2023-10-10 |
| 25 | Falcon-180B | 54.4 | unverified· optimizedT1 | 2023-09-06 |
| 26 | Falcon 2 11B | 53.83 | self-reported· optimizedT1 | 2024-05-09 |
| 27 | Baichuan2-13B | 52.8 | self-reported· optimizedT1 | 2023-09-06 |
| 28 | Qwen-7B | 51.7 | unverified· optimizedT1 | 2023-09-28 |
| 29 | Gemma 7B | 46.4 | unverified· optimizedT1 | 2024-02-21 |
| 30 | Nemotron-4 15B | 46 | self-reported· optimizedT1 | 2024-02-27 |
| 31 | Yi 6B | 44.9 | unverified· optimizedT1 | 2023-11-02 |
| 32 | LLaMA-33B | 44.1 | unverified· optimizedT1 | 2023-02-27 |
| 33 | Llama 2-34B | 42.2 | self-reported· optimizedT1 | 2023-07-18 |
| 34 | INTELLECT-1 | 38.58 | self-reported· optimizedT1 | 2024-11-29 |
| 35 | CodeQwen1.5-7B | 37.7 | self-reported· optimizedT1 | 2024-04-15 |
| 36 | Llama 2-13B | 36.9 | unverified· optimizedT1 | 2023-07-18 |
| 37 | DeepSeek Coder 33B | 35.4 | self-reported· optimizedT1 | 2024-01-25 |
| 38 | Qwen2.5-Coder-0.5B | 34.5 | self-reported· optimizedT1 | 2024-09-18 |
| 39 | MPT-30B | 34.4 | unverified· optimizedT1 | 2023-06-22 |
| 40 | Falcon-40B | 33.8 | unverified· optimizedT1 | 2023-03-15 |
| 41 | StarCoder 2 7B | 32.7 | self-reported· optimizedT1 | 2024-02-29 |
| 42 | chatglm2-6b | 32.4 | unverified· optimizedT1 | 2023-06-24 |
| 43 | internlm-7b | 31.2 | self-reported· optimizedT1 | 2023-07-05 |
| 44 | vicuna-13b-v1.1 | 28.13 | unverified· optimizedT1 | 2023-04-12 |
| 45 | Baichuan 2-7B | 24.6 | unverified· optimizedT1 | 2023-09-20 |
| 46 | StarCoder 2 3B | 21.6 | self-reported· optimizedT1 | 2024-02-29 |
| 47 | DeepSeek Coder 6.7B | 21.3 | self-reported· optimizedT1 | 2024-01-25 |
| 48 | Qwen-1_8B | 21.2 | self-reported· optimizedT1 | 2023-11-30 |
| 49 | LLaMA-13B | 20.55 | unverified· optimizedT1 | 2023-02-27 |
| 50 | Gemma 2B | 17.7 | unverified· optimizedT1 | 2024-02-21 |
| 51 | Llama 2-7B | 16.7 | unverified· optimizedT1 | 2023-07-18 |
| 52 | LLaMA-7B | 11 | unverified· optimizedT1 | 2023-02-24 |
| 53 | Baichuan1-7B | 9.2 | unverified· optimizedT1 | 2023-06-01 |
| 54 | MPT-7B | 9.1 | unverified· optimizedT1 | 2023-05-05 |
| 55 | Falcon-7B | 6.8 | unverified· optimizedT1 | 2023-04-24 |
| 56 | DeepSeek Coder 1.3B | 4.4 | self-reported· optimizedT1 | 2024-01-25 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
Little ambiguity here: 87.6 points across 56 models means wide, readable gaps, and score differences can be taken at face value. This wide spread indicates that rank differences likely reflect genuine capability separation.
3 cited facts
The benchmark's Benchmark Integrity Index is 69 (grade C), ranking 41st out of 51 benchmarks, reflecting its health as a discriminator of frontier models under a disclosed harness. The weakest component is contamination, which narrows how much to trust top scores.
5 cited facts
Three models crowd the top near 92.0 with 8.0 points of ceiling left; the flag reads unsaturated, but ranking inside that leading cluster is close to noise — treat them as effectively tied. This means that today's ranking among top models is not yet close to noise; there remains genuine room to separate the field, as the cluster band is within a couple of points.
5 cited facts
Scores from this harness are not directly comparable to those from other harnesses, as the single-shot chain-of-thought setup is near-ceiling and no longer differentiates. The test set is public, so high scores require extra scrutiny for potential contamination; the benchmark has been public since 2021 and a clone showed contamination-indicative drops.
4 cited facts
Grade-school arithmetic sits at the easy end of any difficulty scale; scores here separate models on basic multi-step reliability and say nothing about mathematics beyond word problems. Final-answer grading cannot see the working, so a right number reached by a broken route counts the same as sound reasoning. That gap is the score's main soft spot. The sharpest caveat is that performance is saturated, the benchmark ignores flawed reasoning, a related earlier benchmark indicated overfitting, and the problems become trivial with tools.
4 cited facts
GSM8K (Grade School Math 8K): Grade School Math 8K — Multi-step grade-school arithmetic word problems; scored as Final-answer exact match.
GPT-4 (Mar 2023) leads GSM8K at 92 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.
56 models have recorded scores on GSM8K, spanning a score spread of 87.6.
tensor.news grades GSM8K C for integrity (score 68/100), ranking #51 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — GSM8K still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on GSM8K is not fully apples-to-apples.
Every GSM8K measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2110.14168.