Zhipu AI·released 2026-04-071 source
GLM-5.1 benchmark scores: 12 benchmarks tracked. 58% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 12 benchmarks·best result 93.33 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 12 independently reproduced·$1.4/$4.4 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 93.33accuracy (%) | reproduced· optimizedT1 | 2026-04-07Epoch AI |
| GPQA diamond | 86.53accuracy (%) | reproduced· optimizedT1 | 2026-04-07Epoch AI |
| SWE-Bench verified | 74.17accuracy (%) | reproduced· optimizedT1 | 2026-04-07Epoch AI |
| FrontierMath-2025-02-28-Private | 58.68accuracy (%) | reproduced· optimizedT1 | 2026-04-07Epoch AI |
| WeirdML | 57.1accuracy (%) | unverifiedT1 | 2026-04-07https://htihle.github.io/weirdml.html |
| SimpleBench | 46.12accuracy (%) | unverifiedT1 | 2026-04-07SimpleBench Leaderboard |
| SimpleQA Verified | 37.29accuracy (%) | reproducedT1 | 2026-04-07Epoch AI |
| ProofBench | 22.22accuracy (%) | unverifiedT1 | 2026-04-07https://www.vals.ai/benchmarks/proof_bench |
| FrontierMath-Tier-4-2025-07-01-Private | 20.83accuracy (%) | reproducedT1 | 2026-04-07Epoch AI |
| ExploitBench | 18.1accuracy (%) | unverifiedT2 | 2026-04-07 |
| Chess Puzzles | 14.77accuracy (%) | reproducedT1 | 2026-04-07Epoch AI |
| CritPt | 4.57accuracy (%) | unverifiedT2 | 2026-04-07 |
Claims drawn from cited facts, not live model generation.
This model was developed by Zhipu AI and released in 2026.
2 cited facts
This model is scored on 12 tracked benchmarks and currently holds no top score on any of them. Because the record is independently reproduced, outside evaluation backs these scores up, so readers can trust them as verified.
3 cited facts
Based on its average gap of 27.57 points under the disclosed harness, this model sits well behind the front-of-pack leaders, a wide gap rather than a competitive or frontier-level margin. It also holds the top score on 0 of the tracked benchmarks, meaning it currently shows no genuine category leadership on those tasks.
2 cited facts
6 facts cross-checked across data sources: 2 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GLM-5.1 is an AI model developed by Zhipu AI, released 2026-04-07. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GLM-5.1 has recorded scores on 12 benchmarks, each shown with its evidence status.
GLM-5.1 has recorded scores on 12 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, FrontierMath-2025-02-28-Private, and 6 more. The full table above shows each score with its evidence status.
7 of 12 recorded scores (58%) are independently reproduced rather than self-reported by the lab.
GLM-5.1 has 12 tracked claims: 7 independently reproduced, 5 unverified.
Listed API pricing: $1.4 per million input tokens, $4.4 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.