Zhipu AI·released 2026-02-11✓ 2 sources
GLM-5 benchmark scores: 14 benchmarks tracked. 43% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 83.76 on GPQA diamond (reproduced)·6 of 14 independently reproduced·$1/$3.2 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GPQA diamond | 83.76accuracy (%) | reproduced· optimizedT1 | 2026-02-11Epoch AI |
| OTIS Mock AIME 2024-2025 | 79.98accuracy (%) | reproduced· optimizedT1 | 2026-02-11Epoch AI |
| SWE-Bench verified | 72.08accuracy (%) | reproduced· optimizedT1 | 2026-02-11Epoch AI |
| Terminal Bench | 52.4accuracy (%) | unverified· optimizedT2 | 2026-02-11 |
| WeirdML | 48.17accuracy (%) | unverifiedT1 | 2026-02-11WeirdML Leaderboard |
| ARC-AGI | 44.67accuracy (%) | unverified· optimizedT2 | 2026-02-11 |
| SimpleBench | 43.84accuracy (%) | unverifiedT1 | 2026-02-11SimpleBench Leaderboard |
| FrontierMath-2025-02-28-Private | 28.83accuracy (%) | reproduced· optimizedT1 | 2026-02-11Epoch AI |
| CL-bench | 18.7accuracy (%) | unverifiedT2 | 2026-02-11 |
| APEX-Agents | 17.2accuracy (%) | unverifiedT2 | 2026-02-11 |
| PostTrainBench | 13.88accuracy (%) | unverifiedT2 | 2026-02-11 |
| Chess Puzzles | 5.3accuracy (%) | reproducedT1 | 2026-02-11Epoch AI |
| ARC-AGI-2 | 4.86accuracy (%) | unverifiedT2 | 2026-02-11 |
| FrontierMath-Tier-4-2025-07-01-Private | 3.5accuracy (%) | reproducedT1 | 2026-02-11Epoch AI |
Claims drawn from cited facts, not live model generation.
Developed by Zhipu AI, this model is a frontier-scale system with approximately 753.91 billion parameters, released in 2026. That scale implies substantial capability headroom but also correspondingly high computational cost and resource demands.
4 cited facts
This model is scored on 14 tracked benchmarks and currently holds no top score in any of them. Because the record is independently reproduced by outside evaluation, these scores can be treated as verified rather than vendor-claimed.
3 cited facts
Based on the available average gap metric, it trails the leader by 34.41 points on average — a wide gap that places it well behind the leaders rather than at the frontier. It holds the top score on none of the benchmarks, so no genuine category leadership is demonstrated here.
2 cited facts
7 facts cross-checked across data sources: 1 corroborated, 2 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models, litellm_prices, modelsdev_models +1 more
GLM-5 is an AI model developed by Zhipu AI, released 2026-02-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GLM-5 has recorded scores on 14 benchmarks, each shown with its evidence status.
GLM-5 has recorded scores on 14 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, FrontierMath-2025-02-28-Private, and 8 more. The full table above shows each score with its evidence status.
6 of 14 recorded scores (43%) are independently reproduced rather than self-reported by the lab.
GLM-5 has 14 tracked claims: 6 independently reproduced, 8 unverified.
Listed API pricing: $1 per million input tokens, $3.2 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.