Zhipu AI·released 2026-06-16✓ 2 sources
GLM-5.2 benchmark scores: 18 benchmarks tracked. 44% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 18 benchmarks·best result 89.14 on GPQA diamond (reproduced)·8 of 18 independently reproduced·$1.4/$4.4 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GPQA diamond | 89.14accuracy (%) | reproduced· optimizedT1 | 2026-06-16Epoch AI |
| OTIS Mock AIME 2024-2025 | 86.38accuracy (%) | reproduced· optimizedT1 | 2026-06-16Epoch AI |
| SWE-Bench verified | 78.7accuracy (%) | reproduced· optimizedT1 | 2026-06-16Epoch AI |
| ARC-AGI | 77accuracy (%) | unverified· optimizedT1 | 2026-06-16https://arcprize.org/leaderboard |
| WeirdML | 70.12accuracy (%) | unverifiedT1 | 2026-06-16https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 59.21accuracy (%) | reproducedT1 | 2026-06-16Epoch AI |
| Surface Evolver Bench | 55.62accuracy (%) | unverifiedT1 | 2026-06-16https://yhenon.github.io/surface-evolver-llm-eval/ |
| DeepSWE | 43.78accuracy (%) | unverifiedT1 | 2026-06-16https://deepswe.datacurve.ai/ |
| SimpleQA Verified | 38.1accuracy (%) | reproducedT1 | 2026-06-16Epoch AI |
| APEX-Agents | 35.6accuracy (%) | unverifiedT2 | 2026-06-16 |
| ProofBench | 35accuracy (%) | unverifiedT1 | 2026-06-16https://www.vals.ai/benchmarks/proof_bench |
| PostTrainBench | 34.29accuracy (%) | unverifiedT2 | 2026-06-16 |
| FrontierMath-Tier-4-v2-Private | 29.27accuracy (%) | reproducedT1 | 2026-06-16Epoch AI |
| FrontierCode | 24.5accuracy (%) | unverifiedT1 | 2026-06-16https://cognition.com/frontiercode |
| ARC-AGI-2 | 22.78accuracy (%) | unverifiedT2 | 2026-06-16 |
| CritPt | 20.86accuracy (%) | unverifiedT2 | 2026-06-16 |
| Chess Puzzles | 16.88accuracy (%) | reproducedT1 | 2026-06-16Epoch AI |
| EBR-bench | 9.52accuracy (%) | reproducedT1 | 2026-06-16Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Zhipu AI and released in 2026. With 753.38 billion parameters, it is frontier-scale, offering substantial headroom but also considerable expense.
3 cited facts
This model is tracked across 18 benchmarks and currently holds no top score in any of them. Because these results have been independently reproduced, the record can be treated as reliable.
3 cited facts
Relative to the best reported score on its benchmarks, this model trails by an average of 29.96 points, a wide gap that places it well behind the leaders under the disclosed harness. It holds the top score on none of the tracked benchmarks, so there is no current category leadership to point to.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 1 single-source, 3 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, huggingface_models, modelsdev_models +1 more
GLM-5.2 is an AI model developed by Zhipu AI, released 2026-06-16. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GLM-5.2 has recorded scores on 18 benchmarks, each shown with its evidence status.
GLM-5.2 has recorded scores on 18 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 12 more. The full table above shows each score with its evidence status.
8 of 18 recorded scores (44%) are independently reproduced rather than self-reported by the lab.
GLM-5.2 has 18 tracked claims: 8 independently reproduced, 10 unverified.
Listed API pricing: $1.4 per million input tokens, $4.4 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.