Zhipu AI·released 2026-08-141 source
GLM-5.3 benchmark scores: 10 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 91.1 on OTIS Mock AIME 2024-2025 (unverified)·0 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 91.1accuracy (%) | unverifiedT2 | 2026-08-14 |
| GPQA diamond | 87.88accuracy (%) | unverifiedT2 | 2026-08-14 |
| WeirdML | 75.4accuracy (%) | unverifiedT1 | 2026-08-14https://htihle.github.io/weirdml.html |
| DeepSWE | 68.96accuracy (%) | unverifiedT1 | 2026-08-14https://deepswe.datacurve.ai/ |
| FrontierMath-Tiers-1-3-v2-Private | 68.77accuracy (%) | unverifiedT2 | 2026-08-14 |
| ProofBench | 49accuracy (%) | unverifiedT1 | 2026-08-14https://www.vals.ai/benchmarks/proof_bench |
| SimpleQA Verified | 41accuracy (%) | unverifiedT2 | 2026-08-14 |
| FrontierMath-Tier-4-v2-Private | 29.27accuracy (%) | unverifiedT2 | 2026-08-14 |
| Mystery Game Puzzles | 26.2accuracy (%) | unverifiedT2 | 2026-08-14 |
| Chess Puzzles | 16.88accuracy (%) | unverifiedT2 | 2026-08-14 |
Claims drawn from cited facts, not live model generation.
This model was developed by Zhipu AI and released in 2026.
2 cited facts
This model is scored on 10 tracked benchmarks. It holds no current top score on any of them. Its record remains unverified, meaning outside evaluation has neither confirmed nor disputed the reported numbers, so they should be read with that uncertainty in mind.
3 cited facts
On average this model trails the leading score by 32.67 points across its benchmarks, a wide gap that reads as well behind the front of the pack rather than at the frontier. Consistently, it holds the top score on 0 benchmarks, so there is no category leadership to report yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GLM-5.3 is an AI model developed by Zhipu AI, released 2026-08-14. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GLM-5.3 has recorded scores on 10 benchmarks, each shown with its evidence status.
GLM-5.3 has recorded scores on 10 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, SimpleQA Verified, Chess Puzzles, FrontierMath-Tiers-1-3-v2-Private, and 4 more. The full table above shows each score with its evidence status.
0 of 10 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
GLM-5.3 has 10 tracked claims: 10 unverified.