Zhipu AI·released 2026-08-201 source
GLM-5.3-Flash benchmark scores: 9 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 93.88 on OTIS Mock AIME 2024-2025 (unverified)·0 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 93.88accuracy (%) | unverifiedT2 | 2026-08-20 |
| GPQA diamond | 86.87accuracy (%) | unverifiedT2 | 2026-08-20 |
| DeepSWE | 63.39accuracy (%) | unverifiedT1 | 2026-08-20https://deepswe.datacurve.ai/ |
| FrontierMath-Tiers-1-3-v2-Private | 55.79accuracy (%) | unverifiedT2 | 2026-08-20 |
| Surface Evolver Bench | 52.5accuracy (%) | unverifiedT1 | 2026-08-20https://yhenon.github.io/surface-evolver-llm-eval/ |
| ProofBench | 21accuracy (%) | unverifiedT1 | 2026-08-20https://www.vals.ai/benchmarks/proof_bench |
| FrontierMath-Tier-4-v2-Private | 17.07accuracy (%) | unverifiedT2 | 2026-08-20 |
| Chess Puzzles | 9.51accuracy (%) | unverifiedT2 | 2026-08-20 |
| Mystery Game Puzzles | 0accuracy (%) | unverifiedT2 | 2026-08-20 |
Claims drawn from cited facts, not live model generation.
This model originates from the developer Zhipu AI. Its vintage places it in a recent release cycle, having arrived in the year 2026.
2 cited facts
This model is scored on 9 tracked benchmarks, but it holds no current top score among them. Its record is unverified: outside evaluation has neither confirmed nor disputed the reported numbers, so they should be read as unconfirmed claims rather than established results.
3 cited facts
On average this model trails the best benchmark score by 45.3 points, a wide gap that places it well behind the leaders rather than at or near the frontier. Consistently, it holds the top score on 0 benchmarks, meaning there is no category in which it currently leads.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GLM-5.3-Flash is an AI model developed by Zhipu AI, released 2026-08-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GLM-5.3-Flash has recorded scores on 9 benchmarks, each shown with its evidence status.
GLM-5.3-Flash has recorded scores on 9 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, FrontierMath-Tiers-1-3-v2-Private, FrontierMath-Tier-4-v2-Private, Mystery Game Puzzles, and 3 more. The full table above shows each score with its evidence status.
0 of 9 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
GLM-5.3-Flash has 9 tracked claims: 9 unverified.