Zhipu AI·released 2025-12-221 source
GLM-4.7 benchmark scores: 14 benchmarks tracked. 43% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 83.32 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 14 independently reproduced·$0.6/$2.2 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 83.32accuracy (%) | reproduced· optimizedT1 | 2025-12-22Epoch AI |
| GPQA diamond | 77.78accuracy (%) | reproduced· optimizedT1 | 2025-12-22Epoch AI |
| SimpleBench | 37.24accuracy (%) | unverifiedT1 | 2025-12-22SimpleBench Leaderboard |
| Terminal Bench | 33.4accuracy (%) | unverified· optimizedT2 | 2025-12-22 |
| SimpleQA Verified | 31.5accuracy (%) | reproducedT1 | 2025-12-22Epoch AI |
| CL-bench | 15.9accuracy (%) | unverifiedT2 | 2025-12-22 |
| CL-bench Life | 10.9accuracy (%) | unverifiedT2 | 2025-12-22 |
| APEX-Agents | 8.7accuracy (%) | unverifiedT2 | 2025-12-22 |
| PostTrainBench | 7.48accuracy (%) | unverifiedT2 | 2025-12-22 |
| ProofBench | 6accuracy (%) | unverifiedT1 | 2025-12-22https://www.vals.ai/benchmarks/proof_bench |
| FrontierMath-2025-02-28-Private | 4.28accuracy (%) | reproduced· optimizedT1 | 2025-12-22Epoch AI |
| CritPt | 1.71accuracy (%) | unverifiedT2 | 2025-12-22 |
| Chess Puzzles | 1.09accuracy (%) | reproducedT1 | 2025-12-22Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | 0accuracy (%) | reproducedT1 | 2025-12-22Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Zhipu AI and released in 2025.
2 cited facts
This model is scored on 14 tracked benchmarks and currently holds no top score on any of them. Because the results have been independently reproduced, the record can be trusted as externally verified.
3 cited facts
Based on the derived average SOTA gap, this model trails the leader by an average of 38.86 points under the disclosed evaluation harness, a wide gap that places it well behind the front-of-pack. It holds the top score on 0 benchmarks, so it has no demonstrated category leadership on the measured suites.
2 cited facts
6 facts cross-checked across data sources: 2 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GLM-4.7 is an AI model developed by Zhipu AI, released 2025-12-22. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GLM-4.7 has recorded scores on 14 benchmarks, each shown with its evidence status.
GLM-4.7 has recorded scores on 14 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleBench, FrontierMath-2025-02-28-Private, SimpleQA Verified, and 8 more. The full table above shows each score with its evidence status.
6 of 14 recorded scores (43%) are independently reproduced rather than self-reported by the lab.
GLM-4.7 has 14 tracked claims: 6 independently reproduced, 8 unverified.
Listed API pricing: $0.6 per million input tokens, $2.2 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.