Zhipu AI·released 2025-09-301 source
GLM-4.6 benchmark scores: 5 benchmarks tracked. 40% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 24.5 on Terminal Bench (unverified)·2 of 5 independently reproduced·$0.6/$2.2 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Terminal Bench | 24.5accuracy (%) | unverified· optimizedT1 | 2025-09-30Terminal-Bench v2 Leaderboard |
| FrontierMath-2025-02-28-Private | 6.7accuracy (%) | reproduced· optimizedT1 | 2025-09-30Epoch AI |
| APEX-Agents | 4accuracy (%) | unverifiedT2 | 2025-09-30 |
| FrontierMath-Tier-4-2025-07-01-Private | 3.55accuracy (%) | reproducedT1 | 2025-09-30Epoch AI |
| CritPt | 1.14accuracy (%) | unverifiedT2 | 2025-09-30 |
Claims drawn from cited facts, not live model generation.
This model originates from Zhipu AI. It was released in 2025, marking the developer's newest generation.
3 cited facts
This model is scored on five tracked benchmarks. It currently holds no top score on any of them. Because the results have been independently reproduced, its reported scores are backed by outside evaluation.
3 cited facts
It trails the SOTA leader by an average of 44.36 points under the disclosed harness, a wide gap indicating it sits well behind the front-runners. It holds the top score on none of the tracked benchmarks, so there is no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 2 single-source, 3 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GLM-4.6 is an AI model developed by Zhipu AI, released 2025-09-30. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GLM-4.6 has recorded scores on 5 benchmarks, each shown with its evidence status.
GLM-4.6 has recorded scores on 5 benchmarks — FrontierMath-2025-02-28-Private, CritPt, FrontierMath-Tier-4-2025-07-01-Private, APEX-Agents, Terminal Bench. The full table above shows each score with its evidence status.
2 of 5 recorded scores (40%) are independently reproduced rather than self-reported by the lab.
GLM-4.6 has 5 tracked claims: 2 independently reproduced, 3 unverified.
Listed API pricing: $0.6 per million input tokens, $2.2 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.