Benchmarking Agentic LLM and VLM Reasoning On Games
Integrity rank #28 of 61 · 24 models scored · top score 58.1 · Gemini 3 Pro
Long-horizon agentic planning, spatial reasoning, and sequential decision-making in games
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Gemini 3 Pro | 58.1 | unverifiedT1 | 2025-11-18 |
| 2 | Gemini 3.1 Pro | 57 | unverifiedT1 | 2026-02-19 |
| 3 | Gemini 3 Flash | 48.1 | unverifiedT1 | 2025-12-17 |
| 4 | Grok 4 | 43.6 | unverifiedT1 | 2025-07-09 |
| 5 | Claude Opus 4.5 | 43.5 | unverifiedT1 | 2025-11-24 |
| 6 | Gemini 2.5 Pro (Mar 2025) | 43.3 | unverifiedT1 | 2025-03-25 |
| 7 | DeepSeek-R1 | 34.9 | unverifiedT1 | 2025-01-20 |
| 8 | Gemini 2.5 Flash (Jun 2025) | 33.5 | unverifiedT1 | 2025-06-17 |
| 9 | GPT-5 | 32.8 | unverifiedT1 | 2025-08-07 |
| 10 | Claude 3.5 Sonnet (October 2024) | 32.6 | unverifiedT1 | 2024-10-22 |
| 11 | GPT-4o (May 2024) | 32.3 | unverifiedT1 | 2024-05-13 |
| 12 | Claude Haiku 4.5 | 31.2 | unverifiedT1 | 2025-10-15 |
| 13 | Grok 3 | 29.5 | unverifiedT1 | 2025-02-17 |
| 14 | Llama 3.1-70B | 27.9 | unverifiedT1 | 2024-07-23 |
| 15 | Llama 3.2 90B | 27.3 | unverifiedT1 | 2024-09-24 |
| 16 | Llama 3.3 70B | 23 | unverifiedT1 | 2024-12-06 |
| 17 | Gemini 1.5 Pro (Sept 2024) | 21 | unverifiedT1 | 2024-09-24 |
| 18 | Claude 3.5 Haiku | 19.3 | unverifiedT1 | 2024-10-22 |
| 19 | Mistral NeMo | 17.6 | unverifiedT1 | 2024-07-18 |
| 20 | GPT-4o mini | 17.4 | unverifiedT1 | 2024-07-18 |
| 21 | Qwen2.5-72B | 16.2 | unverifiedT1 | 2024-09-19 |
| 22 | Llama 3.1-8B | 15.1 | unverifiedT1 | 2024-07-23 |
| 23 | Gemini 1.5 Flash (Sep 2024) | 14.6 | unverifiedT1 | 2024-05-10 |
| 24 | Phi-4 | 11.6 | unverifiedT1 | 2024-12-12 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark encompasses 24 models, and the score spread of 46.5 points suggests that ranking differences may reflect genuine capability variations. The 46.5-point score spread indicates that a wide dispersion of scores reduces confidence in ranking-based trustworthiness assertions.
3 cited facts
This benchmark is not saturated, indicating that current top model rankings still reflect meaningful capability differences. With 41.9 points still unclaimed above the 58.1 leader, the benchmark retains real separating power — gaps at the top can widen or close without bumping the ceiling. Only two models cluster near the top score, suggesting limited competition at the highest performance levels.
4 cited facts
Balrog (Benchmarking Agentic LLM and VLM Reasoning On Games): Benchmarking Agentic LLM and VLM Reasoning On Games (BALROG) drops models into six procedurally-generated reinforcement-learning game environments (NetHack, MiniHack, BabyAI, Crafter, TextWorld, Baba Is AI) and measures how far they progress via natural-language actions. Far from saturated as of 2026.
Gemini 3 Pro leads Balrog at 58.1. The full leaderboard above lists every recorded measurement, not just the headline number.
24 models have recorded scores on Balrog, spanning a score spread of 46.5.
tensor.news grades Balrog A for integrity (score 89/100), ranking #28 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — Balrog still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on Balrog are reasonably apples-to-apples.
Every Balrog measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites balrogai.com.