GeoBench
Integrity rank #30 of 61 · 26 models scored · top score 88 · Gemini 3 Flash
unknown
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Gemini 3 Flash | 88 | unverifiedT1 | 2025-12-17 |
| 2 | Gemini 2.5 Pro (May 2025) | 86 | unverifiedT1 | 2025-05-06 |
| 3 | Gemini 3 Pro | 84 | unverifiedT1 | 2025-11-18 |
| 4 | Gemini 2.5 Pro (Mar 2025) | 81 | unverifiedT1 | 2025-03-25 |
| 5 | GPT-5 | 81 | unverifiedT1 | 2025-08-07 |
| 6 | o1 | 80 | unverifiedT1 | 2024-12-05 |
| 7 | Gemini 2.0 Flash (Feb 2025) | 77 | unverifiedT1 | 2024-12-11 |
| 8 | Gemini 2.5 Flash (May 2025) | 76 | unverifiedT1 | 2025-05-20 |
| 9 | Gemini 1.5 Flash (Sep 2024) | 76 | unverifiedT1 | 2024-05-10 |
| 10 | Claude Opus 4.5 | 75 | unverifiedT1 | 2025-11-24 |
| 11 | o3 | 74 | unverifiedT1 | 2024-12-20 |
| 12 | Gemini 2.5 Flash (Apr 2025) | 73 | unverifiedT1 | 2025-04-17 |
| 13 | GPT-4.1 | 72 | unverifiedT1 | 2025-04-14 |
| 14 | GPT-4o (Nov 2024) | 71 | unverifiedT1 | 2024-05-13 |
| 15 | Claude 3.7 Sonnet | 68 | unverifiedT1 | 2025-02-24 |
| 16 | o4-mini | 64 | unverifiedT1 | 2025-04-16 |
| 17 | GPT-4o mini | 64 | unverifiedT1 | 2024-07-18 |
| 18 | Claude 3.5 Sonnet (October 2024) | 62 | unverifiedT1 | 2024-10-22 |
| 19 | Qwen2.5-72B | 62 | unverifiedT1 | 2024-09-19 |
| 20 | Gemma 3 27B | 52 | unverifiedT1 | 2025-03-12 |
| 21 | Llama 3.2 90B | 52 | unverifiedT1 | 2024-09-24 |
| 22 | Llama 4 Maverick | 52 | unverifiedT1 | 2025-04-05 |
| 23 | Claude Opus 4 | 49 | unverifiedT1 | 2025-05-22 |
| 24 | Grok 4 | 45 | unverifiedT1 | 2025-07-09 |
| 25 | Claude Sonnet 4 | 37 | unverifiedT1 | 2025-05-22 |
| 26 | Claude 3.5 Haiku | 34 | unverifiedT1 | 2024-10-22 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
The benchmark’s health is quantified by an integrity score of 89, giving it an A grade and ranking 20th out of 51 among its peers. The weakest component is contamination, which limits confidence in the highest scores for readers.
4 cited facts
Scores under this harness are directly comparable due to standardized evaluation conditions and consistent metrics, as noted in harness comparability. The test set used for evaluation is public, indicating that high scores should be scrutinized for potential data contamination risks more rigorously than in held-out setups. Harnesses and optimized configurations may exhibit sensitivity to contamination, particularly with a public test set and no known contamination history yet.
4 cited facts
This benchmark's measured aspects are not documented. Unspecified construction leaves difficulty and provenance unverifiable. Ranks within this table still mean something; the absolute scores are not auditable. Key assumptions underlying this benchmark are not provided. The benchmark's limitations or key assumptions remain undocumented, potentially limiting reliable application.
4 cited facts
GeoBench: GeoBench as reported in Epoch AI's Capabilities Index CSV.
Gemini 3 Flash leads GeoBench at 88. The full leaderboard above lists every recorded measurement, not just the headline number.
26 models have recorded scores on GeoBench, spanning a score spread of 54.
tensor.news grades GeoBench A for integrity (score 88/100), ranking #30 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — GeoBench still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on GeoBench are reasonably apples-to-apples.
Every GeoBench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.