Aider polyglot
Integrity rank #31 of 61 · 43 models scored · top score 88 · GPT-5
unknown
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-5 | 88 | unverified· optimizedT1 | 2025-08-07 |
| 2 | o3-pro | 84.9 | unverified· optimizedT1 | 2025-06-10 |
| 3 | Gemini 2.5 Pro (Jun 2025) | 83.1 | unverified· optimizedT1 | 2025-06-05 |
| 4 | o3 | 81.3 | unverified· optimizedT1 | 2024-12-20 |
| 5 | Grok 4 | 79.6 | unverified· optimizedT1 | 2025-07-09 |
| 6 | Gemini 2.5 Pro (May 2025) | 76.9 | unverified· optimizedT1 | 2025-05-06 |
| 7 | DeepSeek-V3.2 | 74.2 | unverified· optimizedT1 | 2025-12-01 |
| 8 | DeepSeek-V3.2-Exp | 74.2 | unverified· optimizedT1 | 2025-09-29 |
| 9 | Gemini 2.5 Pro (Mar 2025) | 72.9 | unverified· optimizedT1 | 2025-03-25 |
| 10 | Claude Opus 4 | 72 | unverified· optimizedT1 | 2025-05-22 |
| 11 | o4-mini | 72 | unverified· optimizedT1 | 2025-04-16 |
| 12 | DeepSeek-R1 (May 2025) | 71.4 | unverified· optimizedT1 | 2025-05-28 |
| 13 | Claude 3.7 Sonnet | 64.9 | unverified· optimizedT1 | 2025-02-24 |
| 14 | o1 | 61.7 | unverified· optimizedT1 | 2024-12-05 |
| 15 | Claude Sonnet 4 | 61.3 | unverified· optimizedT1 | 2025-05-22 |
| 16 | o3-mini | 60.4 | unverified· optimizedT1 | 2025-01-31 |
| 17 | Qwen3-235B-A22B | 59.6 | unverified· optimizedT1 | 2025-04-28 |
| 18 | Qwen3-235B-A22B-Instruct (Jul 2025) | 59.6 | unverified· optimizedT1 | 2025-07-25 |
| 19 | Kimi K2 (Jul 2025) | 59.1 | unverified· optimizedT1 | 2025-07-11 |
| 20 | DeepSeek-R1 | 56.9 | unverified· optimizedT1 | 2025-01-20 |
| 21 | Gemini 2.5 Flash (May 2025) | 55.1 | unverified· optimizedT1 | 2025-05-20 |
| 22 | DeepSeek-V3 (Mar 2025) | 55.1 | unverified· optimizedT1 | 2025-03-24 |
| 23 | Grok 3 | 53.3 | unverified· optimizedT1 | 2025-02-17 |
| 24 | GPT-4.1 | 52.4 | unverified· optimizedT1 | 2025-04-14 |
| 25 | Claude 3.5 Sonnet (October 2024) | 51.6 | unverified· optimizedT1 | 2024-10-22 |
| 26 | Grok-3 mini | 49.3 | unverified· optimizedT1 | 2025-02-19 |
| 27 | DeepSeek-V3 | 48.4 | unverified· optimizedT1 | 2024-12-24 |
| 28 | Gemini 2.5 Flash (Apr 2025) | 47.1 | unverified· optimizedT1 | 2025-04-17 |
| 29 | GPT-4.5 | 44.9 | unverified· optimizedT1 | 2025-02-27 |
| 30 | gpt-oss-120b | 41.8 | unverified· optimizedT1 | 2025-08-05 |
| 31 | Gemini 2.0 Flash (Dec 2024) | 38.2 | unverified· optimizedT1 | 2024-12-11 |
| 32 | Gemini 2.0 Pro | 35.6 | unverified· optimizedT1 | 2024-12-11 |
| 33 | o1-mini | 32.9 | unverified· optimizedT1 | 2024-09-12 |
| 34 | GPT-4.1 mini | 32.4 | unverified· optimizedT1 | 2025-04-14 |
| 35 | Claude 3.5 Haiku | 28 | unverified· optimizedT1 | 2024-10-22 |
| 36 | GPT-4o (Aug 2024) | 23.1 | unverified· optimizedT1 | 2024-05-13 |
| 37 | Qwen2.5-Max | 21.8 | unverified· optimizedT1 | 2025-01-25 |
| 38 | GPT-4o (Nov 2024) | 18.2 | unverified· optimizedT1 | 2024-05-13 |
| 39 | Gemini 2.0 Flash Thinking (Jan 2025) | 18.2 | unverified· optimizedT1 | 2025-01-21 |
| 40 | Llama 4 Maverick | 15.6 | unverified· optimizedT1 | 2025-04-05 |
| 41 | GPT-4.1 nano | 8.9 | unverified· optimizedT1 | 2025-04-14 |
| 42 | Gemma 3 27B | 4.9 | unverified· optimizedT1 | 2025-03-12 |
| 43 | GPT-4o mini | 3.6 | unverified· optimizedT1 | 2024-07-18 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 43 models, with a spread of 84.4 points between best and worst. This wide spread suggests that rank differences likely reflect genuine capability gaps.
2 cited facts
The benchmark's Benchmark Integrity Index is 87 (grade A), ranking it 25th out of 53 benchmarks, reflecting its health as a discriminator of frontier models under a disclosed harness. The weakest component of its integrity breakdown is contamination, which narrows how much to trust top scores.
5 cited facts
A top score of 88.0 with 12.0 points of open ceiling means the leader can still be beaten cleanly — a new high here would register as an outright, measurable advance rather than jitter on a spent scale. Only one model sits at the top, meaning there is no cluster of models within a couple of points; thus, the ranking retains genuine discriminative power.
4 cited facts
The harness is consistent, meaning scores from this benchmark are directly comparable across runs under the same evaluation protocol. The test set privacy is unknown, so it cannot be determined whether the benchmark's test set is public or held out; consequently, the potential for contamination from public exposure cannot be assessed. No contamination history has been recorded for this benchmark, and the harness reports model details and score sources, providing some transparency into evaluation conditions.
4 cited facts
This benchmark's measured quantity, construction, key assumptions, and limitations are not documented. The absence of documented limitations is the sharpest caveat, as it means the benchmark's scope and potential blind spots are opaque.
5 cited facts
Aider polyglot: Aider polyglot as reported in Epoch AI's Capabilities Index CSV.
GPT-5 leads Aider polyglot at 88. The full leaderboard above lists every recorded measurement, not just the headline number.
43 models have recorded scores on Aider polyglot, spanning a score spread of 84.4.
tensor.news grades Aider polyglot A for integrity (score 87/100), ranking #31 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — Aider polyglot still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on Aider polyglot are reasonably apples-to-apples.
Every Aider polyglot measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.