The Agent Company
Integrity rank #27 of 61 · 14 models scored · top score 42.9 · DeepSeek-V3.2-Exp
unknown
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | DeepSeek-V3.2-Exp | 42.9 | unverifiedT1 | 2025-09-29 |
| 2 | Gemini 2.5 Flash (Sep 2025) | 41.1 | unverifiedT1 | 2025-09-25 |
| 3 | Claude Sonnet 4 | 33.1 | unverifiedT1 | 2025-05-22 |
| 4 | Claude 3.7 Sonnet | 30.9 | unverifiedT1 | 2025-02-24 |
| 5 | Gemini 2.5 Pro (May 2025) | 30.3 | unverifiedT1 | 2025-05-06 |
| 6 | Claude 3.5 Sonnet (October 2024) | 24 | unverifiedT1 | 2024-10-22 |
| 7 | Gemini 2.0 Flash (Feb 2025) | 11.4 | unverifiedT1 | 2024-12-11 |
| 8 | GPT-4o (Nov 2024) | 8.6 | unverifiedT1 | 2024-05-13 |
| 9 | Llama 3.1-405B | 7.4 | unverifiedT1 | 2024-07-23 |
| 10 | Llama 3.1-70B | 6.9 | unverifiedT1 | 2024-07-23 |
| 11 | Qwen2.5-72B | 5.7 | unverifiedT1 | 2024-09-19 |
| 12 | Gemini 1.5 Pro (Sept 2024) | 3.4 | unverifiedT1 | 2024-09-24 |
| 13 | Amazon Nova Pro | 1.7 | unverifiedT1 | 2024-12-03 |
| 14 | Qwen2-72B | 1.1 | unverifiedT1 | 2024-06-07 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 14 models, with a spread of 41.8 points between the best and worst scores. A spread this wide supports real separation between models, indicating that rank differences reflect genuine capability gaps.
2 cited facts
The benchmark achieves a Benchmark Integrity Index of 92, earning a grade of A, and ranks 16th out of 51 benchmarks. Its weakest component is contamination, which narrows how much to trust the top scores. Overall, this index represents the benchmark's health as a discriminator of frontier models under a disclosed harness.
5 cited facts
Far from ceiling-bound, this benchmark retains the range to separate frontier models — its top ordering carries information rather than compression artifacts. The leader has claimed 42.9 points against 57.1 still on the table — with this much unconquered range, expect the ordering to keep shifting; today's ranks are provisional. Only 2 models cluster near the top within a couple of points, meaning there is genuine room to separate the field at the top.
4 cited facts
The harness for this benchmark is labeled 'consistent harness,' indicating that scores are directly comparable across runs, which avoids the issue of low comparability where scores from different harnesses should not be compared head-to-head. The test set privacy status is unknown, meaning it is unclear whether the test set is public or held out; a public test set would require greater scrutiny for potential contamination than a held-out one, and this ambiguity warrants caution. Contamination history is unknown, providing no additional qualitative context regarding possible data leakage.
3 cited facts
The specific capability or behavior this benchmark measures is not documented. Details of how the benchmark was constructed, along with its key assumptions and limitations, are also not documented. Consequently, the most important caveat for readers is that without documented limitations or assumptions, it is impossible to know where this benchmark may mislead or what it cannot assess.
6 cited facts
The Agent Company: The Agent Company as reported in Epoch AI's Capabilities Index CSV.
DeepSeek-V3.2-Exp leads The Agent Company at 42.9. The full leaderboard above lists every recorded measurement, not just the headline number.
14 models have recorded scores on The Agent Company, spanning a score spread of 41.8.
tensor.news grades The Agent Company A for integrity (score 91/100), ranking #27 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — The Agent Company still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on The Agent Company are reasonably apples-to-apples.
Every The Agent Company measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.