APEX-Agents
Integrity rank #5 of 61 · 47 models scored · top score 45 · Claude Fable 5
unknown
Strongest on discrimination, weakest on saturation headroom. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 45 | unverifiedT2 | 2026-06-09 |
| 2 | Claude Opus 5 | 43.5 | unverifiedT2 | 2026-07-24 |
| 3 | Claude Opus 4.8 | 42.5 | unverifiedT2 | 2026-05-28 |
| 4 | Muse Spark 1.1 | 41.9 | unverifiedT2 | 2026-07-09 |
| 5 | Grok 4.6 | 41.2 | unverifiedT2 | 2026-08-12 |
| 6 | GPT-5.6 Sol | 40 | unverifiedT2 | 2026-07-09 |
| 7 | Kimi K3 | 39.3 | unverifiedT2 | 2026-07-16 |
| 8 | GPT-5.5 | 38.5 | unverifiedT2 | 2026-04-23 |
| 9 | GPT-5.4 | 36 | unverifiedT2 | 2026-03-05 |
| 10 | GLM-5.2 | 35.6 | unverifiedT2 | 2026-06-16 |
| 11 | GPT-5.2 | 34.4 | unverifiedT2 | 2025-12-11 |
| 12 | Grok 4.5 | 34.2 | unverifiedT2 | 2026-07-08 |
| 13 | Claude Opus 4.7 | 33.9 | unverifiedT2 | 2026-04-16 |
| 14 | Gemini 3.1 Pro | 33.5 | unverifiedT2 | 2026-02-19 |
| 15 | Claude Sonnet 5 | 32.5 | unverifiedT2 | 2026-06-30 |
| 16 | Claude Opus 4.6 | 32.4 | unverifiedT2 | 2026-02-05 |
| 17 | GPT-5.3 Codex | 31.8 | unverifiedT2 | 2026-02-05 |
| 18 | Gemini 3 Pro | 31.5 | unverifiedT2 | 2025-11-18 |
| 19 | Kimi K2.7 Code | 27.6 | unverifiedT2 | 2026-06-12 |
| 20 | GPT-5.4 Mini | 24.6 | unverifiedT2 | 2026-03-17 |
| 21 | Gemini 3 Flash | 24 | unverifiedT2 | 2025-12-17 |
| 22 | Claude Sonnet 4.6 | 23.7 | unverifiedT2 | 2026-02-17 |
| 23 | Claude Opus 4.5 | 20.7 | unverifiedT2 | 2025-11-24 |
| 24 | Kimi K2.6 | 18.9 | unverifiedT2 | 2026-04-20 |
| 25 | GPT-5 | 18.3 | unverifiedT2 | 2025-08-07 |
| 26 | GPT-5.1 | 17.5 | unverifiedT2 | 2025-11-13 |
| 27 | GLM-5 | 17.2 | unverifiedT2 | 2026-02-11 |
| 28 | o3 | 17.2 | unverifiedT2 | 2024-12-20 |
| 29 | GPT-5.4 Nano | 16.9 | unverifiedT2 | 2026-03-17 |
| 30 | Grok 4 | 15.2 | unverifiedT2 | 2025-07-09 |
| 31 | Kimi K2.5 | 14.4 | unverifiedT2 | 2026-02-02 |
| 32 | Qwen 3.5 Plus (hosted 397B-A17B) | 13.6 | unverifiedT2 | 2026-02-16 |
| 33 | Gemini 3.1 Flash-Lite | 13 | unverifiedT2 | 2026-03-03 |
| 34 | Nemotron 3 Ultra | 11.5 | unverifiedT2 | 2026-06-04 |
| 35 | Claude Sonnet 4 | 9.3 | unverifiedT2 | 2025-05-22 |
| 36 | Claude Haiku 4.5 | 8.9 | unverifiedT2 | 2025-10-15 |
| 37 | GLM-4.7 | 8.7 | unverifiedT2 | 2025-12-22 |
| 38 | DeepSeek-V3.2 | 7 | unverifiedT2 | 2025-12-01 |
| 39 | Gemini 2.5 Pro (Jun 2025) | 6.6 | unverifiedT2 | 2025-06-05 |
| 40 | MiniMax-M2.5 | 6.2 | unverifiedT2 | 2026-02-12 |
| 41 | gpt-oss-120b | 4.7 | unverifiedT2 | 2025-08-05 |
| 42 | Kimi K2 Thinking | 4.1 | unverifiedT2 | 2025-11-06 |
| 43 | GLM-4.6 | 4 | unverifiedT2 | 2025-09-30 |
| 44 | Grok 3 | 2.1 | unverifiedT2 | 2025-02-17 |
| 45 | Gemini 2.5 Flash (Jun 2025) | 1.8 | unverifiedT2 | 2025-06-17 |
| 46 | o1 | 1.1 | unverifiedT2 | 2024-12-05 |
| 47 | GPT-4o (Nov 2024) | 1.1 | unverifiedT2 | 2024-05-13 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark evaluates 47 models, and the gap between its highest and lowest scores is 43.9 points. A spread of this size indicates that the observed score differences reflect genuine separation between models rather than noise.
3 cited facts
The benchmark's integrity score is 99, graded A, ranking 5th out of 61 in the Benchmark Integrity Index, indicating strong health as a discriminator of frontier models under a disclosed harness. Its weakest component is saturation, meaning the ceiling is crowded and top scores should be read with less confidence.
5 cited facts
This benchmark is not saturated, so the ranking at the top is not yet noise-dominated. The top score leaves 55 points of headroom below the ceiling, and only two models cluster within a couple of points of the top. Because it is not saturated, there is genuine room to separate the field at the top; small differences among leaders may still be meaningful.
5 cited facts
Scores are directly comparable within this benchmark because a consistent harness was used throughout. Test set privacy is unknown, so it is unclear whether the test set is public or held out; if it is public, high scores would warrant extra scrutiny for possible contamination, though contamination history is also unknown. In practice, head-to-head comparisons should only be made under this same disclosed harness, and results should be interpreted as task performance under that harness rather than as deployed capability.
4 cited facts
This benchmark's intended scope is not documented: what it measures is unknown, so the manual does not state what capability or behavior it evaluates. The construction procedure is also undocumented, so how the test was built or assembled cannot be described. Likewise, the benchmark records no key assumptions or limitations, leaving no documented caveat about where it might mislead or what it cannot tell you. Accordingly, the sharpest caveat a reader can hold is that this benchmark's validity and failure modes are entirely unspecified in the available record.
6 cited facts
APEX-Agents: APEX-Agents as reported in Epoch AI's Capabilities Index CSV.
Claude Fable 5 leads APEX-Agents at 45. The full leaderboard above lists every recorded measurement, not just the headline number.
47 models have recorded scores on APEX-Agents, spanning a score spread of 43.9.
tensor.news grades APEX-Agents A for integrity (score 99/100), ranking #5 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — APEX-Agents still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on APEX-Agents are reasonably apples-to-apples.
Every APEX-Agents measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.