METR Time Horizons
Integrity rank #26 of 61 · 37 models scored · top score 78.86 · Claude Opus 4.6
unknown
Strongest on discrimination, weakest on harness comparability. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Opus 4.6 | 78.86 | unverifiedT2 | 2026-02-05 |
| 2 | Gemini 3.1 Pro | 77.03 | self-reportedT1 | 2026-02-19 |
| 3 | GPT-5.2 | 75.29 | unverifiedT2 | 2025-12-11 |
| 4 | Claude Opus 4.5 | 74.97 | unverifiedT1 | 2025-11-24 |
| 5 | GPT-5.3 Codex | 74.54 | unverifiedT2 | 2026-02-05 |
| 6 | GPT-5.4 | 74.34 | self-reportedT1 | 2026-03-05 |
| 7 | GPT-5.1-Codex-Max | 71.58 | unverifiedT1 | 2025-11-19 |
| 8 | Gemini 3 Pro | 70.98 | unverifiedT2 | 2025-11-18 |
| 9 | GPT-5 | 69.61 | unverifiedT1 | 2025-08-07 |
| 10 | Claude Sonnet 4.5 | 67.38 | unverifiedT1 | 2025-09-29 |
| 11 | Claude Opus 4.1 | 66.81 | unverifiedT1 | 2025-08-05 |
| 12 | Grok 4 | 66.58 | unverifiedT1 | 2025-07-09 |
| 13 | o3 | 65.44 | unverifiedT1 | 2024-12-20 |
| 14 | Claude Opus 4 | 63.93 | unverifiedT1 | 2025-05-22 |
| 15 | o4-mini | 63.93 | unverifiedT1 | 2025-04-16 |
| 16 | Claude Sonnet 4 | 61.97 | unverifiedT1 | 2025-05-22 |
| 17 | Claude 3.7 Sonnet | 59.97 | unverifiedT1 | 2025-02-24 |
| 18 | Kimi K2 Thinking | 59.19 | unverifiedT1 | 2025-11-06 |
| 19 | gpt-oss-120b | 56.63 | unverifiedT1 | 2025-08-05 |
| 20 | o1 | 55.93 | unverifiedT1 | 2024-12-05 |
| 21 | Gemini 2.5 Pro (Jun 2025) | 55.44 | unverifiedT1 | 2025-06-05 |
| 22 | DeepSeek-R1 (May 2025) | 53.78 | unverifiedT1 | 2025-05-28 |
| 23 | Claude 3.5 Sonnet (October 2024) | 52.7 | unverifiedT1 | 2024-10-22 |
| 24 | DeepSeek-R1 | 51.93 | unverifiedT1 | 2025-01-20 |
| 25 | DeepSeek-V3 (Mar 2025) | 49.58 | unverifiedT1 | 2025-03-24 |
| 26 | o1-preview | 49.3 | unverifiedT1 | 2024-09-12 |
| 27 | Claude 3.5 Sonnet | 47.68 | unverifiedT1 | 2024-06-20 |
| 28 | DeepSeek-V3 | 47.36 | unverifiedT1 | 2024-12-24 |
| 29 | GPT-4o (Nov 2024) | 40.76 | unverifiedT1 | 2024-05-13 |
| 30 | GPT-4 Turbo (Nov 2023) | 40.43 | unverifiedT1 | 2023-11-06 |
| 31 | Claude 3 Opus | 37.75 | unverifiedT1 | 2024-03-04 |
| 32 | GPT-4 Turbo (Apr 2024) | 36.74 | unverifiedT1 | 2024-04-09 |
| 33 | GPT-4 (Mar 2023) | 36.11 | unverifiedT1 | 2023-03-15 |
| 34 | Qwen2.5-72B | 35.78 | unverifiedT1 | 2024-09-19 |
| 35 | GPT-4o (Aug 2024) | 33.84 | unverifiedT2 | 2024-05-13 |
| 36 | Qwen2-72B | 29.9 | unverifiedT1 | 2024-06-07 |
| 37 | GPT-4 (Jun 2023) | 29.3 | unverifiedT2 | 2023-06-13 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark evaluates 37 models, and the gap between its best and worst scores is 49.56 points. That spread is wide, meaning observed score differences likely reflect genuine capability separation rather than mere noise.
3 cited facts
This benchmark's Benchmark Integrity Index is 91 (grade A), ranking 26th among 61 benchmarks. Its weakest integrity component is harness, and that harness variance complicates cross-model comparisons. Taken together, the score measures the benchmark's health as a discriminator of frontier models under a disclosed harness, not any model's capability.
6 cited facts
This benchmark is not saturated, with a top score of 78.86 points and 21.14 points of headroom remaining to the ceiling. Only a couple of models sit within a couple of points of the top, so there is still genuine room to separate the field at the top. For the reader, this means the current ranking among top models is not yet close to noise, and small differences are not the whole story while the headroom still allows real capability gaps to emerge.
6 cited facts
Harness comparability is mixed and not directly comparable, so task performance scores from different harnesses should not be compared head-to-head. Test-set privacy is unknown; if the test set for this benchmark is public, high scores would warrant more scrutiny for possible contamination than a held-out set would, but this cannot be confirmed. Contamination history is unknown and the harness notes disclose only metadata fields, leaving sensitivity to harness configuration and contamination unassessed; scores should therefore be interpreted as task performance under the disclosed harness, not deployed capability.
4 cited facts
This benchmark's intended measurement target is not documented in the available facts. Its construction process, stated key assumptions, and any human baseline are likewise not documented, so the test's design rationale cannot be reconstructed from this record. The sharpest caveat is that without documented limitations or assumptions, this benchmark's outputs should not be interpreted as evidence of any specific capability or failure mode.
6 cited facts
METR Time Horizons: METR Time Horizons as reported in Epoch AI's Capabilities Index CSV.
Claude Opus 4.6 leads METR Time Horizons at 78.86. The full leaderboard above lists every recorded measurement, not just the headline number.
37 models have recorded scores on METR Time Horizons, spanning a score spread of 49.56.
tensor.news grades METR Time Horizons A for integrity (score 91/100), ranking #26 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — METR Time Horizons still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on METR Time Horizons is not fully apples-to-apples.
Every METR Time Horizons measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.