Terminal Bench
Integrity rank #36 of 61 · 36 models scored · top score 84.7 · GPT-5.5
unknown
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-5.5 | 84.7 | unverified· optimizedT1 | 2026-04-23 |
| 2 | GPT-5.4 | 81.8 | unverified· optimizedT1 | 2026-03-05 |
| 3 | Gemini 3.1 Pro | 80.2 | unverified· optimizedT1 | 2026-02-19 |
| 4 | Claude Opus 4.7 | 80.2 | unverified· optimizedT1 | 2026-04-16 |
| 5 | Claude Opus 4.6 | 79.8 | unverified· optimizedT1 | 2026-02-05 |
| 6 | GPT-5.3 Codex | 78.4 | unverified· optimizedT1 | 2026-02-05 |
| 7 | Gemini 3 Pro | 69.4 | unverified· optimizedT1 | 2025-11-18 |
| 8 | GPT-5.2 | 64.9 | unverified· optimizedT1 | 2025-12-11 |
| 9 | Gemini 3 Flash | 64.3 | unverified· optimizedT1 | 2025-12-17 |
| 10 | Claude Opus 4.5 | 63.1 | unverified· optimizedT1 | 2025-11-24 |
| 11 | GPT-5.1-Codex-Max | 60.4 | unverified· optimizedT1 | 2025-11-19 |
| 12 | Grok 4.20 | 57.3 | unverified· optimizedT1 | 2026-02-17 |
| 13 | Claude Sonnet 4.6 | 53.4 | unverified· optimizedT1 | 2026-02-17 |
| 14 | GLM-5 | 52.4 | unverified· optimizedT2 | 2026-02-11 |
| 15 | GPT-5 | 49.6 | unverified· optimizedT1 | 2025-08-07 |
| 16 | GPT-5.1 | 47.6 | unverified· optimizedT1 | 2025-11-13 |
| 17 | Claude Sonnet 4.5 | 46.5 | unverified· optimizedT1 | 2025-09-29 |
| 18 | MiniMax-M2.7 | 45.1 | unverified· optimizedT1 | 2026-03-18 |
| 19 | Kimi K2.5 | 43.2 | unverified· optimizedT2 | 2026-02-02 |
| 20 | MiniMax-M2.5 | 42.7 | unverified· optimizedT1 | 2026-02-12 |
| 21 | DeepSeek-V3.2 | 39.6 | unverified· optimizedT2 | 2025-12-01 |
| 22 | Claude Opus 4.1 | 38 | unverified· optimizedT1 | 2025-08-05 |
| 23 | Kimi K2 Thinking | 35.7 | unverified· optimizedT1 | 2025-11-06 |
| 24 | Claude Haiku 4.5 | 35.5 | unverified· optimizedT1 | 2025-10-15 |
| 25 | GPT-5 mini | 34.8 | unverified· optimizedT1 | 2025-08-07 |
| 26 | GLM-4.7 | 33.4 | unverified· optimizedT2 | 2025-12-22 |
| 27 | Gemini 2.5 Pro (Jun 2025) | 32.6 | unverified· optimizedT1 | 2025-06-05 |
| 28 | Kimi K2 (Jul 2025) | 27.8 | unverified· optimizedT1 | 2025-07-11 |
| 29 | Grok 4 | 27.2 | unverified· optimizedT1 | 2025-07-09 |
| 30 | Qwen 3.6 35B-A3B | 24.6 | unverified· optimizedT1 | 2026-04-14 |
| 31 | GLM-4.6 | 24.5 | unverified· optimizedT1 | 2025-09-30 |
| 32 | GPT-5 nano | 21.8 | unverified· optimizedT1 | 2025-08-07 |
| 33 | gpt-oss-120b | 18.7 | unverified· optimizedT1 | 2025-08-05 |
| 34 | Gemini 2.5 Flash (Jun 2025) | 17.1 | unverified· optimizedT2 | 2025-06-17 |
| 35 | Gemini 2.5 Flash (Sep 2025) | 17.1 | unverified· optimizedT1 | 2025-09-25 |
| 36 | gpt-oss-20b | 3.4 | unverified· optimizedT1 | 2025-08-05 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark evaluates 36 models, providing a broad basis for comparison. The 81.3-point spread between the highest and lowest scores is wide, so rank differences reflect genuine capability gaps rather than noise.
2 cited facts
The benchmark's integrity score is 84, a B grade, indicating its health as a discriminator of frontier models under a disclosed harness. This integrity rank places it 36th among the 61 benchmarks in the universe. Contamination is the weakest component of the integrity breakdown, costing readers reduced trust in top scores.
5 cited facts
The benchmark is not saturated: the top score is 84.7, leaving 15.3 points of headroom to the ceiling. Only a single model sits near the top, within a couple of points of the leading result. Because the field is not saturated, small differences among today's top models still reflect genuine capability gaps rather than noise.
5 cited facts
Scores from this benchmark represent task performance under a disclosed harness, and because the harness is consistent, results are directly comparable across reported evaluations. The privacy of this benchmark's test set is unknown, so if it is public, high scores warrant closer scrutiny for possible contamination than they would if the set were held out. Contamination history is unknown, and the harness notes only describe metadata fields in the reporting format, so any sensitivity of task performance to harness specifics cannot be assessed from available facts.
4 cited facts
This benchmark's methodology is not documented: what it measures is unspecified, so its intended scope is unknown. Its construction is also undocumented, meaning how the test items were built cannot be assessed, and its key assumptions are not recorded either. The benchmark's limitations are likewise undocumented, and no human baseline is provided, so the sharpest caveat is that no caveat can be stated from the available facts—readers cannot determine where this benchmark is most likely to mislead.
5 cited facts
Terminal Bench: Terminal Bench as reported in Epoch AI's Capabilities Index CSV.
GPT-5.5 leads Terminal Bench at 84.7. The full leaderboard above lists every recorded measurement, not just the headline number.
36 models have recorded scores on Terminal Bench, spanning a score spread of 81.3.
tensor.news grades Terminal Bench B for integrity (score 84/100), ranking #36 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — Terminal Bench still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on Terminal Bench are reasonably apples-to-apples.
Every Terminal Bench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.