Surface Evolver Bench
Integrity rank #21 of 61 · 19 models scored · top score 95 · Kimi K3
unknown
Strongest on discrimination, weakest on saturation headroom. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Kimi K3 | 95 | unverifiedT1 | 2026-07-16 |
| 2 | Claude Fable 5 | 95 | unverifiedT1 | 2026-06-09 |
| 3 | GPT-5.6 Sol | 93.12 | unverifiedT1 | 2026-07-09 |
| 4 | GPT-5.5 | 88.12 | unverifiedT1 | 2026-04-23 |
| 5 | Claude Opus 4.8 | 87.5 | unverifiedT1 | 2026-05-28 |
| 6 | GPT-5.6 Terra | 83.75 | unverifiedT1 | 2026-07-09 |
| 7 | Grok 4.5 | 74.38 | unverifiedT1 | 2026-07-08 |
| 8 | Claude Sonnet 5 | 60 | unverifiedT1 | 2026-06-30 |
| 9 | GPT-5.6 Luna | 60 | unverifiedT1 | 2026-07-09 |
| 10 | Gemini 3.5 Flash | 58.13 | unverifiedT1 | 2026-05-19 |
| 11 | GLM-5.2 | 55.62 | unverifiedT1 | 2026-06-16 |
| 12 | MiniMax-M3 | 53.12 | unverifiedT1 | 2026-06-01 |
| 13 | Muse Spark 1.1 | 52.5 | unverifiedT1 | 2026-07-09 |
| 14 | Kimi K2.7 Code | 48.75 | unverifiedT1 | 2026-06-12 |
| 15 | Qwen 3.6 35B-A3B | 44.38 | unverifiedT1 | 2026-04-14 |
| 16 | DeepSeek-V4-Pro | 40 | unverifiedT1 | 2026-04-24 |
| 17 | Gemma 4 31B IT | 30.62 | unverifiedT1 | 2026-04-02 |
| 18 | Mistral Medium 3.5 | 26.88 | unverifiedT1 | 2026-04-29 |
| 19 | gpt-oss-120b | 25 | unverifiedT1 | 2025-08-05 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark evaluates 19 models, and the best score exceeds the worst by 70.0 points. Because the spread is wide, the observed rank differences are likely to reflect genuine capability separation rather than noise.
3 cited facts
The benchmark's integrity score is 96, yielding an integrity grade of A and ranking 21st among 61 benchmarks in the universe. The weakest component of the integrity breakdown is saturation, which costs the reader a crowded ceiling that blurs differences among top scores. This score tracks the benchmark's integrity as a discriminator of frontier models under a disclosed harness, rather than any model's capability.
6 cited facts
The benchmark is saturated, so differences among the leading models should be treated as noise rather than real capability gaps. The top score is 95.0, leaving 5.0 points of headroom to the ceiling, and three models cluster within a couple of points at the top, indicating little genuine separation at the top.
4 cited facts
Because the harness is consistent, task performance scores are directly comparable across runs under this same harness, while scores from different harnesses should not be compared head-to-head. The test set's privacy status and contamination history are both unknown, so no extra scrutiny for public exposure or prior contamination can be warranted; the disclosed harness notes contain only reporting metadata, not sensitivity indicators.
4 cited facts
This benchmark's methodology is not documented; what it measures is unspecified. How the test was built and its underlying assumptions are likewise not documented. The sharpest caveat is that its limitations and human-comparison baseline are unavailable, so it cannot show where it might mislead or how it relates to human performance.
5 cited facts
Surface Evolver Bench: Surface Evolver Bench as reported in Epoch AI's Capabilities Index CSV.
Kimi K3 leads Surface Evolver Bench at 95. The full leaderboard above lists every recorded measurement, not just the headline number.
19 models have recorded scores on Surface Evolver Bench, spanning a score spread of 70.
tensor.news grades Surface Evolver Bench A for integrity (score 96/100), ranking #21 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
Yes — top scores are clustering near the ceiling (saturation ratio 0.16), so Surface Evolver Bench no longer separates leading models well. Treat small gaps at the top with caution.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on Surface Evolver Bench are reasonably apples-to-apples.
Every Surface Evolver Bench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.