SWE-bench Verified
Integrity rank #37 of 61 · 32 models scored · top score 83.47 · Claude Opus 4.7
Resolve real GitHub issues so hidden repo tests pass
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Opus 4.7 | 83.47 | reproduced· optimizedT1 | 2026-04-16 |
| 2 | GPT-5.5 | 80.58 | reproduced· optimizedT1 | 2026-04-23 |
| 3 | Gemini 3.5 Flash | 79.34 | reproduced· optimizedT1 | 2026-05-19 |
| 4 | Claude Opus 4.6 | 78.72 | reproduced· optimizedT1 | 2026-02-05 |
| 5 | GLM-5.2 | 78.7 | reproduced· optimizedT1 | 2026-06-16 |
| 6 | DeepSeek-V4-Pro | 77.64 | reproduced· optimizedT1 | 2026-04-24 |
| 7 | Qwen3.7-Max | 77.27 | reproduced· optimizedT1 | 2026-05-19 |
| 8 | GPT-5.4 | 76.86 | reproduced· optimizedT1 | 2026-03-05 |
| 9 | Kimi K2.6 | 76.65 | reproduced· optimizedT1 | 2026-04-20 |
| 10 | Qwen 3.6 Max (Preview) | 76.65 | reproduced· optimizedT1 | 2026-04-20 |
| 11 | Claude Opus 4.5 | 76.65 | reproduced· optimizedT1 | 2025-11-24 |
| 12 | Gemini 3.1 Pro | 75.62 | reproduced· optimizedT1 | 2026-02-19 |
| 13 | Gemini 3 Flash | 75.41 | reproduced· optimizedT1 | 2025-12-17 |
| 14 | Claude Sonnet 4.6 | 75.21 | reproduced· optimizedT1 | 2026-02-17 |
| 15 | GPT-5.3 Codex | 74.79 | reproduced· optimizedT1 | 2026-02-05 |
| 16 | GLM-5.1 | 74.17 | reproduced· optimizedT1 | 2026-04-07 |
| 17 | Kimi K2.5 | 73.76 | reproduced· optimizedT1 | 2026-02-02 |
| 18 | GPT-5.2 | 73.76 | reproduced· optimizedT1 | 2025-12-11 |
| 19 | GPT-5 | 73.55 | reproduced· optimizedT1 | 2025-08-07 |
| 20 | Claude Opus 4.1 | 73.35 | reproduced· optimizedT1 | 2025-08-05 |
| 21 | Gemini 3 Pro | 72.93 | reproduced· optimizedT1 | 2025-11-18 |
| 22 | GLM-5 | 72.08 | reproduced· optimizedT1 | 2026-02-11 |
| 23 | Claude Sonnet 4.5 | 71.28 | reproduced· optimizedT1 | 2025-09-29 |
| 24 | Claude Opus 4 | 70.66 | reproduced· optimizedT1 | 2025-05-22 |
| 25 | GPT-5.1 | 67.98 | reproduced· optimizedT1 | 2025-11-13 |
| 26 | GPT-5 mini | 64.67 | reproduced· optimizedT1 | 2025-08-07 |
| 27 | o3 | 62.32 | reproduced· optimizedT1 | 2024-12-20 |
| 28 | Claude 3.7 Sonnet | 60.95 | reproduced· optimizedT1 | 2025-02-24 |
| 29 | Qwen 3.6 Plus | 57.85 | reproduced· optimizedT1 | 2026-04-01 |
| 30 | Gemini 2.5 Pro (Jun 2025) | 57.56 | reproduced· optimizedT1 | 2025-06-05 |
| 31 | GPT-4.1 | 48.54 | reproduced· optimizedT1 | 2025-04-14 |
| 32 | GPT-4o (Nov 2024) | 30.99 | reproduced· optimizedT1 | 2024-05-13 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
For the 32 models scored, a 52.48-point spread is enough separation to make large gaps trustworthy, while near-ties still deserve skepticism. This wide spread supports real separation between models, meaning the score differences likely reflect genuine capability gaps.
3 cited facts
The benchmark's Integrity Index of 83 (grade B) reflects its health as a discriminator of frontier models under its disclosed harness, ranking 28th among 51 benchmarks. The weakest integrity component is contamination, which narrows how much to trust top scores.
4 cited facts
This benchmark is not saturated, meaning there is still genuine room to separate the field at the top. 16.53 points of ceiling above the 83.47 top score keep this benchmark discriminating at the frontier — a lead can still stretch or vanish within the measured range, so top-of-table movement remains informative. Only one model clusters near the top, so small differences in ranking are meaningful.
4 cited facts
Under this benchmark's consistent harness, task performance scores are directly comparable, while the public test set means high scores warrant additional scrutiny for possible contamination. The harness exhibits high sensitivity—the same model can experience large performance swings across different scaffolds—indicating a core integrity problem, and contamination has been documented via public PRs that predate cutoffs, including instances of cheating by reading future commit state.
4 cited facts
Real GitHub issues anchor the task in work that actually happened. What the score certifies, though, is passing hidden tests — adjacent to, not identical with, resolving the issue a maintainer had. The tasks were derived from an earlier benchmark, validated by human reviewers and an AI filter. Hidden tests are the arbiter, and tests are an approximation: a patch can pass while missing the issue's actual intent, so the pass rate is an upper bound on true resolution. Scaffolding sensitivity is the quiet caveat: the same model under different agent setups can post different numbers, which alongside its limit to Python and test-overfitting risk means scores travel poorly outside their disclosed setup.
4 cited facts
SWE-Bench verified (SWE-bench Verified): SWE-bench Verified — a human-validated subset of SWE-bench in which a model must resolve real GitHub issues from Python repositories end-to-end; scored as the percent of issues resolved.
Claude Opus 4.7 leads SWE-Bench verified at 83.47 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
32 models have recorded scores on SWE-Bench verified, spanning a score spread of 52.48.
tensor.news grades SWE-Bench verified B for integrity (score 83/100), ranking #37 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — SWE-Bench verified still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on SWE-Bench verified are reasonably apples-to-apples.
Every SWE-Bench verified measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites openai.com/index/introducing-swe-bench-verified.