Benchmark explainer
What is SWE-bench Verified?
Resolve real GitHub issues so hidden repo tests pass
SWE-bench Verified — a human-validated subset of SWE-bench in which a model must resolve real GitHub issues from Python repositories end-to-end; scored as the percent of issues resolved.
How it's scored
- Metric
- % resolved (pass@1, test-verified)
- Score ceiling
- 100
- Construction
- 500 human-validated (OpenAI-filtered) tasks from 2294 SWE-bench items across 12 repos
- Human baseline
- N/A (real merged fix = 100% by construction)
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
97.4/100How far from ceiling / clustered at the top?
Contamination resistance
3.1/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
79.4/100How old is the benchmark?
Contamination history: Public GitHub PRs predate cutoffs; documented 'cheating' by reading future commit state
Limitations: Gameable w/ overfit patches; Python-only; score dominated by agent scaffold not model
Who leads SWE-bench Verified
| Model | Score | Evidence |
|---|---|---|
| Claude Opus 4.7 | 83.47 | reproduced |
| GPT-5.5 | 80.58 | reproduced |
| Gemini 3.5 Flash | 79.34 | reproduced |
| Claude Opus 4.6 | 78.72 | reproduced |
| GLM-5.2 | 78.7 | reproduced |
Frequently asked questions
SWE-Bench verified (SWE-bench Verified): SWE-bench Verified — a human-validated subset of SWE-bench in which a model must resolve real GitHub issues from Python repositories end-to-end; scored as the percent of issues resolved.
Claude Opus 4.7 leads SWE-Bench verified at 83.47 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
32 models have recorded scores on SWE-Bench verified, spanning a score spread of 52.48.
tensor.news grades SWE-Bench verified B for integrity (score 83/100), ranking #37 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — SWE-Bench verified still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on SWE-Bench verified are reasonably apples-to-apples.
Every SWE-Bench verified measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites openai.com/index/introducing-swe-bench-verified.