Skip to content

Benchmark explainer

What is SWE-bench Verified?

Resolve real GitHub issues so hidden repo tests pass

SWE-bench Verified — a human-validated subset of SWE-bench in which a model must resolve real GitHub issues from Python repositories end-to-end; scored as the percent of issues resolved.

BINTEGRITY 83 / 100tensor.news
consistent harnesspublic test set

How it's scored

Metric
% resolved (pass@1, test-verified)
Score ceiling
100
Construction
500 human-validated (OpenAI-filtered) tasks from 2294 SWE-bench items across 12 repos
Human baseline
N/A (real merged fix = 100% by construction)

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

97.4/100

How far from ceiling / clustered at the top?

Contamination resistance

3.1/100

Public vs held-out; training-leak risk.

Harness comparability

100/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

79.4/100

How old is the benchmark?

Contamination history: Public GitHub PRs predate cutoffs; documented 'cheating' by reading future commit state

Limitations: Gameable w/ overfit patches; Python-only; score dominated by agent scaffold not model

Who leads SWE-bench Verified

ModelScoreEvidence
Claude Opus 4.783.47reproduced
GPT-5.580.58reproduced
Gemini 3.5 Flash79.34reproduced
Claude Opus 4.678.72reproduced
GLM-5.278.7reproduced

Frequently asked questions

SWE-Bench verified (SWE-bench Verified): SWE-bench Verified — a human-validated subset of SWE-bench in which a model must resolve real GitHub issues from Python repositories end-to-end; scored as the percent of issues resolved.

Claude Opus 4.7 leads SWE-Bench verified at 83.47 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.

32 models have recorded scores on SWE-Bench verified, spanning a score spread of 52.48.

tensor.news grades SWE-Bench verified B for integrity (score 83/100), ranking #37 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

No — SWE-Bench verified still has headroom and continues to discriminate between models rather than bunching them at the ceiling.

Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.

Scores are reported under a consistent harness, so comparisons on SWE-Bench verified are reasonably apples-to-apples.

Every SWE-Bench verified measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites openai.com/index/introducing-swe-bench-verified.

Follow the record