GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
Integrity rank #29 of 61 · 24 models scored · top score 47.06 · Claude Opus 4.8
SWE-agents' ability to write high-performance code optimizations
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Opus 4.8 | 47.06 | unverifiedT1 | 2026-05-28 |
| 2 | Claude Opus 4.7 | 44.12 | unverifiedT1 | 2026-04-16 |
| 3 | Claude Opus 4.6 | 41.2 | unverifiedT1 | 2026-02-05 |
| 4 | GPT-5.5 | 40.2 | unverifiedT1 | 2026-04-23 |
| 5 | Claude Sonnet 5 | 37.25 | unverifiedT1 | 2026-06-30 |
| 6 | GPT-5.4 | 31.37 | unverifiedT1 | 2026-03-05 |
| 7 | GPT-5.2 | 27.4 | unverifiedT1 | 2025-12-11 |
| 8 | Claude Opus 4.5 | 26.5 | unverifiedT1 | 2025-11-24 |
| 9 | Gemini 3.1 Pro | 22.55 | unverifiedT1 | 2026-02-19 |
| 10 | Gemini 3 Pro | 18.6 | unverifiedT1 | 2025-11-18 |
| 11 | Claude Sonnet 4.5 | 14.71 | unverifiedT1 | 2025-09-29 |
| 12 | GPT-5.1 | 13.73 | unverifiedT1 | 2025-11-13 |
| 13 | Gemini 3 Flash | 9.8 | unverifiedT1 | 2025-12-17 |
| 14 | o3 | 8.8 | unverifiedT1 | 2024-12-20 |
| 15 | Claude Opus 4 | 6.9 | unverifiedT1 | 2025-05-22 |
| 16 | GPT-5 | 6.9 | unverifiedT1 | 2025-08-07 |
| 17 | Kimi K2 (Jul 2025) | 4.9 | unverifiedT1 | 2025-07-11 |
| 18 | Claude Sonnet 4 | 4.9 | unverifiedT1 | 2025-05-22 |
| 19 | Claude 3.5 Sonnet (October 2024) | 4.6 | unverifiedT1 | 2024-10-22 |
| 20 | Gemini 2.5 Pro (Jun 2025) | 3.92 | unverifiedT1 | 2025-06-05 |
| 21 | Claude 3.7 Sonnet | 3.8 | unverifiedT1 | 2025-02-24 |
| 22 | o4-mini | 3.6 | unverifiedT1 | 2025-04-16 |
| 23 | o3-mini | 1.3 | unverifiedT1 | 2025-01-31 |
| 24 | GPT-4o (Nov 2024) | 0 | unverifiedT1 | 2024-05-13 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 24 models, and the spread between its best and worst scores is 47.06 points. That wide spread supports a real separation between model capabilities rather than narrow differences that could be indistinguishable from noise.
2 cited facts
The benchmark's integrity score is 89, corresponding to an A grade, and it ranks 29th out of 61 benchmarks. The weakest integrity component is contamination, which narrows how much to trust top scores.
4 cited facts
This benchmark is not saturated. The leading score is 47.06 with 52.94 points of headroom, and only one model clusters near the top within a couple of points. For the reader, the absence of saturation means today's top rankings still carry real signal, so differences at the top are not yet close to noise.
5 cited facts
Scores from this benchmark are produced under a consistent harness, so results are directly comparable within this setup; however, because the test set is public, high scores warrant extra scrutiny for possible contamination relative to a held-out test set. The harness is sensitive to hardware and noise due to runtime-based grading, and although the benchmark has a contamination history from public commits, the Nov-2025 Hack Detector rejects deceptive optimizations that memorization alone would not satisfy.
4 cited facts
This benchmark measures whether an agent can produce code optimizations that match an expert's performance, using tasks built from the commit history of several codebases. Each task pairs a codebase with performance tests, generated by an automated pipeline, and treats a human expert optimization commit as the success target. Its core assumption is that reproducing most of an expert's speedup while maintaining correctness is a valid proxy for real optimization skill, and that the performance tests precisely specify what is being optimized. The sharpest caveat is that near-floor performance scores on this small task set, combined with runtime-measurement noise, severely limit what it can tell you about an agent's real-world optimization ability.
6 cited facts
GSO-Bench (GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents): GSO (canonical name 'GSO'; the acronym is not expanded in the paper, titled 'Challenging Software Optimization Tasks for Evaluating SWE-Agents') tasks coding agents with modifying real-world codebases to achieve expert-level runtime speedups, verified by performance tests. Leading SWE-agents solve less than 5%.
Claude Opus 4.8 leads GSO-Bench at 47.06. The full leaderboard above lists every recorded measurement, not just the headline number.
24 models have recorded scores on GSO-Bench, spanning a score spread of 47.06.
tensor.news grades GSO-Bench A for integrity (score 89/100), ranking #29 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — GSO-Bench still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on GSO-Bench are reasonably apples-to-apples.
Every GSO-Bench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2505.23671.