GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Integrity rank #15 of 61 · 11 models scored · top score 49.7 · GPT-5.2
Ability to produce expert-quality professional work products across 44 occupations
Strongest on contamination resistance, weakest on saturation headroom.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-5.2 | 49.7 | unverifiedT2 | 2025-12-11 |
| 2 | Claude Opus 4.5 | 45.5 | unverifiedT2 | 2025-11-24 |
| 3 | Claude Opus 4.1 | 43.6 | unverifiedT2 | 2025-08-05 |
| 4 | Claude Sonnet 4.5 | 42.5 | unverifiedT2 | 2025-09-29 |
| 5 | Gemini 3 Pro | 40.3 | unverifiedT2 | 2025-11-18 |
| 6 | GPT-5 | 34.8 | unverifiedT2 | 2025-08-07 |
| 7 | o3 | 30.8 | unverifiedT2 | 2024-12-20 |
| 8 | o4-mini | 25.3 | unverifiedT2 | 2025-04-16 |
| 9 | Gemini 2.5 Pro (Jun 2025) | 23.3 | unverifiedT2 | 2025-06-05 |
| 10 | Grok 4 | 21.1 | unverifiedT2 | 2025-07-09 |
| 11 | GPT-4o (Nov 2024) | 9.9 | unverifiedT2 | 2024-05-13 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
Visible separation, thin base: the 39.8-point range across 11 models makes the gaps believable while leaving each rank resting on few entrants. The wide spread supports real separation between models, indicating that rank differences can be trusted.
3 cited facts
This benchmark achieves an Integrity Index of 98 with grade A, ranking 12 among 51 benchmarks, reflecting strong overall health as a discriminator of frontier models under its disclosed harness. Saturation is the weakest integrity component, meaning the performance ceiling is crowded and limits differentiation among top models.
4 cited facts
More scale remains above the leader (50.3 points) than below it (49.7) — this benchmark is far from done separating models, so read the current ranking as an early ordering, not a settled one. Only one model clusters near the top, which is within a couple of points, indicating limited variation at the leading edge.
3 cited facts
A score compressed from many occupations speaks to breadth, not to any particular profession, and should not be read as competence in the reader's own field. Building from real professional deliverables anchors the test in work that actually happens; the cost is that correctness becomes a matter of matching professional judgment rather than checking an answer. The design assumes that blind human preference over one-shot deliverables serves as a proxy for real economic task value. One-shot deliverables are the easy slice of professional work — the multi-day, iterative part is exactly what this test omits — so scores overstate readiness for real jobs, and say nothing outside white-collar work.
4 cited facts
GDPval (GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks): GDPval evaluates AI models on real-world, economically valuable deliverables (documents, slides, spreadsheets, diagrams, CAD) drawn from 44 occupations across the 9 largest U.S. GDP-contributing sectors; scored as the rate at which model outputs beat ('wins') or match ('ties') expert human deliverables under blind pairwise grading. Best model (Claude Opus 4.1) reached 47.6% wins+ties as of Sept 2025.
GPT-5.2 leads GDPval at 49.7. The full leaderboard above lists every recorded measurement, not just the headline number.
11 models have recorded scores on GDPval, spanning a score spread of 39.8.
tensor.news grades GDPval A for integrity (score 98/100), ranking #15 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — GDPval still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is semi-private, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on GDPval are reasonably apples-to-apples.
Every GDPval measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2510.04374.