PostTrainBench
Integrity rank #22 of 61 · 25 models scored · top score 41.79 · Claude Fable 5
Autonomous agentic post-training/fine-tuning skill under a fixed compute budget
Strongest on contamination resistance, weakest on discrimination. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 41.79 | unverifiedT2 | 2026-06-09 |
| 2 | GPT-5.6 Sol | 36.23 | unverifiedT2 | 2026-07-09 |
| 3 | GLM-5.2 | 34.29 | unverifiedT2 | 2026-06-16 |
| 4 | Claude Opus 4.8 | 34.08 | unverifiedT2 | 2026-05-28 |
| 5 | Claude Opus 5 | 34.06 | unverifiedT2 | 2026-07-24 |
| 6 | Kimi K3 | 31.96 | unverifiedT2 | 2026-07-16 |
| 7 | Claude Opus 4.7 | 28.56 | unverifiedT2 | 2026-04-16 |
| 8 | GPT-5.5 | 25.02 | unverifiedT2 | 2026-04-23 |
| 9 | Claude Opus 4.6 | 24.82 | unverifiedT2 | 2026-02-05 |
| 10 | Grok 4.5 | 23.45 | unverifiedT2 | 2026-07-08 |
| 11 | Gemini 3.1 Pro | 21.59 | unverifiedT2 | 2026-02-19 |
| 12 | GPT-5.2 | 21.38 | unverifiedT2 | 2025-12-11 |
| 13 | GPT-5.4 | 20.23 | unverifiedT2 | 2026-03-05 |
| 14 | GPT-5.1-Codex-Max | 19.68 | unverifiedT2 | 2025-11-19 |
| 15 | Gemini 3 Pro | 18.12 | unverifiedT2 | 2025-11-18 |
| 16 | GPT-5.3 Codex | 17.76 | unverifiedT2 | 2026-02-05 |
| 17 | Claude Opus 4.5 | 17.29 | unverifiedT2 | 2025-11-24 |
| 18 | Claude Sonnet 4.6 | 16.42 | unverifiedT2 | 2026-02-17 |
| 19 | GLM-5 | 13.88 | unverifiedT2 | 2026-02-11 |
| 20 | Kimi K2.5 | 10.26 | unverifiedT2 | 2026-02-02 |
| 21 | Claude Sonnet 4.5 | 9.94 | unverifiedT2 | 2025-09-29 |
| 22 | MiniMax-M2.5 | 9.5 | unverifiedT2 | 2026-02-12 |
| 23 | GLM-4.7 | 7.48 | unverifiedT2 | 2025-12-22 |
| 24 | Qwen3-Max | 7.42 | unverifiedT2 | 2025-09-05 |
| 25 | Kimi K2 Thinking | 7.25 | unverifiedT2 | 2025-11-06 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark evaluates 25 models, with scores spanning 34.54 points from best to worst. A spread of this size indicates that score differences between models are likely to reflect genuine capability gaps rather than noise.
3 cited facts
The benchmark's integrity score is 95, earning an A grade and ranking 22nd out of 61 benchmarks. Its weakest integrity component is discrimination, which narrows how distinctly the benchmark can separate frontier models.
5 cited facts
This benchmark is not saturated, as the leading score leaves substantial headroom before the ceiling. Only a single model occupies the top band, so the field is not compressed near the maximum. For the reader, this means differences among top models reflect genuine capability rather than noise.
5 cited facts
Because the harness used for this benchmark is consistent, task performance scores are directly comparable across submissions under that harness. The test set is public rather than held out, so high scores deserve more scrutiny for possible contamination than they would if the set were private, especially given the benchmark's public nature and potential for agent overfitting. Performance is highly sensitive to the agent scaffold, so results reflect task performance under the disclosed harness rather than any deployed capability.
4 cited facts
This benchmark measures an autonomous agent's skill at post-training or fine-tuning a language model within a fixed compute budget, treating that capability as the object of evaluation. This benchmark is built by giving the agent a single GPU, a limited time budget, and free choice of data and post-training technique, then scoring its output on a fixed set of downstream benchmarks. Its core assumptions are that improvement on those fixed downstream benchmarks under a fixed budget proxies general post-training competence and that the agent's chosen methods will generalize. The sharpest caveat is that the setup uses small models and a single GPU with a short budget, so this benchmark's results may not reflect frontier post-training; moreover, the public target benchmarks are partly gameable and the outcome is heavily dependent on scaffolding, while any human baseline is not documented.
5 cited facts
PostTrainBench: PostTrainBench tests a CLI coding agent's ability to autonomously post-train small base LLMs: given a 1-4B base model, a single H100, and a 10-hour window, the agent must improve the model across 7 downstream benchmarks using post-training techniques of its own choosing. Best submission approx 34.1% average.
Claude Fable 5 leads PostTrainBench at 41.79. The full leaderboard above lists every recorded measurement, not just the headline number.
25 models have recorded scores on PostTrainBench, spanning a score spread of 34.54.
tensor.news grades PostTrainBench A for integrity (score 95/100), ranking #22 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — PostTrainBench still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on PostTrainBench are reasonably apples-to-apples.
Every PostTrainBench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites posttrainbench.com.