Codeforces contest rating (model-estimated)
Integrity rank #20 of 61 · 8 models scored · top score 2029 · DeepSeek-R1
Competitive programming strength on a human-calibrated rating scale
Strongest on discrimination, weakest on saturation headroom. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | DeepSeek-R1 | 2029 | self-reported· optimizedT1 | 2025-01-22 |
| 2 | DeepSeek-R1-Distill-Qwen-32B | 1691 | self-reported· optimizedT1 | 2025-01-22 |
| 3 | DeepSeek-R1-Distill-Llama-70B | 1633 | self-reported· optimizedT1 | 2025-01-22 |
| 4 | DeepSeek-R1-Distill-Qwen-14B | 1481 | self-reported· optimizedT1 | 2025-01-22 |
| 5 | DeepSeek-R1-Distill-Llama-8B | 1205 | self-reported· optimizedT1 | 2025-01-22 |
| 6 | DeepSeek-R1-Distill-Qwen-7B | 1189 | self-reported· optimizedT1 | 2025-01-22 |
| 7 | DeepSeek-V3 | 1134 | self-reported· optimizedT1 | 2025-01-22 |
| 8 | DeepSeek-R1-Distill-Qwen-1.5B | 954 | self-reported· optimizedT1 | 2025-01-22 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark evaluates 8 models, resulting in a performance spread of 1075.0 points between the highest and lowest scores. This wide spread supports the reliability of rank differences, as the gap is large enough to indicate real separation between models rather than negligible capability variations.
3 cited facts
The benchmark achieves an integrity score of 96 (grade A), ranking 18 out of 58 evaluated benchmarks. This metric reflects the benchmark's health as a discriminator of frontier models under a disclosed harness. Saturation is the weakest component, meaning the ceiling is crowded.
5 cited facts
The benchmark is not currently saturated, with a top score of 2029.0. A single model clusters near the peak, leaving the upper tier well-separated from the rest of the field. This indicates there is still genuine room to separate the models at the top, so readers should not interpret minor score differences as definitive capability gaps.
3 cited facts
This benchmark employs a consistent harness, meaning scores are directly comparable across evaluations. Because the test set is public, high scores deserve more scrutiny for possible contamination than a held-out one. Qualitative assessments of harness sensitivity should note that these metrics reflect task performance under a disclosed harness rather than deployed capability.
3 cited facts
This benchmark measures competitive programming strength by mapping model performance onto a rating scale calibrated against human players. The test was constructed by estimating ratings from a specific sample of contests selected under a laboratory-defined protocol rather than tracking continuous play. Readers should note that relying on a limited contest sample introduces high variance, making single-point ratings potentially misleading indicators of consistent skill.
3 cited facts
Codeforces rating (Codeforces contest rating (model-estimated)): Codeforces rating as benchmark — model solutions to Codeforces contest problems converted into an Elo-style rating comparable to human competitor tiers; higher is better, no fixed ceiling.
DeepSeek-R1 leads Codeforces rating at 2029 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.
8 models have recorded scores on Codeforces rating, spanning a score spread of 1075.
tensor.news grades Codeforces rating A for integrity (score 96/100), ranking #20 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — Codeforces rating still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on Codeforces rating are reasonably apples-to-apples.
Every Codeforces rating measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites codeforces.com.