released 2026-07-151 source
Inkling benchmark scores: 12 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 12 benchmarks·best result 88.88 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 12 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 88.88accuracy (%) | reproduced· optimizedT1 | 2026-07-15Epoch AI |
| GPQA diamond | 84.34accuracy (%) | reproduced· optimizedT1 | 2026-07-15Epoch AI |
| ARC-AGI | 79.5accuracy (%) | unverified· optimizedT1 | 2026-07-15https://arcprize.org/leaderboard |
| SimpleQA Verified | 40.2accuracy (%) | reproducedT1 | 2026-07-15Epoch AI |
| ARC-AGI-2 | 36.53accuracy (%) | unverifiedT2 | 2026-07-15 |
| FrontierMath-Tiers-1-3-v2-Private | 33.33accuracy (%) | reproducedT1 | 2026-07-15Epoch AI |
| WeirdML | 32.29accuracy (%) | unverifiedT1 | 2026-07-15https://htihle.github.io/weirdml.html |
| Chess Puzzles | 16.88accuracy (%) | reproducedT1 | 2026-07-15Epoch AI |
| FrontierCode | 14accuracy (%) | unverifiedT1 | 2026-07-15https://cognition.com/frontiercode |
| CritPt | 5.43accuracy (%) | unverifiedT2 | 2026-07-15 |
| FrontierMath-Tier-4-v2-Private | 4.88accuracy (%) | reproducedT1 | 2026-07-15Epoch AI |
| ProofBench | 0accuracy (%) | unverifiedT1 | 2026-07-15https://www.vals.ai/benchmarks/proof_bench |
Claims drawn from cited facts, not live model generation.
This model is evaluated on 12 tracked benchmarks and currently holds no top score in any of them. Because those scores are independently reproduced by outside evaluation, the results can be treated as verified rather than merely vendor-claimed.
3 cited facts
With an average SOTA gap of 45.08 points, the model sits well behind the leaders on its benchmarks, a wide gap that signals it is not front-of-pack. It holds the top score on none of the tracked benchmarks, so it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Inkling is an AI model released 2026-07-15. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Inkling has recorded scores on 12 benchmarks, each shown with its evidence status.
Inkling has recorded scores on 12 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 6 more. The full table above shows each score with its evidence status.
6 of 12 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Inkling has 12 tracked claims: 6 independently reproduced, 6 unverified.