released 2026-07-301 source
Inkling-Small benchmark scores: 11 benchmarks tracked. 64% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 89.99 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 89.99accuracy (%) | reproduced· optimizedT1 | 2026-07-30Epoch AI |
| GPQA diamond | 84.68accuracy (%) | reproduced· optimizedT1 | 2026-07-30Epoch AI |
| ARC-AGI | 84accuracy (%) | unverified· optimizedT1 | 2026-07-30https://arcprize.org/leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 46.32accuracy (%) | reproducedT1 | 2026-07-30Epoch AI |
| ARC-AGI-2 | 40.14accuracy (%) | unverifiedT2 | 2026-07-30 |
| SimpleQA Verified | 19.5accuracy (%) | reproducedT1 | 2026-07-30Epoch AI |
| FrontierMath-Tier-4-v2-Private | 17.07accuracy (%) | reproducedT1 | 2026-07-30Epoch AI |
| Chess Puzzles | 13.72accuracy (%) | reproducedT1 | 2026-07-30Epoch AI |
| CritPt | 8.29accuracy (%) | unverifiedT2 | 2026-07-30 |
| ProofBench | 6accuracy (%) | unverifiedT1 | 2026-07-30https://www.vals.ai/benchmarks/proof_bench |
| Mystery Game Puzzles | 0accuracy (%) | reproducedT1 | 2026-07-30Epoch AI |
Claims drawn from cited facts, not live model generation.
This model is evaluated across 11 tracked benchmarks, though it holds no current top score on any of them. Because the results have been independently reproduced, the record can be considered trustworthy rather than merely a vendor claim.
3 cited facts
This model sits well behind the front of the pack, trailing the leader by an average of 43.35 points on its benchmarks. It holds the top score on none of the benchmarks, so it shows no current category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Inkling-Small is an AI model released 2026-07-30. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Inkling-Small has recorded scores on 11 benchmarks, each shown with its evidence status.
Inkling-Small has recorded scores on 11 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleQA Verified, CritPt, ARC-AGI, and 5 more. The full table above shows each score with its evidence status.
7 of 11 recorded scores (64%) are independently reproduced rather than self-reported by the lab.
Inkling-Small has 11 tracked claims: 7 independently reproduced, 4 unverified.