Benchmark explainer
What is DTBench?
LLM capability on extracting structured tables from unstructured documents (PDF, HTML, mixed layouts).
DTBench as reported in Epoch AI's Capabilities Index CSV.
How it's scored
- Metric
- normalized performance in Epoch CSV, rendered as accuracy (%)
- Score ceiling
- 100
- Construction
- Synthetic benchmark of document-to-table pairs; published at KDD 2026 with public GitHub dataset.
- Human baseline
- Human annotators benchmarked in paper §4; baseline accuracy reported as comparison.
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
97.9/100How far from ceiling / clustered at the top?
Contamination resistance
99.3/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
100/100How old is the benchmark?
Contamination history: Synthetic, published 2026; some templates drawn from public document datasets (notable: authors flag specific overlap risks in §6).
Limitations: Synthetic generation may not match real-world document noise/distribution; LLM extraction accuracy is the only published metric — interpretability scoring absent.
Who leads DTBench
| Model | Score | Evidence |
|---|---|---|
| Claude Fable 5 | 97.33 | unverified |
| Claude Opus 5 | 96 | unverified |
| Grok 4.6 | 95.55 | unverified |
| Gemini 3.7 Flash | 94.67 | unverified |
| Grok 4.5 | 94.22 | unverified |
Compare the top DTBench scorers
Frequently asked questions
DTBench: DTBench as reported in Epoch AI's Capabilities Index CSV.
Claude Fable 5 leads DTBench at 97.33. The full leaderboard above lists every recorded measurement, not just the headline number.
142 models have recorded scores on DTBench, spanning a score spread of 94.58.
tensor.news grades DTBench A for integrity (score 99/100), ranking #5 of 62 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
Yes — top scores are clustering near the ceiling (saturation ratio 0.02), so DTBench no longer separates leading models well. Treat small gaps at the top with caution.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on DTBench are reasonably apples-to-apples.
Every DTBench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2602.13812.