Skip to content

Benchmark explainer

What is DTBench?

LLM capability on extracting structured tables from unstructured documents (PDF, HTML, mixed layouts).

DTBench as reported in Epoch AI's Capabilities Index CSV.

AINTEGRITY 99 / 100tensor.news
consistent harnesspublic test setsaturated

How it's scored

Metric
normalized performance in Epoch CSV, rendered as accuracy (%)
Score ceiling
100
Construction
Synthetic benchmark of document-to-table pairs; published at KDD 2026 with public GitHub dataset.
Human baseline
Human annotators benchmarked in paper §4; baseline accuracy reported as comparison.

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

97.9/100

How far from ceiling / clustered at the top?

Contamination resistance

99.3/100

Public vs held-out; training-leak risk.

Harness comparability

100/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

100/100

How old is the benchmark?

Contamination history: Synthetic, published 2026; some templates drawn from public document datasets (notable: authors flag specific overlap risks in §6).

Limitations: Synthetic generation may not match real-world document noise/distribution; LLM extraction accuracy is the only published metric — interpretability scoring absent.

Who leads DTBench

ModelScoreEvidence
Claude Fable 597.33unverified
Claude Opus 596unverified
Grok 4.695.55unverified
Gemini 3.7 Flash94.67unverified
Grok 4.594.22unverified

Compare the top DTBench scorers

Frequently asked questions

DTBench: DTBench as reported in Epoch AI's Capabilities Index CSV.

Claude Fable 5 leads DTBench at 97.33. The full leaderboard above lists every recorded measurement, not just the headline number.

142 models have recorded scores on DTBench, spanning a score spread of 94.58.

tensor.news grades DTBench A for integrity (score 99/100), ranking #5 of 62 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

Yes — top scores are clustering near the ceiling (saturation ratio 0.02), so DTBench no longer separates leading models well. Treat small gaps at the top with caution.

Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.

Scores are reported under a consistent harness, so comparisons on DTBench are reasonably apples-to-apples.

Every DTBench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2602.13812.

Follow the record