DTBench
Integrity rank #5 of 62 · 142 models scored · top score 97.33 · Claude Fable 5
LLM capability on extracting structured tables from unstructured documents (PDF, HTML, mixed layouts).
Strongest on discrimination, weakest on saturation headroom. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 97.33 | unverifiedT1 | 2026-06-09 |
| 2 | Claude Opus 5 | 96 | unverifiedT1 | 2026-07-24 |
| 3 | Grok 4.6 | 95.55 | unverifiedT1 | 2026-08-12 |
| 4 | Gemini 3.7 Flash | 94.67 | unverifiedT1 | 2026-08-13 |
| 5 | Grok 4.5 | 94.22 | unverifiedT1 | 2026-07-08 |
| 6 | GPT-5.5 | 93.33 | unverifiedT1 | 2026-04-23 |
| 7 | GPT-5.5 Pro | 93.33 | unverifiedT1 | 2026-04-23 |
| 8 | GPT-5.6 Sol | 93.33 | unverifiedT1 | 2026-07-09 |
| 9 | Gemini 3.1 Pro | 92.45 | unverifiedT1 | 2026-02-19 |
| 10 | Gemini 3.6 Flash | 92.45 | unverifiedT1 | 2026-07-21 |
| 11 | Claude Opus 4.8 | 91.55 | unverifiedT1 | 2026-05-28 |
| 12 | Gemini 3.5 Flash | 91.12 | unverifiedT1 | 2026-05-19 |
| 13 | Muse Spark 1.2 | 91.12 | unverifiedT1 | 2026-08-05 |
| 14 | Claude Opus 4.7 | 91.08 | unverifiedT1 | 2026-04-16 |
| 15 | Muse Spark 1.1 | 90.67 | unverifiedT1 | 2026-07-09 |
| 16 | GPT-5.4 | 90.67 | unverifiedT1 | 2026-03-05 |
| 17 | GLM-5.2 | 89.33 | unverifiedT1 | 2026-06-16 |
| 18 | GPT-5.6 Terra | 87.55 | unverifiedT1 | 2026-07-09 |
| 19 | Claude Sonnet 5 | 87.47 | unverifiedT1 | 2026-06-30 |
| 20 | Qwen3.7-Max | 87.12 | unverifiedT1 | 2026-05-19 |
| 21 | Qwen 3.8 Max | 86.67 | unverifiedT1 | 2026-08-02 |
| 22 | Kimi K3 | 85.33 | unverifiedT1 | 2026-07-16 |
| 23 | Claude Opus 4.6 | 85.33 | unverifiedT1 | 2026-02-05 |
| 24 | Kimi K2.6 | 84.88 | unverifiedT1 | 2026-04-20 |
| 25 | DeepSeek V4 Flash 0731 | 84.88 | unverifiedT1 | 2026-07-31 |
| 26 | GPT-5.2 | 84.88 | unverifiedT1 | 2025-12-11 |
| 27 | DeepSeek-V4-Pro | 84.45 | unverifiedT1 | 2026-04-24 |
| 28 | Grok 4.3 Beta | 84.45 | unverifiedT1 | 2026-04-17 |
| 29 | GPT-5 | 84.45 | unverifiedT1 | 2025-08-07 |
| 30 | Grok 4.20 | 83.55 | unverifiedT1 | 2026-02-17 |
| 31 | Nemotron 3 Ultra | 83.55 | unverifiedT1 | 2026-06-04 |
| 32 | GPT-5.1 | 83.55 | unverifiedT1 | 2025-11-13 |
| 33 | Claude Opus 4.5 | 83.12 | unverifiedT1 | 2025-11-24 |
| 34 | Claude Sonnet 4.6 | 83.12 | unverifiedT1 | 2026-02-17 |
| 35 | Gemini 3 Flash | 81.78 | unverifiedT1 | 2025-12-17 |
| 36 | GPT-5.6 Luna | 81.78 | unverifiedT1 | 2026-07-09 |
| 37 | DeepSeek-V3.2-Exp | 79.55 | unverifiedT1 | 2025-09-29 |
| 38 | Inkling | 79.12 | unverifiedT1 | 2026-07-15 |
| 39 | Qwen3.5 397B-A17B | 79.12 | unverifiedT1 | 2026-02-13 |
| 40 | Qwen 3.6 Max (Preview) | 78.67 | unverifiedT1 | 2026-04-20 |
| 41 | o3-pro | 78.22 | unverifiedT1 | 2025-06-10 |
| 42 | DeepSeek-V4-Flash | 77.33 | unverifiedT1 | 2026-04-24 |
| 43 | DeepSeek-V3.2 | 76 | unverifiedT1 | 2025-12-01 |
| 44 | o3 | 74.67 | unverifiedT1 | 2025-04-16 |
| 45 | Qwen3.7-Plus | 73.33 | unverifiedT1 | 2026-06-02 |
| 46 | Gemini 3.5 Flash-Lite | 72.45 | unverifiedT1 | 2026-07-21 |
| 47 | Claude Sonnet 4.5 | 72 | unverifiedT1 | 2025-09-29 |
| 48 | Qwen 3.5 Flash (hosted 35B-A3B) | 71.55 | unverifiedT1 | 2026-02-25 |
| 49 | Gemma 4 31B IT | 71.12 | unverifiedT1 | 2026-04-02 |
| 50 | Grok 4 Fast | 71.12 | unverifiedT1 | 2025-09-19 |
| 51 | DeepSeek-V3.1 | 71.12 | unverifiedT1 | 2025-08-21 |
| 52 | Gemini 2.5 Pro (Jun 2025) | 70.67 | unverifiedT1 | 2025-06-05 |
| 53 | Qwen3-Max | 70.22 | unverifiedT1 | 2025-09-24 |
| 54 | Qwen 3.6 Plus | 69.78 | unverifiedT1 | 2026-03-31 |
| 55 | Claude Opus 4 | 69.33 | unverifiedT1 | 2025-05-22 |
| 56 | Qwen 3.5 Plus (hosted 397B-A17B) | 67.55 | unverifiedT1 | 2026-02-16 |
| 57 | GPT-5 mini | 67.55 | unverifiedT1 | 2025-08-07 |
| 58 | Qwen3-235B-A22B-Thinking (Jul 2025) | 67.12 | unverifiedT1 | 2025-07-25 |
| 59 | GPT-5.4 Nano | 67.12 | unverifiedT1 | 2026-03-17 |
| 60 | Claude Opus 4.1 | 66.67 | unverifiedT1 | 2025-08-05 |
| 61 | Qwen3.5-35B-A3B | 66.67 | unverifiedT1 | 2026-02-24 |
| 62 | GPT-5.4 Mini | 66.67 | unverifiedT1 | 2026-03-17 |
| 63 | MiniMax-M3 | 64.88 | unverifiedT1 | 2026-06-01 |
| 64 | Qwen3-235B-A22B-Instruct (Jul 2025) | 64 | unverifiedT1 | 2025-07-25 |
| 65 | Qwen3.6 27B | 63.55 | unverifiedT1 | 2026-04-22 |
| 66 | o4-mini | 62.67 | unverifiedT1 | 2025-04-16 |
| 67 | Qwen 3.6 Flash | 61.78 | unverifiedT1 | 2026-04-27 |
| 68 | Claude Sonnet 4 | 61.78 | unverifiedT1 | 2025-05-22 |
| 69 | Gemini 3.1 Flash-Lite | 61.33 | unverifiedT1 | 2026-03-03 |
| 70 | Gemini 2.5 Flash (Jun 2025) | 60.88 | unverifiedT1 | 2025-06-17 |
| 71 | gpt-oss-120b | 60.45 | unverifiedT1 | 2025-08-05 |
| 72 | Qwen3-235B-A22B | 59.55 | unverifiedT1 | 2025-04-28 |
| 73 | Mistral Medium 3.5 | 59.12 | unverifiedT1 | 2026-04-28 |
| 74 | Gemma 4 26B A4B | 58.22 | unverifiedT1 | 2026-04-02 |
| 75 | o1 | 57.78 | unverifiedT1 | 2024-12-17 |
| 76 | Qwen 3.6 35B-A3B | 56.45 | unverifiedT1 | 2026-04-14 |
| 77 | Claude Haiku 4.5 | 56 | unverifiedT1 | 2025-10-15 |
| 78 | Qwen3.5-9B | 52 | unverifiedT1 | 2026-02-24 |
| 79 | Qwen3-30B-A3B-Thinking (Jul 2025) | 48.88 | unverifiedT1 | 2025-07-30 |
| 80 | o3-mini | 48 | unverifiedT1 | 2025-01-31 |
| 81 | GPT-4.1 mini | 48 | unverifiedT1 | 2025-04-14 |
| 82 | GPT-4.1 | 47.12 | unverifiedT1 | 2025-04-14 |
| 83 | gpt-oss-20b | 46.67 | unverifiedT1 | 2025-08-05 |
| 84 | Claude 3.5 Sonnet (October 2024) | 46.4 | unverifiedT1 | 2024-10-22 |
| 85 | Claude 3.5 Sonnet | 46.28 | unverifiedT1 | 2024-06-20 |
| 86 | Qwen3-32B | 45.78 | unverifiedT1 | 2025-04-28 |
| 87 | Qwen3-30B-A3B-Instruct (Jul 2025) | 45.33 | unverifiedT1 | 2025-07-29 |
| 88 | Grok-2 (Dec 2024) | 42.02 | unverifiedT1 | 2024-12-12 |
| 89 | DeepSeek-V3 (Mar 2025) | 41.33 | unverifiedT1 | 2025-03-24 |
| 90 | GPT-4o (May 2024) | 40.88 | unverifiedT1 | 2024-05-13 |
| 91 | Qwen3-14B | 40 | unverifiedT1 | 2025-04-29 |
| 92 | Gemini 2.0 Flash (Feb 2025) | 38.67 | unverifiedT1 | 2025-02-05 |
| 93 | Qwen2.5-72B | 38.22 | unverifiedT1 | 2024-09-19 |
| 94 | Gemini 2.5 Flash-Lite (Jun 2025) | 38.05 | unverifiedT1 | 2025-06-17 |
| 95 | GPT-4 (Jun 2023) | 37.78 | unverifiedT1 | 2023-06-13 |
| 96 | GPT-5 nano | 37.78 | unverifiedT1 | 2025-08-07 |
| 97 | Mistral Medium 3 | 37.12 | unverifiedT1 | 2025-05-07 |
| 98 | Llama 4 Maverick | 36.45 | unverifiedT1 | 2025-04-05 |
| 99 | GPT-4 Turbo (Nov 2023) | 36 | unverifiedT1 | 2023-11-06 |
| 100 | Claude 3 Opus | 36 | unverifiedT1 | 2024-02-29 |
| 101 | Llama 3.1-405B | 35.63 | unverifiedT1 | 2024-07-23 |
| 102 | Magistral Small 1.2 | 35.52 | unverifiedT1 | 2025-09-18 |
| 103 | Mistral Large 2 (Jul 2024) | 35.35 | unverifiedT1 | 2024-07-24 |
| 104 | Mistral Large 2 (Nov 2024) | 34.7 | unverifiedT1 | 2024-11-18 |
| 105 | Qwen3-30B-A3B | 33.78 | unverifiedT1 | 2025-04-28 |
| 106 | Llama 3.1-70B | 33.33 | unverifiedT1 | 2024-07-23 |
| 107 | Mistral Small 3.2 | 33.12 | unverifiedT1 | 2025-06-20 |
| 108 | Qwen3-8B | 32.88 | unverifiedT1 | 2025-04-28 |
| 109 | Llama 3.3 70B | 32.45 | unverifiedT1 | 2024-12-06 |
| 110 | Gemini 1.5 Pro (Sept 2024) | 31.63 | unverifiedT1 | 2024-09-24 |
| 111 | Mistral Small 3.1 | 31.07 | unverifiedT1 | 2025-03-17 |
| 112 | Llama 4 Scout | 29.78 | unverifiedT1 | 2025-04-05 |
| 113 | Claude 3.5 Haiku | 27.77 | unverifiedT1 | 2024-10-22 |
| 114 | Gemini 1.5 Pro (May 2024) | 27.45 | unverifiedT1 | 2024-05-14 |
| 115 | Mistral Large | 26.5 | unverifiedT1 | 2024-02-26 |
| 116 | Mixtral 8x22B | 25.23 | unverifiedT1 | 2024-04-17 |
| 117 | Command R+ | 24.88 | unverifiedT1 | 2024-08-30 |
| 118 | GPT-4o mini | 24 | unverifiedT1 | 2024-07-18 |
| 119 | Llama 3-70B | 23.6 | unverifiedT1 | 2024-04-18 |
| 120 | Gemini 1.5 Flash (Sep 2024) | 22.98 | unverifiedT1 | 2024-09-24 |
| 121 | Claude 3 Sonnet | 22.67 | unverifiedT1 | 2024-02-29 |
| 122 | Gemma 3 27B | 20.88 | unverifiedT1 | 2025-03-12 |
| 123 | GPT-4.1 nano | 20.88 | unverifiedT1 | 2025-04-14 |
| 124 | Claude 2 | 19.75 | unverifiedT1 | 2023-07-11 |
| 125 | Ministral 3B | 19.55 | unverifiedT1 | 2024-10-16 |
| 126 | Gemini 1.5 Flash (May 2024) | 18.97 | unverifiedT1 | 2024-05-23 |
| 127 | Claude 2.1 | 18.25 | unverifiedT1 | 2023-11-21 |
| 128 | Llama 3.1-8B | 18.22 | unverifiedT1 | 2024-07-23 |
| 129 | Gemma 3 4B | 18.22 | unverifiedT1 | 2025-03-12 |
| 130 | Claude 3 Haiku | 16.88 | unverifiedT1 | 2024-03-07 |
| 131 | Mixtral 8x7B | 16.02 | unverifiedT1 | 2023-12-11 |
| 132 | Gemma 3 12B | 14.67 | unverifiedT1 | 2025-03-12 |
| 133 | Mistral NeMo | 14.35 | unverifiedT1 | 2024-07-18 |
| 134 | GPT-3.5 Turbo (Jan 2024) | 14.22 | unverifiedT1 | 2024-01-25 |
| 135 | Gemma 2 27B | 13.33 | unverifiedT1 | 2024-06-24 |
| 136 | Qwen2.5-7B | 12.88 | unverifiedT1 | 2024-09-19 |
| 137 | Gemini 1.0 Pro | 9.77 | unverifiedT1 | 2023-12-13 |
| 138 | Claude Instant | 9.63 | unverifiedT1 | 2023-08-09 |
| 139 | Llama 3-8B | 6.55 | unverifiedT1 | 2024-04-18 |
| 140 | Mistral 7B v0.3 | 4.23 | unverifiedT1 | 2023-09-27 |
| 141 | Llama 2-13B | 3.73 | unverifiedT1 | 2023-07-18 |
| 142 | Llama 2-70B | 2.75 | unverifiedT1 | 2023-07-18 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 142 models, with a spread of 94.58 points between the best and worst results. That wide spread supports real separation between models, since close scores near the top are unlikely to reflect a genuine capability gap.
3 cited facts
This benchmark's Benchmark Integrity Index stands at 99, with a letter grade of A, ranking 5th among 62 benchmarks assessed. The weakest component of its integrity breakdown is saturation, which means the ceiling is crowded and top-end scores leave less room to separate leading models. The score should be read as the benchmark's health as a discriminator of frontier models under a disclosed harness, not as any claim about a model's capability.
7 cited facts
This benchmark is saturated, with a top score of 97.33 and only 2.67 points of headroom to the ceiling. Three models cluster near the top, within a couple of points of one another. That means rankings among the leading models are close to noise, so small differences there should not be read as real capability gaps.
5 cited facts
Reported scores for this benchmark come from a consistent harness, and the accompanying records disclose the model, model version, performance, score source, benchmark release date, and model date, with optimized-scaffold flags carried forward from prior exports, so results are comparable across models evaluated under that same setup. The privacy status of the test set is unknown, so it cannot be confirmed whether high scores rest on a public set that invites contamination scrutiny or on a held-out one, and the contamination history notes that some templates were drawn from public document datasets with overlap risks flagged by the authors. Because comparability holds only within the same harness and the test set's exposure is unresolved, scores should be read qualitatively as task performance under a disclosed harness rather than as evidence of deployed capability, with harness sensitivity and possible contamination treated as open considerations.
7 cited facts
This benchmark measures how well large language models extract structured tables from unstructured documents such as PDFs, HTML, and mixed layouts. It is built as a synthetic collection of document-to-table pairs, published publicly alongside its paper and dataset, resting on the assumption that synthetic pairs allow controllable difficulty and ground-truth evaluation at scale. Human annotators were also benchmarked in the paper as a comparison point for extraction accuracy. The sharpest caveat is that synthetic generation may not match real-world document noise or distribution, and because extraction accuracy is the only published metric, the benchmark cannot tell you anything about interpretability.
5 cited facts
DTBench: DTBench as reported in Epoch AI's Capabilities Index CSV.
Claude Fable 5 leads DTBench at 97.33. The full leaderboard above lists every recorded measurement, not just the headline number.
142 models have recorded scores on DTBench, spanning a score spread of 94.58.
tensor.news grades DTBench A for integrity (score 99/100), ranking #5 of 62 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
Yes — top scores are clustering near the ceiling (saturation ratio 0.02), so DTBench no longer separates leading models well. Treat small gaps at the top with caution.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on DTBench are reasonably apples-to-apples.
Every DTBench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2602.13812.