Text-to-parametric-CAD generation and geometric fidelity
Strongest on discrimination, weakest on freshness. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | o3 | 74 | unverifiedT1 | 2024-12-20 |
| 2 | Gemini 2.5 Pro (Mar 2025) | 64 | unverifiedT1 | 2025-03-25 |
| 3 | o4-mini | 62 | unverifiedT1 | 2025-04-16 |
| 4 | o1 | 56 | unverifiedT1 | 2024-12-05 |
| 5 | Claude 3.7 Sonnet | 54 | unverifiedT1 | 2025-02-24 |
| 6 | o3-mini | 54 | unverifiedT1 | 2025-01-31 |
| 7 | Claude 3.5 Sonnet (October 2024) | 48 | unverifiedT1 | 2024-10-22 |
| 8 | GPT-4.1 | 42 | unverifiedT1 | 2025-04-14 |
| 9 | Gemini 1.5 Pro (Sept 2024) | 34 | unverifiedT1 | 2024-09-24 |
| 10 | Claude 3.5 Haiku | 32 | unverifiedT1 | 2024-10-22 |
| 11 | Gemini 2.0 Flash (Feb 2025) | 30 | unverifiedT1 | 2024-12-11 |
| 12 | GPT-4o (Aug 2024) | 26 | unverifiedT1 | 2024-05-13 |
| 13 | GPT-4.1 mini | 16 | unverifiedT1 | 2025-04-14 |
| 14 | Claude 3 Haiku | 12 | unverifiedT1 | 2024-03-04 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
Spread between best and worst is the number to check on this table: a wide gap means the benchmark still separates models, while a narrow gap means close ranks may not mark a real capability difference. A wide score spread indicates that rank differences are trustworthy, as a wide spread supports real separation between models.
3 cited facts
Scores under this harness are considered comparable due to its consistent configuration, but performance may vary with prompt template or geometric tolerance adjustments. The public nature of the test set requires careful evaluation of high scores due to potential contamination risks from the open-source benchmark's community usage. This harness shows sensitivity to prompt template variations and tolerance thresholds, which could affect cross-harness comparisons despite its consistent design.
4 cited facts
CadEval: CadEval measures end-to-end text-to-CAD synthesis: given a textual prompt, a model must produce a CAD program (OpenSCAD-style) whose rendered geometry matches the specification. Generated code is executed in a sandbox and the resulting mesh is compared to a reference via geometric checks.
o3 leads CadEval at 74. The full leaderboard above lists every recorded measurement, not just the headline number.
14 models have recorded scores on CadEval, spanning a score spread of 62.
tensor.news grades CadEval A for integrity (score 98/100), ranking #13 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — CadEval still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on CadEval are reasonably apples-to-apples.
Every CadEval measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites github.com/wgpatrick/cadeval.