Technology Innovation Institute·released 2023-03-151 source
Falcon-40B benchmark scores: 10 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 80.37 on HellaSwag (self-reported)·0 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| HellaSwag | 80.37accuracy (%) | self-reported· optimizedT1 | 2023-03-15Falcon2-11B Technical Report |
| TriviaQA | 79.9accuracy (%) | unverified· optimizedT1 | 2023-03-15Llama 2: Open Foundation and Fine-Tuned Chat Models |
| LAMBADA | 77.3accuracy (%) | unverified· optimizedT1 | 2023-03-15The Falcon Series of Open Language Models |
| PIQA | 66accuracy (%) | unverified· optimizedT1 | 2023-03-15Llama 2: Open Foundation and Fine-Tuned Chat Models |
| Winogrande | 53.8accuracy (%) | unverified· optimizedT1 | 2023-03-15Llama 2: Open Foundation and Fine-Tuned Chat Models |
| ARC AI2 | 49.15accuracy (%) | unverified· optimizedT1 | 2023-03-15The Falcon Series of Open Language Models |
| MMLU | 42.53accuracy (%) | self-reported· optimizedT1 | 2023-03-15Falcon2-11B Technical ReportFalcon2-11B Technical Report |
| OpenBookQA | 42.13accuracy (%) | unverified· optimizedT1 | 2023-03-15Llama 2: Open Foundation and Fine-Tuned Chat Models |
| GSM8K | 33.8accuracy (%) | unverified· optimizedT1 | 2023-03-15Stanford HELM |
| BBH | 16.13accuracy (%) | self-reported· optimizedT1 | 2023-03-15Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Technology Innovation Institute and released in 2023.
2 cited facts
The 10-benchmark spread without a first place gives a dependable mid-field read; coverage that wide rules out the placement being a quirk of test selection. Because the scores are self-reported, they are vendor-claimed and not independently confirmed, so these numbers should be read as claims rather than verified results.
3 cited facts
The model trails the leader by an average of 31.7 points, a wide gap that places it well behind the frontier.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Falcon-40B is an AI model developed by Technology Innovation Institute, released 2023-03-15. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Falcon-40B has recorded scores on 10 benchmarks, each shown with its evidence status.
Falcon-40B has recorded scores on 10 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, MMLU, ARC AI2, and 4 more. The full table above shows each score with its evidence status.
0 of 10 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Falcon-40B has 10 tracked claims: 3 self-reported, 7 unverified.