Technology Innovation Institute·released 2023-09-061 source
Falcon-180B benchmark scores: 8 benchmarks tracked, leading 1. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 8 benchmarks·best result 85.33 on HellaSwag (unverified)·0 of 8 independently reproduced
Head-to-headFalcon-180B vs Llama 2-70BLeads on: LAMBADA
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| HellaSwag | 85.33accuracy (%) | unverified· optimizedT1 | 2023-09-06The Falcon Series of Open Language Models |
| LAMBADA | 79.8accuracy (%) | unverified· optimizedT1 | 2023-09-06The Falcon Series of Open Language Models |
| Winogrande | 74.2accuracy (%) | unverified· optimizedT1 | 2023-09-06The Falcon Series of Open Language Models |
| PIQA | 69.8accuracy (%) | unverified· optimizedT1 | 2023-09-06The Falcon Series of Open Language Models |
| MMLU | 60.8accuracy (%) | unverified· optimizedT1 | 2023-09-06The Falcon Series of Open Language Models |
| ARC AI2 | 57.07accuracy (%) | unverified· optimizedT1 | 2023-09-06The Falcon Series of Open Language Models |
| GSM8K | 54.4accuracy (%) | unverified· optimizedT1 | 2023-09-06Yi: Open Foundation Models by 01.AI |
| OpenBookQA | 52.27accuracy (%) | unverified· optimizedT1 | 2023-09-06The Falcon Series of Open Language Models |
Claims drawn from cited facts, not live model generation.
This model was developed by Technology Innovation Institute and released in 2023.
2 cited facts
An 8-benchmark record is enough to take seriously and too little to be definitive — mid-band coverage, in short. Its 1 top score is a concrete lead on a single test — narrow leadership, but the checkable kind, tied to a specific harness rather than a general claim. The record is unverified, meaning it has neither been confirmed nor disputed by outside evaluation, so treat the scores as claims rather than verified results.
3 cited facts
The model trails the average state-of-the-art score by 18.9 points, indicating a wide gap and placing it well behind the frontier. It holds the top score on only one benchmark, so it does not demonstrate genuine category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Falcon-180B is an AI model developed by Technology Innovation Institute, released 2023-09-06. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Falcon-180B has recorded scores on 8 benchmarks, each shown with its evidence status.
Falcon-180B has recorded scores on 8 benchmarks — PIQA, OpenBookQA, Winogrande, MMLU, ARC AI2, GSM8K, and 2 more. The full table above shows each score with its evidence status.
0 of 8 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Falcon-180B has 8 tracked claims: 8 unverified.