Technology Innovation Institute·released 2023-04-24✓ 2 sources
Falcon-7B benchmark scores: 10 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 74.9 on LAMBADA (unverified)·0 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| LAMBADA | 74.9accuracy (%) | unverified· optimizedT1 | 2023-04-24The Falcon Series of Open Language Models |
| HellaSwag | 70.8accuracy (%) | self-reported· optimizedT1 | 2023-04-24Falcon2-11B Technical Report |
| TriviaQA | 64.6accuracy (%) | unverified· optimizedT1 | 2023-04-24Llama 2: Open Foundation and Fine-Tuned Chat Models |
| PIQA | 60.6accuracy (%) | self-reported· optimizedT1 | 2023-04-24Qwen Technical Report |
| OpenBookQA | 35.47accuracy (%) | unverified· optimizedT1 | 2023-04-24Llama 2: Open Foundation and Fine-Tuned Chat Models |
| Winogrande | 34.4accuracy (%) | unverified· optimizedT1 | 2023-04-24Llama 2: Open Foundation and Fine-Tuned Chat Models |
| ARC AI2 | 30.48accuracy (%) | unverified· optimizedT1 | 2023-04-24The Falcon Series of Open Language Models |
| MMLU | 13.33accuracy (%) | self-reported· optimizedT1 | 2023-04-24Falcon2-11B Technical Report |
| GSM8K | 6.8accuracy (%) | unverified· optimizedT1 | 2023-04-24Stanford HELM |
| BBH | 5.03accuracy (%) | unverified· optimizedT1 | 2023-04-24Baichuan 2: Open Large-scale Language Models |
Claims drawn from cited facts, not live model generation.
Developed by Technology Innovation Institute, this model was released in 2023 with 7.22 billion parameters, placing it at the compact-to-mid-size scale. That scale implies a favorable trade-off between efficiency and cost, offering lower inference expense than frontier-scale systems while retaining substantial headroom for many tasks.
4 cited facts
This model is evaluated across 10 tracked benchmarks, but it currently holds no top score on any of them. Because the recorded performance is self-reported and not yet independently confirmed, treat the numbers as vendor claims rather than verified results.
3 cited facts
The model trails the benchmark leader by an average of 46.17 points, a wide gap that places it well behind the front of the pack. It holds the top score on none of the tracked benchmarks, so it shows no current category leadership.
2 cited facts
2 facts cross-checked across data sources: 1 corroborated, 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models
Falcon-7B is an AI model developed by Technology Innovation Institute, released 2023-04-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Falcon-7B has recorded scores on 10 benchmarks, each shown with its evidence status.
Falcon-7B has recorded scores on 10 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, MMLU, ARC AI2, and 4 more. The full table above shows each score with its evidence status.
0 of 10 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Falcon-7B has 10 tracked claims: 3 self-reported, 7 unverified.