Cohere·released 2024-08-301 source
Command R+ benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 59.2 on MMLU (unverified)·0 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 59.2accuracy (%) | unverifiedT1 | 2024-08-30Stanford CRFM Leaderboard |
| DTBench | 24.88accuracy (%) | unverifiedT1 | 2024-08-30https://conceptualreasoning.ai/dtbench |
| LMCA | 5.89accuracy (%) | unverifiedT1 | 2024-08-30https://conceptualreasoning.ai/lmca |
| SimpleBench | 0.88accuracy (%) | unverifiedT1 | 2024-08-30SimpleBench Leaderboard |
Claims drawn from cited facts, not live model generation.
This model was developed by Cohere and released in 2024. That pairing places it in the current generation of models from an established developer, so its vintage reflects recent design and training practice rather than legacy architecture.
4 cited facts
This model is scored on 4 tracked benchmarks. It holds no current top score on any of them. Its record is unverified: neither confirmed nor disputed by outside evaluation, so the scores should be read as neither independently backed nor contradicted.
3 cited facts
On average this model trails the state-of-the-art leader by 60.84 points across its benchmarks, a wide gap that places it well behind the front of the pack rather than at the frontier or merely mid-pack. It holds the top score on none of its benchmarks (sota leader count of 0), so there is no category in which it can claim genuine leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Command R+ is an AI model developed by Cohere, released 2024-08-30. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Command R+ has recorded scores on 4 benchmarks, each shown with its evidence status.
Command R+ has recorded scores on 4 benchmarks — DTBench, MMLU, LMCA, SimpleBench. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Command R+ has 4 tracked claims: 4 unverified.