← Concept IndexDEFINITION WHY IT MATTERS COMMONLY CONFUSED WITH SOURCES
Benchmarks and their limits
Also called: leaderboards, eval contamination
Benchmarks are standard test sets used to compare models. They're useful for rough capability, but can be gamed, can leak into training data, and rarely match your real workload.
A top leaderboard model can still be wrong for your task. Treat benchmarks as a starting filter, then judge on your own cases — the ones you actually care about getting right.
Proof of fitness. A high score is evidence of general ability, not that the model handles your data, tone, and edge cases.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1U.S. National Institute of Standards and Technology · 2023-01-26
First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.