← Concept Index

Evaluation & observability

Benchmarks and their limits

Also called: leaderboards, eval contamination

DEFINITION

Benchmarks are standard test sets used to compare models. They're useful for rough capability, but can be gamed, can leak into training data, and rarely match your real workload.

WHY IT MATTERS

A top leaderboard model can still be wrong for your task. Treat benchmarks as a starting filter, then judge on your own cases — the ones you actually care about getting right.

COMMONLY CONFUSED WITH

Proof of fitness. A high score is evidence of general ability, not that the model handles your data, tone, and edge cases.

SOURCES

First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.