Human review of outputs
Also called: sampling outputs, rating quality, spot-checking
Regularly having people read a sample of what a live AI system actually produced, and rate it against real criteria — the ground truth that automated metrics and the model's own confidence can't give you.
Automated scores drift from what users actually experience, and a model can't reliably grade itself. Sampling real outputs is how teams catch quality decay, new failure modes, and 'technically passing but unhelpful' answers before users do.
Waiting for complaints. Complaints are a biased, late signal from the few who bother; deliberate sampling checks the typical output, not just the ones that made someone angry.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1U.S. National Institute of Standards and Technology · 2023-01-26
First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.