← Concept Index

Evaluation & observability

Human review of outputs

Also called: sampling outputs, rating quality, spot-checking

DEFINITION

Regularly having people read a sample of what a live AI system actually produced, and rate it against real criteria — the ground truth that automated metrics and the model's own confidence can't give you.

WHY IT MATTERS

Automated scores drift from what users actually experience, and a model can't reliably grade itself. Sampling real outputs is how teams catch quality decay, new failure modes, and 'technically passing but unhelpful' answers before users do.

COMMONLY CONFUSED WITH

Waiting for complaints. Complaints are a biased, late signal from the few who bother; deliberate sampling checks the typical output, not just the ones that made someone angry.

SOURCES

First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.