Monitoring and drift
Also called: production monitoring, quality decay, online evaluation
Watching a live system's quality over time, because a model that passed every test at launch can slowly degrade as real inputs, user behavior, the world, or an upgraded model underneath it shift away from what you tested.
Evaluation before launch is a snapshot; production is a moving target. Without ongoing signals — sampled human ratings, complaint rates, refusal and fallback rates, latency — drift is invisible until users have already been getting worse answers for weeks.
One-time evaluation. Passing your eval set proves the system worked on launch day; monitoring is how you learn it still works today.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1U.S. National Institute of Standards and Technology · 2023-01-26
First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.