← Concept IndexDEFINITION WHY IT MATTERS COMMONLY CONFUSED WITH SOURCES
Evaluating AI output
Also called: testing an AI system, how good is it
Evaluation is checking, in a structured and repeatable way, whether an AI system does what you need — using a fixed set of realistic cases with known-good answers, not a few lucky demos.
Impressions from a handful of prompts are misleading. A real eval set catches regressions when you change the prompt or model, and turns 'it feels good' into evidence you can act on.
A public benchmark score. Benchmarks measure general ability; your evaluation measures fitness for your specific task.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1U.S. National Institute of Standards and Technology · 2023-01-26
First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.