← Concept Index

Evaluation & observability

Evaluating AI output

Also called: testing an AI system, how good is it

DEFINITION

Evaluation is checking, in a structured and repeatable way, whether an AI system does what you need — using a fixed set of realistic cases with known-good answers, not a few lucky demos.

WHY IT MATTERS

Impressions from a handful of prompts are misleading. A real eval set catches regressions when you change the prompt or model, and turns 'it feels good' into evidence you can act on.

COMMONLY CONFUSED WITH

A public benchmark score. Benchmarks measure general ability; your evaluation measures fitness for your specific task.

SOURCES

First-pass citations, limited to primary sources; a reviewer will broaden and verify these before this entry leaves draft.