LANDMARK · 8 MIN · UNDERSTAND

How evaluation works

After this landmark, you can build a small evaluation set for a real task and explain what it can and cannot tell you.

EU context
This concept is shared, but the Everyday lens is less central here.

You can still explore it. We’re showing the shared explanation and a related practical view without hiding the knowledge.

SHARED FOUNDATION

Because model output varies between runs and degrades in ways that aren’t obvious, judging quality by trying a few prompts is unreliable. Evaluation replaces that with something repeatable: a fixed set of representative cases, a definition of what a good response looks like for each, and a way of scoring. Then a change — a new prompt, a new model version, a new retrieval setting — can be measured rather than felt. The set doesn’t need to be large to be transformative; twenty to fifty real cases, including the awkward ones, will catch most regressions that vibes miss. The essential discipline is that cases come from reality rather than imagination, that the awkward and rare cases are deliberately over-represented relative to how often they occur, and that you write down what a good answer is *before* you look at what the system produced.

This is the canonical concept. It stays the same across learner lenses so personalization never changes the underlying facts.

EVERYDAY LENS

What this looks like for you

The personal version: keep a few questions you know the answer to and try them on any new assistant. It’s a private benchmark nobody can optimise for, and it will tell you more than a launch announcement.

MAKE A DECISION

A prompt change makes a summarisation feature noticeably better on the three examples the team tried. What should happen before it ships?

CARRY THISWrite down ten real cases from your work with the answer you would accept for each. That file is your evaluation set — run it against your current tool today.
CONCEPTS IN THE INDEX