LANDMARK · 8 MIN · JUDGE
Vendor evidence
After this landmark, you can ask an AI vendor for the specific evidence that would change your decision, and recognise the answers that sound like evidence but aren’t.
↗
This concept is shared, but the Everyday lens is less central here.
You can still explore it. We’re showing the shared explanation and a related practical view without hiding the knowledge.
AI procurement goes wrong in a predictable way: the demo is excellent, the pilot is promising, and the deployment disappoints. That isn’t usually dishonesty — it’s that a demo is a curated best case and your work is not. The fix is to ask for evidence with four properties. It should be measured on tasks like yours, not on a public benchmark. It should report failures, not just successes — what does the system get wrong, and how would you notice? It should be reproducible by you, on your own examples, before you commit. And it should describe the conditions: which model version, what data, who checked. Anything that fails all four is marketing. The most useful question you can ask a vendor is not “how accurate is it?” but “show me the cases where it was wrong, and tell me how a customer found out.”
This is the canonical concept. It stays the same across learner lenses so personalization never changes the underlying facts.
◌ EVERYDAY LENS
What this looks like for you
Treat app-store claims and launch-post numbers as advertising. Before you rely on a tool for something that matters, try it on a handful of cases where you already know the right answer — that private test tells you more than any published score.
MAKE A DECISION
A vendor pitching a document-summarisation tool cites a top-three position on a public summarisation benchmark. What should you ask for next?
CARRY THISAssemble five real examples from your own work with known good answers. That set is your evaluation kit — use it on the next AI tool anyone proposes.