LANDMARK · 7 MIN · UNDERSTAND
Uncertainty and confidence
After this landmark, you can explain why stated confidence is unreliable, and name techniques that produce a signal you can actually act on.
Ask a model how sure it is and you get a number that sounds like a probability and behaves like prose. It is generated the same way as everything else — as a plausible continuation — so “I’m 95% confident” is a statement in the style of confident statements, not a measurement. The model’s internal token probabilities are a genuine measurement, but of something narrower than you want: how likely that token was given the training distribution, which is a claim about fluency, not truth. A fabricated citation can be composed of high-probability tokens throughout, because it looks exactly like the citations the model has seen. The practical position is that uncertainty must be obtained from outside the model’s self-report: from agreement between independent attempts, from whether a claim traces to a source, or from a check performed by something that isn’t the model. Hedging language is a style setting; treating it as a risk signal is the mistake this landmark exists to prevent.
This is the canonical concept. It stays the same across learner lenses so personalization never changes the underlying facts.
◌ EVERYDAY LENS
What this looks like for you
Confidence tells you nothing. The useful question is not “are you sure?” — it will simply say yes — but “what would I have to check to know this is right?”. If you can’t check it and it matters, don’t rely on it.
MAKE A DECISION
You need a way to flag which AI-generated case summaries need human review. Which signal is most useful?
CARRY THISAsk an assistant the same factual question three times in separate conversations. Where the answers differ is where its confidence was worth the least.