TOPIC 06 OF 06
Deciding 'good enough'
Matching the bar to the stakesEvaluation isn't chasing 100% — it's deciding what level of quality is acceptable given what's at stake, and being honest when you're below it. A brainstorming aid can ship at 'usually helpful.' A tool that touches money, health, or people's rights needs a far higher bar, human review in the loop, and often a specialist's sign-off. The test set tells you where you are; the stakes tell you where you need to be.
Your eval shows a medical-summary tool is right 92% of the time. For idea generation that'd be great; for clinical use it's nowhere near enough, and the honest conclusion is 'not ready to act on unaided,' not 'ship it, 92% is high.'
A high score is not automatically a passing score — 'good enough' depends entirely on the cost of the errors. Setting that bar before you look at results keeps you from rationalising a number into acceptability after the fact. This is judgment, not a legal or clinical determination; high-stakes uses still need qualified review.
For one AI use, write the bar first: 'this is good enough to act on unaided when accuracy on our test set is at least ___ and the failures are ___.' Then measure against it.