TRAIL

Evaluate an AI product

Look beyond demos to evidence, access, limits, and consequences.

Every AI product demos well. This trail is the checklist for looking past the demo — at what the system actually is, what evidence backs its claims, what it touches, and what happens when it's wrong.

By the end you'll be able to interrogate a tool the way a skeptical buyer should: separating model from marketing, tracing claims to sources, probing for bias and security holes, and weighing the licensing and consequences before you commit.

12 landmarks · 79 min · 0/12 explored

Begin trail
  1. 1OrientModel, product, system

    First question about any product: what are you actually buying? The same underlying model can be wrapped into wildly different systems with different guarantees.

  2. 2OrientFluent is not true

    A slick demo proves the output is fluent, not that it's correct. Keep the two apart while you evaluate.

    Practice · Reality Lab: Repair a confident draft Practice · Consequence Room: The number nobody checked

  3. 3OrientTraining and use are different

    What a model learned in training and what it's given at use time are different things — and it decides how much a vendor's "it knows your domain" claim is worth.

  4. 4JudgeMatch checking to risk

    Decide how much evidence this purchase deserves based on what it will be trusted to do. A toy and a decision system are not evaluated the same way.

    Practice · Reality Lab: Ground an answer in its sources Practice · Reality Lab: Match the check to the risk Practice · Reality Lab: Verify a message that might be fake Practice · Consequence Room: The number nobody checked Practice · Consequence Room: The snippet that shipped

  5. 5JudgeTrace a claim

    Take one concrete claim the product makes and follow it to a source you can verify. Claims that can't be traced are marketing.

    Practice · Reality Lab: Repair a confident draft Practice · Reality Lab: Break a big ask into steps Practice · Reality Lab: Ground an answer in its sources Practice · Reality Lab: Make the draft your own Practice · Reality Lab: Match the check to the risk Practice · Consequence Room: The number nobody checked

  6. 6JudgeBias and representation

    Probe who the system works less well for. A product that performs unevenly across people is a liability, not a feature.

    Practice · Consequence Room: The shortlist machine

  7. 7JudgePrompt injection

    If the tool reads outside content or calls tools, it can be hijacked by what it reads. Ask how the vendor separates instructions from untrusted input.

    Practice · Consequence Room: The assistant that reads the web

  8. 8JudgeHuman oversight

    Check whether the product actually leaves room for a human to understand and override it — or whether oversight is theater.

    Practice · Reality Lab: Turn a process into an agent workflow Practice · Consequence Room: The assistant that reads the web Practice · Consequence Room: The snippet that shipped Practice · Consequence Room: The shortlist machine

  9. 9JudgeVendor evidence

    Now ask the vendor for the thing that would actually change your mind: results on your own examples, and an honest account of what the system gets wrong.

  10. 10JudgeConflicting sources

    Before you trust any summary the tool produces, check whether it flattened a real disagreement into one smooth answer. The disagreement was the information.

  11. 11JudgeReality Lab: evidence

    A hands-on check: work through the evidence for a claim yourself, so evaluating stops being abstract.

  12. 12JudgeCopyright and permission

    The last dimension buyers forget: what the tool is trained on and what you're licensed to do with its output. Rights and permission close the evaluation.

    Practice · Reality Lab: Make the draft your own Practice · Consequence Room: The image you didn’t own