LANDMARK · 7 MIN · UNDERSTAND
Data provenance
After this landmark, you can ask precise questions about where a model’s data came from, and read provenance signals on content without over-trusting them.
↗
This concept is shared, but the Everyday lens is less central here.
You can still explore it. We’re showing the shared explanation and a related practical view without hiding the knowledge.
Provenance is the history of a piece of data: where it came from, under what terms, and what has happened to it since. It matters at both ends of a model. On the input side, large training corpora are assembled largely by scraping the public web, which means they contain copyrighted work, personal data, and material whose authors never agreed to this use — the substance of ongoing litigation and of transparency rules requiring providers to publish a summary of training content. On the output side, the question reverses: given a piece of text, an image, or a recording, can anyone tell whether it was generated? Content credentials and watermarking exist to answer that, by attaching signed metadata or embedding a detectable signal. Both are worth supporting and neither is reliable as a test, because metadata is stripped by ordinary processing and detectors have real error rates in both directions.
This is the canonical concept. It stays the same across learner lenses so personalization never changes the underlying facts.
◌ EVERYDAY LENS
What this looks like for you
The absence of an “AI-generated” label means nothing — most synthetic content carries no marker at all, and screenshotting removes what markers exist. Judge by source and corroboration rather than by looking for a badge, and don’t trust a detector tool’s verdict either way.
MAKE A DECISION
A student’s essay is flagged by an AI-detection tool as “98% likely AI-generated”. What is the responsible response?
CARRY THISTake one AI-assisted document you produced recently and write down where its content came from: your words, retrieved sources, and model-generated material.