LANDMARK · 8 MIN · USE
Images and multimodality
After this landmark, you can work across images, audio, and video with the same skepticism you’d apply to text, and treat provenance as part of the output.
Multimodal models can generate and interpret pictures, sound, and video, but they’re doing the same thing they do with text: producing a plausible output, not a verified one. A generated image can invent details that were never specified; a model describing your photo can confidently misread it. Different formats feel more like evidence — a picture looks like proof in a way a paragraph doesn’t — but the reliability hasn’t changed. What has changed is that synthetic media is now cheap, so where something came from matters as much as what it shows.
This is the canonical concept. It stays the same across learner lenses so personalization never changes the underlying facts.
◌ EVERYDAY LENS
What this looks like for you
Treat a model’s reading of a photo or a chart the way you’d treat its reading of a document — check the detail that matters. And remember a convincing image or voice can be generated, so “I saw it” isn’t proof anymore.
MAKE A DECISION
You ask an AI to generate a chart of your quarterly numbers from a table you paste in. It returns a polished, professional-looking bar chart. What do you check first?
CARRY THISNext time AI reads or generates an image for you, name one specific detail — a number, a label, a face — and verify it against the real thing.