← Concept IndexDEFINITION WHY IT MATTERS COMMONLY CONFUSED WITH
Image understanding (vision input)
Also called: reading an image, vision models, describe this photo
The ability to take an image as input and describe, analyse, or answer questions about it — reading a chart, transcribing a sign, or summarising a screenshot — as distinct from generating images.
It unlocks genuinely useful tasks, from turning a photo of a receipt into text to explaining a diagram. But it also inherits the same fallibility: a vision model can misread numbers or invent details with full confidence.
Image generation. Understanding takes a picture in and gives words out; generation takes words in and makes a picture. They're different capabilities that a 'multimodal' model may combine.