← Concept Index

Multimodality & generation

Image understanding (vision input)

Also called: reading an image, vision models, describe this photo

DEFINITION

The ability to take an image as input and describe, analyse, or answer questions about it — reading a chart, transcribing a sign, or summarising a screenshot — as distinct from generating images.

WHY IT MATTERS

It unlocks genuinely useful tasks, from turning a photo of a receipt into text to explaining a diagram. But it also inherits the same fallibility: a vision model can misread numbers or invent details with full confidence.

COMMONLY CONFUSED WITH

Image generation. Understanding takes a picture in and gives words out; generation takes words in and makes a picture. They're different capabilities that a 'multimodal' model may combine.