← Concept Index

Multimodality & generation

Multimodal models

Also called: vision-language, image understanding

DEFINITION

A multimodal model can take in and/or produce more than one kind of data — for example reading an image and answering questions about it, or turning text into speech.

WHY IT MATTERS

It's why you can hand a model a screenshot or a chart, but also why the same trust rules apply across modalities: a confident wrong description of an image is still a hallucination.

COMMONLY CONFUSED WITH

Separate specialised tools bolted together. A true multimodal model handles the modalities within one system, though behaviour still varies by modality.