← Concept IndexDEFINITION WHY IT MATTERS COMMONLY CONFUSED WITH
Multimodal models
Also called: vision-language, image understanding
A multimodal model can take in and/or produce more than one kind of data — for example reading an image and answering questions about it, or turning text into speech.
It's why you can hand a model a screenshot or a chart, but also why the same trust rules apply across modalities: a confident wrong description of an image is still a hallucination.
Separate specialised tools bolted together. A true multimodal model handles the modalities within one system, though behaviour still varies by modality.