← Concept IndexDEFINITION WHY IT MATTERS COMMONLY CONFUSED WITH
Limits of multimodal models
Also called: why it misread the image, OCR errors, spatial reasoning
The characteristic failure modes when models handle images, audio, or video: mis-reading small text or numbers, weak spatial and counting reasoning, and confidently describing things that aren't there.
Multimodal output looks authoritative, so its mistakes are easy to trust. Knowing where these models are weak — exact figures in a chart, how many objects, precise layout — tells you what to double-check.
Human-level perception. These models don't 'see' as we do; they can miss the obvious and invent the plausible, especially on precise, detailed, or unusual inputs.