← Concept Index

Multimodality & generation

Limits of multimodal models

Also called: why it misread the image, OCR errors, spatial reasoning

DEFINITION

The characteristic failure modes when models handle images, audio, or video: mis-reading small text or numbers, weak spatial and counting reasoning, and confidently describing things that aren't there.

WHY IT MATTERS

Multimodal output looks authoritative, so its mistakes are easy to trust. Knowing where these models are weak — exact figures in a chart, how many objects, precise layout — tells you what to double-check.

COMMONLY CONFUSED WITH

Human-level perception. These models don't 'see' as we do; they can miss the obvious and invent the plausible, especially on precise, detailed, or unusual inputs.