← Concept Index

Multimodality & generation

Speech-to-text (transcription)

Also called: transcription, ASR, voice to text

DEFINITION

Turning spoken audio into written text. Modern systems are fast and accurate on clear speech, and are what power meeting notes, captions, and voice assistants.

WHY IT MATTERS

It's one of the most reliable and widely useful AI capabilities — but accuracy drops with accents, jargon, crosstalk, and noise, and a wrong transcript quietly propagates into every summary built on it.

COMMONLY CONFUSED WITH

Understanding what was said. Transcription produces words; it doesn't judge importance or meaning, and it can mis-hear a critical name or number without any signal that it did.