← Concept Index

Multimodality & generation

Speech & audio generation

Also called: voice cloning, text-to-speech, TTS

DEFINITION

Models can generate natural-sounding speech from text and can imitate a specific voice from a short sample. The same technology transcribes and translates audio.

WHY IT MATTERS

Convincing voice cloning from seconds of audio changes the threat model for phone scams and 'I heard them say it' evidence. The everyday defence is a shared code word and calling back on a known number.

COMMONLY CONFUSED WITH

A recording. Synthetic speech can be indistinguishable by ear, so hearing a familiar voice is no longer proof of who is speaking.