Skip to index

GLOSSARY

Speech Synthesis (TTS)

Turning text into natural-sounding speech — the output half of voice AI, now indistinguishable from human narration for many languages.

Modern TTS models generate speech directly from text with realistic prosody, emotion and even laughter — a jump from the robotic voices of a decade ago. The catalog's voice tools split into generic narrators (Murf, PlayHT, WellSaid for e-learning and video), expressive performers (ElevenLabs for audiobooks and games), and specialty tools (voice conversion, dubbing, real-time changelogs).

Selection criteria that actually matter: voice library breadth and licensing, latency for interactive use, fine-grained control (pauses, emphasis, pronunciation), and commercial-use terms. The uncanny valley has essentially closed for narration — which is why disclosure and voice-owner consent are now the industry's live debates.

Related terms

Tools that use this

Related categories