GLOSSARY
Speech Recognition (ASR)
Converting spoken audio into text — the foundation under meeting transcribers, voice assistants and subtitling tools, now near-human in accuracy for clear audio.
Modern automatic speech recognition is a deep-learning pipeline: an audio encoder (often transformer-based, in the Whisper lineage) maps sound to representations that a decoder turns into words, with punctuation, capitalization and even speaker labels. Noise, accents, overlapping speakers and domain jargon remain the accuracy killers.
In this directory, ASR is the invisible engine of a whole product family: meeting notetakers (Otter, Fireflies), podcast editors (Descript), subtitle generators (VEED, Captions) and voice interfaces. Choosing between them is mostly about the workflow around the transcription — summaries, action items, editing — because raw accuracy differences have narrowed dramatically.