Automatic speech recognition turns speech into text. OpenAI's Whisper (open-weight) set a strong bar: trained on diverse multilingual audio, it handles accents, background noise, and many languages, and can translate. It powers captions, voice interfaces, meeting notes, and the audio side of voice assistants. Run it locally or via API; it's the default building block for "voice in".