AI Fundamentals
Speech-to-Text (STT)
Automatic transcription of audio into text. Modern systems (Whisper, Deepgram, AssemblyAI, ElevenLabs) run near-real-time with speaker diarization and cross-language support. STT is the front of most voice-AI pipelines feeding an LLM downstream.
Related terms
Related on this site
Where this fits
Speech-to-Text (STT) is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.