AI Fundamentals

Speech-to-Text (STT)

Automatic transcription of audio into text. Modern systems (Whisper, Deepgram, AssemblyAI, ElevenLabs) run near-real-time with speaker diarization and cross-language support. STT is the front of most voice-AI pipelines feeding an LLM downstream.

Related terms

Next steps

Where this fits

Speech-to-Text (STT) is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.