Three-step process turns sound into text using probabilities, not understanding.
AI transcription systems use neural networks trained on billions of examples of speech paired with text. The process has three steps. First, encoding sound by breaking audio into very short chunks and turning them into numerical features. Second, predicting text where a language model predicts the most likely word that fits both the sound and the context. Third, stringing it together, just like ChatGPT predicts the next word in a sentence. It’s all probabilities. The model doesn’t understand the words, it just plays the odds.
— AI Fail — Trump Says ‘Merz’ —YouTube Transcribes ‘Merkel’ · Forbes