M13.6 CONNECT THE MECHANISM
Map sound into words with an alignment model
A voice note has 120 frames and the transcript has 12 characters, but nobody said which frame goes with which. See how CTC and Whisper solve that, and score the result with word error rate.
LESSON OVERVIEW14 min lesson
Lesson overview
A voice note has 120 frames and the transcript has 12 characters, but nobody said which frame goes with which. See how CTC and Whisper solve that, and score the result with word error rate.
What you’ll explore
- Speech recognition combines acoustic evidence with sequence modeling and decoding; alignment assumptions, text normalization, and error metrics affect what performance means.
GO TO THE SOURCE
Original explanations, connected to the research.
Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks (Graves et al., 2006)Robust Speech Recognition via Large-Scale Weak Supervision (Radford et al., 2022)Sequence Transduction with Recurrent Neural Networks (Graves, 2012)Listen, Attend and Spell (Chan et al., 2015)wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations (Baevski et al., 2020)Speech and Language Processing, 3rd edition draft (Jurafsky & Martin)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.