M13.5 CONNECT THE MECHANISM
From a waveform to time-frequency features
Say "buy oat milk" into your phone and it records 16,000 numbers a second. Follow them into the picture of sound that speech models actually read.
LESSON OVERVIEW14 min lesson
Lesson overview
Say "buy oat milk" into your phone and it records 16,000 numbers a second. Follow them into the picture of sound that speech models actually read.
What you’ll explore
- Waveforms sample amplitude over time; frequency transforms and spectrograms reveal complementary structure, with resolution and sampling assumptions that affect audio learning.
GO TO THE SOURCE
Original explanations, connected to the research.
Speech and Language Processing, 3rd edition draft — chapter on automatic speech recognition (Jurafsky & Martin)Robust Speech Recognition via Large-Scale Weak Supervision (Radford et al., 2022)NumPy documentation — numpy.fft.rfftSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.