M13.7 CONNECT THE MECHANISM
Generate sound from a representation
To read your note aloud, a model must produce 16,000 numbers for every second of speech. See how WaveNet, vocoders, and codec tokens make that fast enough to talk.
LESSON OVERVIEW14 min lesson
Lesson overview
To read your note aloud, a model must produce 16,000 numbers for every second of speech. See how WaveNet, vocoders, and codec tokens make that fast enough to talk.
What you’ll explore
- Speech and audio generators separate conditioning, intermediate representations, and waveform synthesis; vocoders, codecs, and streaming choices introduce different quality and latency tradeoffs.
GO TO THE SOURCE
Original explanations, connected to the research.
WaveNet: A Generative Model for Raw Audio (van den Oord et al., 2016)Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions (Shen et al., 2018)HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis (Kong et al., 2020)SoundStream: An End-to-End Neural Audio Codec (Zeghidour et al., 2021)High Fidelity Neural Audio Compression (EnCodec; Défossez et al., 2022)Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers (VALL-E; Wang et al., 2023)Speech and Language Processing, 3rd edition draft (Jurafsky & Martin)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.