M22.2 CONNECT THE MECHANISM
Generate different kinds of data one unit at a time
A language model writes one word at a time. The same trick can paint a picture one pixel at a time, or one image token at a time. Here's the rule that makes it exact.
LESSON OVERVIEW13 min lesson
Lesson overview
A language model writes one word at a time. The same trick can paint a picture one pixel at a time, or one image token at a time. Here's the rule that makes it exact.
What you’ll explore
- Autoregressive generation factors a joint distribution into conditional predictions; representation and ordering determine whether the generated units are text, pixels, audio codes, or structured fields.
GO TO THE SOURCE
Original explanations, connected to the research.
Pixel Recurrent Neural Networks (van den Oord, Kalchbrenner & Kavukcuoglu, 2016)Conditional Image Generation with PixelCNN Decoders (van den Oord et al., 2016)WaveNet: A Generative Model for Raw Audio (van den Oord et al., 2016)Neural Discrete Representation Learning (VQ-VAE; van den Oord, Vinyals & Kavukcuoglu, 2017)Zero-Shot Text-to-Image Generation (DALL·E; Ramesh et al., 2021)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.