M22.7 CONNECT THE MECHANISM
Why image generators work in latent space
Stable Diffusion doesn't denoise pixels. It denoises a picture 48 times smaller, reads your prompt through attention, and exaggerates what the prompt changes. Follow one request from words to image.
LESSON OVERVIEW15 min lesson
Lesson overview
Stable Diffusion doesn't denoise pixels. It denoises a picture 48 times smaller, reads your prompt through attention, and exaggerates what the prompt changes. Follow one request from words to image.
What you’ll explore
- Trace latent-diffusion training and sampling, and distinguish conditioning from guidance.
GO TO THE SOURCE
Original explanations, connected to the research.
High-Resolution Image Synthesis with Latent Diffusion Models (Rombach et al., 2022)Classifier-Free Diffusion Guidance (Ho & Salimans, 2022)Stable Diffusion v1-4 model card (CompVis)Denoising Diffusion Probabilistic Models (Ho, Jain & Abbeel, 2020)Scalable Diffusion Models with Transformers (Peebles & Xie, 2022)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.