M35.7 CONNECT THE MECHANISM
Compare language, image, audio, and state-space artifacts
A museum guide app that sees, listens, and speaks needs four very different models. Put CLIP, LLaVA, Whisper, and Stable Diffusion on one comparison sheet and see what each one keeps in memory.
LESSON OVERVIEW13 min lesson
Lesson overview
A museum guide app that sees, listens, and speaks needs four very different models. Put CLIP, LLaVA, Whisper, and Stable Diffusion on one comparison sheet and see what each one keeps in memory.
What you’ll explore
- Compare model families by representation, objective, output mechanism, and inference state rather than assuming every generative system is a chat transformer.
GO TO THE SOURCE
Original explanations, connected to the research.
Learning Transferable Visual Models From Natural Language SupervisionVisual Instruction TuningImproved Baselines with Visual Instruction Tuning (LLaVA-1.5)High-Resolution Image Synthesis with Latent Diffusion ModelsScalable Diffusion Models with TransformersWhisperEnCodecMamba: Linear-Time Sequence Modeling with Selective State SpacesSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.