M23.3 CONNECT THE MECHANISM
Connect a vision encoder to a language model
A photo of a pill bottle becomes 576 "words" that no dictionary contains. Follow them into a language model, and find out where the dose on the label got lost.
LESSON OVERVIEW14 min lesson
Lesson overview
A photo of a pill bottle becomes 576 "words" that no dictionary contains. Follow them into a language model, and find out where the dose on the label got lost.
What you’ll explore
- A vision-language model converts visual inputs into representations consumed by a language model through projections, prefixes, or cross-attention; interfaces and training determine what visual evidence survives.
GO TO THE SOURCE
Original explanations, connected to the research.
Visual Instruction Tuning (LLaVA; Liu, Li, Wu & Lee, 2023)Improved Baselines with Visual Instruction Tuning (LLaVA-1.5; Liu, Li, Li & Lee, 2023)Flamingo: a Visual Language Model for Few-Shot Learning (Alayrac et al., 2022)BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (Li et al., 2023)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.