M23.5 CONNECT THE MECHANISM
Train alignment, generation, and instruction behavior across modalities
LLaVA's first stage trains just 0.3% of the model's weights, and it works. Walk through the stages that turn two separate models into one assistant.
LESSON OVERVIEW14 min lesson
Lesson overview
LLaVA's first stage trains just 0.3% of the model's weights, and it works. Walk through the stages that turn two separate models into one assistant.
What you’ll explore
- Multimodal training combines data pairing, modality objectives, interface adaptation, and instruction or preference stages; record frozen components, loss masks, and cross-modal evaluation.
GO TO THE SOURCE
Original explanations, connected to the research.
Visual Instruction Tuning (LLaVA; Liu, Li, Wu & Lee, 2023)Improved Baselines with Visual Instruction Tuning (LLaVA-1.5; Liu, Li, Li & Lee, 2023)Aligning Large Multimodal Models with Factually Augmented RLHF (LLaVA-RLHF; Sun et al., 2023)Flamingo: a Visual Language Model for Few-Shot Learning (Alayrac et al., 2022)The Llama 3 Herd of ModelsSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.