Back to the lesson libraryMECHANISM · 14 MIN
M23.5 CONNECT THE MECHANISM

Train alignment, generation, and instruction behavior across modalities

LLaVA's first stage trains just 0.3% of the model's weights, and it works. Walk through the stages that turn two separate models into one assistant.

LESSON OVERVIEW14 min lesson

Lesson overview

LLaVA's first stage trains just 0.3% of the model's weights, and it works. Walk through the stages that turn two separate models into one assistant.

What you’ll explore

  • Multimodal training combines data pairing, modality objectives, interface adaptation, and instruction or preference stages; record frozen components, loss masks, and cross-modal evaluation.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.