M23.2 CONNECT THE MECHANISM
Bring matching images and descriptions closer in representation space
In 2021, a model trained on 400 million captioned web images learned to recognize things it was never given a label for. Build its loss by hand from a 3 × 3 table.
LESSON OVERVIEW17 min lesson
Lesson overview
In 2021, a model trained on 400 million captioned web images learned to recognize things it was never given a label for. Build its loss by hand from a 3 × 3 table.
What you’ll explore
- CLIP-style contrastive training learns image and text representations by comparing paired and unpaired examples; similarity supports retrieval and prompted classification but is not calibrated truth.
GO TO THE SOURCE
Original explanations, connected to the research.
Learning Transferable Visual Models From Natural Language Supervision (CLIP; Radford et al., 2021)Multimodal Neurons in Artificial Neural Networks (Goh et al., Distill, 2021)Sigmoid Loss for Language Image Pre-Training (SigLIP; Zhai et al., 2023)Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision (ALIGN; Jia et al., 2021)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.