Back to the lesson libraryMECHANISM · 17 MIN
M23.2 CONNECT THE MECHANISM

Bring matching images and descriptions closer in representation space

In 2021, a model trained on 400 million captioned web images learned to recognize things it was never given a label for. Build its loss by hand from a 3 × 3 table.

LESSON OVERVIEW17 min lesson

Lesson overview

In 2021, a model trained on 400 million captioned web images learned to recognize things it was never given a label for. Build its loss by hand from a 3 × 3 table.

What you’ll explore

  • CLIP-style contrastive training learns image and text representations by comparing paired and unpaired examples; similarity supports retrieval and prompted classification but is not calibrated truth.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.