Back to the lesson libraryMECHANISM · 12 MIN
M12.5 CONNECT THE MECHANISM

Treat image patches as a sequence of representations

In 2020, a model that chops a photo into 196 squares and reads them like words matched the best CNNs. Count its tokens, and see why halving the patch size costs 16 times the attention.

LESSON OVERVIEW12 min lesson

Lesson overview

In 2020, a model that chops a photo into 196 squares and reads them like words matched the best CNNs. Count its tokens, and see why halving the patch size costs 16 times the attention.

What you’ll explore

  • Vision transformers embed image patches and use attention to mix their information; patch size, positional information, and training choices determine cost and spatial detail.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.