M12.5 CONNECT THE MECHANISM
Treat image patches as a sequence of representations
In 2020, a model that chops a photo into 196 squares and reads them like words matched the best CNNs. Count its tokens, and see why halving the patch size costs 16 times the attention.
LESSON OVERVIEW12 min lesson
Lesson overview
In 2020, a model that chops a photo into 196 squares and reads them like words matched the best CNNs. Count its tokens, and see why halving the patch size costs 16 times the attention.
What you’ll explore
- Vision transformers embed image patches and use attention to mix their information; patch size, positional information, and training choices determine cost and spatial detail.
GO TO THE SOURCE
Original explanations, connected to the research.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (Dosovitskiy et al., 2020)Training data-efficient image transformers & distillation through attention (Touvron et al., 2020)Dive into Deep Learning — transformers for visionSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.