M14.4 CONNECT THE MECHANISM
How text becomes tokens
A language model never sees your letters. See how text is chopped into pieces and swapped for numbers, and why "東京" can cost three times as many bytes as it has characters.
LESSON OVERVIEW13 min lesson
Lesson overview
A language model never sees your letters. See how text is chopped into pieces and swapped for numbers, and why "東京" can cost three times as many bytes as it has characters.
What you’ll explore
- Distinguish text units from token IDs and diagnose the effects of a tokenizer mismatch.
GO TO THE SOURCE
Original explanations, connected to the research.
Neural Machine Translation of Rare Words with Subword Units (Sennrich, Haddow & Birch, 2016)BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (Devlin et al., 2018)Language Models are Unsupervised Multitask Learners (Radford et al., 2019), the GPT-2 paper, section 2.2 on byte-level BPEPython documentation: Unicode HOWTOSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.