M14.6 CONNECT THE MECHANISM
Learn the units used by a language model
A 1994 file-compression trick now decides how every chatbot reads text. Learn byte-pair encoding by hand, run it on the original paper's example, and see what vocabulary size costs.
LESSON OVERVIEW15 min lesson
Lesson overview
A 1994 file-compression trick now decides how every chatbot reads text. Learn byte-pair encoding by hand, run it on the original paper's example, and see what vocabulary size costs.
What you’ll explore
- Trace one byte-pair merge and distinguish tokenizer training from later text encoding and model training.
GO TO THE SOURCE
Original explanations, connected to the research.
Neural Machine Translation of Rare Words with Subword Units (Sennrich, Haddow & Birch, 2016)Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates (Kudo, 2018)SentencePiece: A simple and language independent subword tokenizer and detokenizer (Kudo & Richardson, 2018)Language Models are Unsupervised Multitask Learners (Radford et al., 2019), the GPT-2 paper, section 2.2Hugging Face LLM course, chapter 6 — byte-pair encoding tokenization (WordPiece and Unigram follow in sections 6 and 7)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.