M14.7 CONNECT THE MECHANISM
One vocabulary does not divide every language equally
Say "good morning" in Hindi and a tokenizer may charge you many times more than in English. See where the gap comes from, what it costs, and how to measure fairly.
LESSON OVERVIEW13 min lesson
Lesson overview
Say "good morning" in Hindi and a tokenizer may charge you many times more than in English. See where the gap comes from, what it costs, and how to measure fairly.
What you’ll explore
- Tokenization affects sequence length, rare strings, multilingual coverage, code, and cost; compare the encoded units and normalization before comparing model metrics.
GO TO THE SOURCE
Original explanations, connected to the research.
Language Model Tokenizers Introduce Unfairness Between Languages (Petrov et al., 2023)Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models (Ahia et al., 2023)LLaMA: Open and Efficient Foundation Language Models (Touvron et al., 2023), section 2.1 on the tokenizerUnicode Standard AnnexSpeech and Language Processing, 3rd edition draft (Jurafsky & Martin)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.