M16.3 CONNECT THE MECHANISM
Choose how much the model sees of each source
Kestrel's encyclopedia is 1% of its data on disk but gets 5% of its training. Work out how many times each source is read, and when repetition stops helping.
LESSON OVERVIEW13 min lesson
Lesson overview
Kestrel's encyclopedia is 1% of its data on disk but gets 5% of its training. Work out how many times each source is read, and when repetition stops helping.
What you’ll explore
- Calculate expected token exposure from mixture weights and distinguish source size, sampling share, oversampling, and validation-based selection.
GO TO THE SOURCE
Original explanations, connected to the research.
Brown et al. — Language Models are Few-Shot Learners (GPT-3, 2020), Table 2.2Muennighoff et al. — Scaling Data-Constrained Language Models (2023)Xie et al. — DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining (2023)The Llama 3 Herd of Models (2024)Groeneveld et al. — OLMo, an open language-model frameworkSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.