M16.2 CONNECT THE MECHANISM
Build a traceable corpus before spending training compute
A crawled web page is mostly menus, cookie banners, and copies of other pages. Follow a snapshot of the web through the filters that turn it into training text.
LESSON OVERVIEW19 min lesson
Lesson overview
A crawled web page is mostly menus, cookie banners, and copies of other pages. Follow a snapshot of the web through the filters that turn it into training text.
What you’ll explore
- Trace source selection, filtering, deduplication, tokenization, and versioned packing while identifying where evaluation leakage can enter.
GO TO THE SOURCE
Original explanations, connected to the research.
Raffel et al. — Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5 and C4, 2019)Penedo et al. — The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale (2024)Lee et al. — Deduplicating Training Data Makes Language Models Better (2021)Dodge et al. — Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (2021)Brown et al. — Language Models are Few-Shot Learners (GPT-3, 2020)Gao et al. — The Pile: An 800GB Dataset of Diverse Text for Language Modeling (2020)Soldaini et al. — Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research (2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.