M16.8 CONNECT THE MECHANISM
Continue pretraining without losing sight of the starting model
Teach Kestrel medicine in a tenth of the original compute, then stretch its context from 2,048 to 8,192 tokens, without letting it forget how to write an ordinary email.
LESSON OVERVIEW12 min lesson
Lesson overview
Teach Kestrel medicine in a tenth of the original compute, then stretch its context from 2,048 to 8,192 tokens, without letting it forget how to write an ordinary email.
What you’ll explore
- Continued pretraining adapts an existing checkpoint to new distributions or contexts; preserve the artifact lineage and evaluate gains alongside forgetting and deployment changes.
GO TO THE SOURCE
Original explanations, connected to the research.
Gururangan et al. — Don't Stop Pretraining: Adapt Language Models to Domains and Tasks (2020)Ibrahim et al. — Simple and Scalable Strategies to Continually Pre-train Large Language Models (2024)Chen et al. — Extending Context Window of Large Language Models via Positional Interpolation (2023)Rozière et al. — Code Llama: Open Foundation Models for Code (2023)The Llama 3 Herd of Models (2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.