M17.8 CONNECT THE MECHANISM
Make a large training run recoverable and auditable
Meta's Llama 3 run was interrupted about every three hours, mostly by hardware. Learn the arithmetic of failures, how often to checkpoint, and how to spot the failures that don't crash anything.
LESSON OVERVIEW14 min lesson
Lesson overview
Meta's Llama 3 run was interrupted about every three hours, mostly by hardware. Learn the arithmetic of failures, how often to checkpoint, and how to spot the failures that don't crash anything.
What you’ll explore
- Reliable training operations connect health monitoring, checkpoint recovery, reproducible configuration, and resource accounting to the identity and limitations of the resulting model.
GO TO THE SOURCE
Original explanations, connected to the research.
A first order approximation to the optimum checkpoint interval (Young, Communications of the ACM, 1974)A higher order estimate of the optimum checkpoint interval for restart dumps (Daly, Future Generation Computer Systems, 2006)The Llama 3 Herd of Models (Llama Team, Meta, 2024)PaLM: Scaling Language Modeling with Pathways (Chowdhery et al., 2022)Silent Data Corruptions at Scale (Dixit et al., 2021)Llama 2: Open Foundation and Fine-Tuned Chat Models (Touvron et al., 2023), table 2 on training energySuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.