M17.6 CONNECT THE MECHANISM
Spend numerical precision where it is needed
Store a gradient of 1 × 10⁻⁹ in fp16 and it becomes exactly zero. See how bits split between range and precision, and how loss scaling rescues small gradients.
LESSON OVERVIEW13 min lesson
Lesson overview
Store a gradient of 1 × 10⁻⁹ in fp16 and it becomes exactly zero. See how bits split between range and precision, and how loss scaling rescues small gradients.
What you’ll explore
- Mixed precision uses different number formats for selected tensors and operations; scaling, accumulation, and stability checks protect training while reducing storage or increasing throughput.
GO TO THE SOURCE
Original explanations, connected to the research.
Mixed Precision Training (Micikevicius et al., 2017)A Study of BFLOAT16 for Deep Learning Training (Kalamkar et al., 2019)FP8 Formats for Deep Learning (Micikevicius et al., 2022)PyTorch automatic mixed precision and GradScaler documentationDeepSeek-V3 Technical ReportSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.