B17 A SMALL STEP SIDEWAYS
Measure surprise and compare distributions
Why can a language model's loss fall to 1.75 bits but never below? Measure surprise in bits, then see how cross-entropy, perplexity, and KL grade a model's guesses.
LESSON OVERVIEW12 min lesson
Lesson overview
Why can a language model's loss fall to 1.75 bits but never below? Measure surprise in bits, then see how cross-entropy, perplexity, and KL grade a model's guesses.
What you’ll explore
- Compute surprise, entropy, cross-entropy, and KL divergence for a small distribution, turn a loss into perplexity, and explain why KL divergence depends on its direction.
GO TO THE SOURCE
Original explanations, connected to the research.
Deep Learning — probability and information theorySuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.