B17 A SMALL STEP SIDEWAYS

Measure surprise and compare distributions

Why can a language model's loss fall to 1.75 bits but never below? Measure surprise in bits, then see how cross-entropy, perplexity, and KL grade a model's guesses.

LESSON OVERVIEW12 min lesson

Lesson overview

Why can a language model's loss fall to 1.75 bits but never below? Measure surprise in bits, then see how cross-entropy, perplexity, and KL grade a model's guesses.

What you’ll explore

  • Compute surprise, entropy, cross-entropy, and KL divergence for a small distribution, turn a loss into perplexity, and explain why KL divergence depends on its direction.
GO TO THE SOURCE

Original explanations, connected to the research.

Deep Learning — probability and information theory
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.