Back to the lesson libraryMECHANISM · 13 MIN
M20.6 CONNECT THE MECHANISM

Train on preference pairs without a separate online reward loop

RLHF juggles four models and a sampling loop. DPO gets the same kind of result from one line of loss and two models, because the assistant turns out to be its own reward model.

LESSON OVERVIEW13 min lesson

Lesson overview

RLHF juggles four models and a sampling loop. DPO gets the same kind of result from one line of loss and two models, because the assistant turns out to be its own reward model.

What you’ll explore

  • DPO adjusts chosen-versus-rejected response likelihoods relative to a reference model using a preference loss; its assumptions and data coverage still constrain resulting behavior.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.