M21.5 CONNECT THE MECHANISM
Choose whether to score the result, the steps, or both
Cross out the 6s in 16/64 and you get 1/4, the right answer by a wrong method. Should the model be rewarded? That question splits outcome rewards from process rewards.
LESSON OVERVIEW13 min lesson
Lesson overview
Cross out the 6s in 16/64 and you get 1/4, the right answer by a wrong method. Should the model be rewarded? That question splits outcome rewards from process rewards.
What you’ll explore
- Outcome rewards score the final answer and process rewards score each step; both can be gamed, so check correctness, shortcuts, and whether the reward matches what you want.
GO TO THE SOURCE
Original explanations, connected to the research.
Let's Verify Step by Step (Lightman et al., 2023)Solving math word problems with process- and outcome-based feedback (Uesato et al., 2022)Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations (Wang et al., 2023)Faulty Reward Functions in the Wild (OpenAI, 2016)The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models (Pan et al., 2022)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.