Back to the lesson libraryMECHANISM · 14 MIN
M21.2 CONNECT THE MECHANISM

Teach a useful starting style for problem solving

A model trained purely by reward learned to reason but mixed languages mid-thought. A few thousand good worked examples fixed that. Learn what makes a demonstration worth copying.

LESSON OVERVIEW14 min lesson

Lesson overview

A model trained purely by reward learned to reason but mixed languages mid-thought. A few thousand good worked examples fixed that. Learn what makes a demonstration worth copying.

What you’ll explore

  • Reasoning demonstrations and cold-start tuning supply examples of problem decomposition and presentation; curate valid, diverse traces and distinguish supervised imitation from later reward optimization.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.