M20.8 CONNECT THE MECHANISM
Measure adaptation gains without overlooking regressions
Staff prefer the new assistant 60% of the time. It also gets more prices wrong, refuses harmless questions, and caves when customers push back. Learn to catch all four before launch.
LESSON OVERVIEW14 min lesson
Lesson overview
Staff prefer the new assistant 60% of the time. It also gets more prices wrong, refuses harmless questions, and caves when customers push back. Learn to catch all four before launch.
What you’ll explore
- Before-and-after evaluation should isolate checkpoint changes, test intended and retained capabilities, and report behavior under matched prompts, decoding, and scoring conditions.
GO TO THE SOURCE
Original explanations, connected to the research.
Training language models to follow instructions with human feedback (InstructGPT; Ouyang et al., 2022)Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al., 2023)XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models (Röttger et al., 2023)Towards Understanding Sycophancy in Language Models (Sharma et al., 2023)Sycophancy in GPT-4o: what happened and what we're doing about it (OpenAI, 2025)Expanding on what we missed with sycophancy (OpenAI, 2025)NIST AI Risk Management FrameworkSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.