M28.3 CONNECT THE MECHANISM
Treat evaluators as fallible measurement instruments
Your two best graders agree 92% of the time, and that can still mean almost nothing. Measure agreement properly, then catch a model judge that prefers whichever answer it reads first.
LESSON OVERVIEW17 min lesson
Lesson overview
Your two best graders agree 92% of the time, and that can still mean almost nothing. Measure agreement properly, then catch a model judge that prefers whichever answer it reads first.
What you’ll explore
- Design and audit a human/model judging procedure with explicit criteria, counterbalanced order, and separate correctness and style judgments.
GO TO THE SOURCE
Original explanations, connected to the research.
Zheng et al. — Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaA Coefficient of Agreement for Nominal Scales (Cohen, 1960)The Measurement of Observer Agreement for Categorical Data (Landis & Koch, 1977)Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference (Chiang et al., 2024)LMSYS blog: Chatbot Arena's move from online Elo to the Bradley–Terry model (December 2023)Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators (Dubois et al., 2024)Liang et al. — Holistic Evaluation of Language ModelsSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.