M28.1 CONNECT THE MECHANISM
Measure the behavior that the task actually needs
The new support bot beats the old one 78% to 75% on 200 questions. Is it better? Learn to score free-text answers and put honest error bars on the result.
LESSON OVERVIEW14 min lesson
Lesson overview
The new support bot beats the old one 78% to 75% on 200 questions. Is it better? Learn to score free-text answers and put honest error bars on the result.
What you’ll explore
- Evaluation connects task goals to metrics, baselines, and uncertainty; accuracy, calibration, coverage, and cost answer different questions and can move in different directions.
GO TO THE SOURCE
Original explanations, connected to the research.
SQuAD: 100,000+ Questions for Machine Comprehension of Text (Rajpurkar et al., 2016), source of the exact-match and token-F1 conventionsAdding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (Miller, 2024)Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design (Sclar et al., 2023)Holistic Evaluation of Language Models (HELM) (Liang et al., 2022)On Calibration of Modern Neural Networks (Guo et al., 2017)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.