M28.2 CONNECT THE MECHANISM
Build a test whose score supports the intended claim
A support bot aces a 500-question quiz about phone plans, then fumbles real customers. Learn to build a benchmark that measures the right thing and stays honest once people start chasing it.
LESSON OVERVIEW12 min lesson
Lesson overview
A support bot aces a 500-question quiz about phone plans, then fumbles real customers. Learn to build a benchmark that measures the right thing and stays honest once people start chasing it.
What you’ll explore
- Specify benchmark tasks, sampling, scoring, resource conditions, and contamination checks before interpreting comparisons.
GO TO THE SOURCE
Original explanations, connected to the research.
Measuring Massive Multitask Language Understanding (MMLU) (Hendrycks et al., 2020)MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark (Wang et al., 2024)GPT-4 Technical Report (OpenAI, 2023)Language Models are Few-Shot Learners (Brown et al., 2020), appendix on 13-gram contamination checksBeyond the Imitation Game: BIG-bench (Srivastava et al., 2022), which introduced a canary stringDynabench: Rethinking Benchmarking in NLP (Kiela et al., 2021)LiveBench: A Challenging, Contamination-Limited LLM Benchmark (White et al., 2024)Liang et al. — Holistic Evaluation of Language ModelsSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.