M21.6 CONNECT THE MECHANISM
Spend extra generation on candidates, checks, and search
A model that's right 60% of the time can be right 68% of the time if it answers five times and votes, or 99% with a good checker. Here's the arithmetic, and the catch.
LESSON OVERVIEW16 min lesson
Lesson overview
A model that's right 60% of the time can be right 68% of the time if it answers five times and votes, or 99% with a good checker. Here's the arithmetic, and the catch.
What you’ll explore
- Self-consistency, best-of-n selection, and tree search use inference compute to explore or select outputs; benefits depend on candidate diversity, evaluator quality, and the total budget.
GO TO THE SOURCE
Original explanations, connected to the research.
Self-Consistency Improves Chain of Thought Reasoning in Language Models (Wang et al., 2022)Tree of Thoughts: Deliberate Problem Solving with Large Language Models (Yao et al., 2023)Training Verifiers to Solve Math Word Problems (Cobbe et al., 2021)Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (Snell et al., 2024)DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek-AI, 2025)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.