M38.4 CONNECT THE MECHANISM
Train and check probabilities
A model that says "90% sure" and is right 60% of the time is dangerous to route by. See how training can reward useful probabilities, how one fitted number changes confidence, and why the result still has to be checked.
LESSON OVERVIEW24 min lesson
Lesson overview
A model that says "90% sure" and is right 60% of the time is dangerous to route by. See how training can reward useful probabilities, how one fitted number changes confidence, and why the result still has to be checked.
What you’ll explore
- Supervised cross-entropy is the negative log score, and a strictly proper scoring rule (log, Brier, and for ordered answers the ranked probability score) is maximized in expectation by reporting the true distribution, while accuracy and linear scores are not. Training on a proper score encourages useful probabilities without guaranteeing them, a fitted temperature changes confidence but not ranking, and calibration must be checked on held-out data, including the cases the policy will automate.
GO TO THE SOURCE
Original explanations, connected to the research.
Laya model card (Convai Innovations, Hugging Face)Laya BENCHMARKS.md (Convai Innovations, GitHub)Nimble training guide (Bespoke Labs, GitHub)CLM-v0.1-8B model card (Hugging Face)Introducing System One models and Jev (TypeSafe AI blog, September 2026)TypeSafe AI documentation, confidence formulasStrictly Proper Scoring Rules, Prediction, and Estimation (Gneiting and Raftery, 2007)On Calibration of Modern Neural Networks (Guo et al., 2017)DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (Shao et al., 2024), which introduced GRPOSuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.