M13.8 CONNECT THE MECHANISM
Judge quality across time, languages, and delay
Is 3 cups of error good? Is a 7% WER good if one language sits at 30%? Learn the scores forecasters and speech teams trust, and what each one hides.
LESSON OVERVIEW14 min lesson
Lesson overview
Is 3 cups of error good? Is a 7% WER good if one language sits at 30%? Learn the scores forecasters and speech teams trust, and what each one hides.
What you’ll explore
- Sequence evaluation must consider alignment, meaning, robustness, timing, and subgroup coverage; a single aggregate metric rarely captures the full user task.
GO TO THE SOURCE
Original explanations, connected to the research.
Another Look at Measures of Forecast Accuracy (Hyndman & Koehler, 2006)Forecasting: Principles and Practice, 3rd edition, §5.8 Evaluating point forecast accuracy (Hyndman & Athanasopoulos, 2021)ITU-T Recommendation P.800: Methods for subjective determination of transmission qualityNatural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions (Shen et al., 2018)Speech and Language Processing, 3rd edition draft (Jurafsky & Martin)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.