M19.7 CONNECT THE MECHANISM
Can the model use a long prompt?
A model that accepts 128K tokens can still miss a fact on page 150. See how RoPE gets stretched to reach long contexts, why the middle gets lost, and how to test honestly.
LESSON OVERVIEW14 min lesson
Lesson overview
A model that accepts 128K tokens can still miss a fact on page 150. See how RoPE gets stretched to reach long contexts, why the middle gets lost, and how to test honestly.
What you’ll explore
- Long-context capability combines positional behavior, training exposure, information access, and serving resources; evaluation must test retrieval, distraction, and reasoning across positions.
GO TO THE SOURCE
Original explanations, connected to the research.
Lost in the Middle: How Language Models Use Long Contexts (Liu et al., 2023)Extending Context Window of Large Language Models via Positional Interpolation (Chen et al., 2023)NTK-Aware Scaled RoPE (bloc97, 2023, r/LocalLLaMA post)YaRN: Efficient Context Window Extension of Large Language Models (Peng et al., 2023)RULER: What's the Real Context Size of Your Long-Context Language Models? (Hsieh et al., 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.