Back to the lesson libraryMECHANISM · 19 MIN
M27.2 CONNECT THE MECHANISM

From a prompt to a stream of tokens

Why does a chatbot pause, then type fast? Time the two phases of generation, count the bytes of the cache between them, and reuse the work on a shared prompt.

LESSON OVERVIEW19 min lesson

Lesson overview

Why does a chatbot pause, then type fast? Time the two phases of generation, count the bytes of the cache between them, and reuse the work on a shared prompt.

What you’ll explore

  • Explain prefill and cached decoding, including how request-specific key/value activations differ from learned model weights.
Suggest a correction

A precise note can make an explanation better.

Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.

The file includes this note, the scene title, and lesson metadata. Your saved progress and quiz responses are excluded. Download before leaving or reloading to keep your note.