All learning paths
YOUR FOUNDATION-FIRST SYLLABUS

Language models

Begin with everyday examples of inputs, data, and learning. Then follow how a language model learns from text, writes a response, and becomes part of a chatbot.

0 of 99 core ideas demonstrated98 lessons1345 estimated minutes remaining

Next: What is AI? Start with an everyday task · 11 min

Already know some of this?

Go straight to an idea’s knowledge check. Passing its different checks without hints carries that evidence into this path. Checking an advanced idea does not award its prerequisites; you can still explore them.

Check what I already know
THE IDEAS, IN ORDER

Build your understanding.

Shared credit is automatic

14 short chapters, in prerequisite order. Open any chapter to explore.

Chapter counts use compatible knowledge-check evidence. Reading a lesson does not mark its ideas as demonstrated.

Chapter 1 · Lessons 1–7 Meet an AI system Start with an everyday task, then follow its inputs, model, and output. 7 lessons · 0 of 7 core ideas demonstrated Up next
Chapter 2 · Lessons 8–14 How computers represent a problem Explore everyday AI ideas; how AI got here; how computers represent a problem; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
Chapter 3 · Lessons 15–21 Turning language into model inputs Explore turning language into model inputs; math preparation when it is needed. 7 lessons · 0 of 7 core ideas demonstrated
Chapter 4 · Lessons 22–28 Following attention through a transformer Explore turning language into model inputs; following attention through a transformer; learning from useful data; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
Chapter 5 · Lessons 29–35 Inside a neural network: Inside an artificial neuron Explore inside a neural network; turning language into model inputs; following attention through a transformer; and supporting ideas. 7 lessons · 0 of 8 core ideas demonstrated
Chapter 6 · Lessons 36–42 Learning from useful data Explore learning from useful data; preparing and training a language model; math preparation when it is needed. 7 lessons · 0 of 7 core ideas demonstrated
Chapter 7 · Lessons 43–49 Helping a model improve Explore helping a model improve; preparing and training a language model; inside a neural network; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
Chapter 8 · Lessons 50–56 Inside a neural network: Trace inputs through layers to a loss Explore inside a neural network; helping a model improve; math preparation when it is needed. 7 lessons · 0 of 7 core ideas demonstrated
Chapter 9 · Lessons 57–63 Inside a neural network: Help information and learning signals travel Explore helping a model improve; inside a neural network; following attention through a transformer; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
Chapter 10 · Lessons 64–70 Testing what a model has learned Explore preparing and training a language model; learning to predict; testing what a model has learned. 7 lessons · 0 of 7 core ideas demonstrated
Chapter 11 · Lessons 71–76 Preparing and training a language model Explore preparing and training a language model; following attention through a transformer; adapting a model after pretraining; and supporting ideas. 6 lessons · 0 of 6 core ideas demonstrated
Chapter 12 · Lessons 77–83 Running a model in the real world Explore running a model in the real world; learning from useful data; learning to predict; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
Chapter 13 · Lessons 84–91 Checking reliability and behavior Explore working with uncertainty; testing what a model has learned; how computers represent a problem; and supporting ideas. 8 lessons · 0 of 8 core ideas demonstrated
Chapter 14 · Lessons 92–98 Building and operating an AI app Explore making responsible system decisions; building and operating an AI app; putting the pieces together. 7 lessons · 0 of 7 core ideas demonstrated
Optional extensions and comparisons (42)

These lessons deepen or compare the reference design. They do not add requirements to this path’s completion.

How to use this school without getting lostOptional orientation: how paths, side lessons, and practice fit together, plus study habits that make learning stick.Practice, progress, and knowing what you understandOptional orientation: use practice, mistakes, and spaced review to find out what you really understand.Misuse in practice: jailbreaks, deepfakes, and provenanceA familiar face or voice is no longer proof of who is speaking. Learn how misuse works and the checks that still hold.Prompting in practice: instructions, examples, and step-by-step requestsGet better answers from any chat assistant today: say what you want, show an example, give it the facts, and check what comes back.Score experts and choose a small setRoute tokens through selected expert networks and compare active work with total model capacity.Represent a sequence through an evolving hidden stateThe recurrence behind S4 and Mamba, which also runs as a convolution and is a classic forecasting tool.Trace the Llama 3.1 dense decoder lifecycleFollow the published Llama 3 training recipe step by step, and learn to tell what the report says from what people guess.Compare Qwen3 dense and MoE models with reasoning modesCompare documented dense and expert designs in the Qwen3 family.Separate DeepSeek-R1, R1-Zero, and distilled studentsConnect documented reasoning post-training stages to the mechanisms you have learned.Inspect the Olmo 3 model flowInspect a documented open training recipe and its reproducibility artifacts.Final project: defend an end-to-end AI system designDefend an architecture choice against explicit task and resource constraints.Project: compare dense and sparse model lifecyclesCompare dense and expert computation with a supplied quantitative fixture.Trace an example through a batchFollow array shapes through a batch: the everyday skill of reading and debugging model code.Why a small choice can create a huge searchWhy trying every possibility explodes, and why language models can't search every sentence.What makes an accelerator useful?Why GPUs made modern AI possible, and why more chips don't always mean more speed.Be a good scientist with a small examplePractice a scientist's habits (counterexamples, exhaustive tests, held-out data) on a model small enough to check completely.Compare sequence architectures with an honest budgetTransformers vs state-space and hybrid models: compare them honestly on quality, memory, and speed.Measure adaptation gains without overlooking regressionsDid fine-tuning help, or quietly break something? Measure before and after.Keep retrieved evidence current, permitted, and resistant to manipulationStale documents, leaked permissions, and poisoned sources: how retrieval-backed chatbots fail.Evaluate the completed task and the trajectory that produced itLanguage models increasingly act. Learn how to check that an agent actually finished the job.Understand what training and deployment can reveal about dataCan a model leak its training data? Memorization, extraction, and differential privacy.Evaluate claims against evidence and recognize missing informationWhy chatbots 'hallucinate', and how to measure and reduce confident wrong answers.Protect tools and sensitive data around the modelPrompt injection and tool misuse: keep a chatbot's text from gaining powers it shouldn't have.Project: trace a grounded answer and a controlled tool actionBuild a small grounded answerer with a guarded tool action, on paper or in code.Neurons, learning rules, and the Dartmouth proposalMeet Turing's test, the Dartmouth workshop, and a checkers program that outplayed its author.Symbols, games, and early language programsSee how Shakey planned with symbols and why ELIZA fooled people.Limits, expectations, and changing supportFind out what the XOR proof really showed and why AI funding froze twice.From knowledge engineering to learning from dataCompare hand-written expert rules with filters that learn from data.Why deeper networks became practicalSee the four ingredients behind AlexNet's 2012 landslide.Games, attention, and broadly reusable modelsTrace the path from AlphaGo's move 37 to ChatGPT.Read the shape of a loss landscapePicture training as walking downhill: valleys, saddles, and why neural networks train at all.Find the first broken assumption in a training runFind the first broken step when a training run goes wrong.Turn text into structured labels and relationsParse sentences into parts of speech, names, and relations: classic NLP that still powers pipelines.One vocabulary does not divide every language equallyWhy some languages cost more tokens, and why models stumble on numbers and spelling.Look back at the information needed for this stepWhy attention was invented: the bottleneck it removed and the idea behind the whole transformer.Rearrange attention around a compressed running summaryAttention that stays fast on very long inputs, and what it gives up to get there.Adapt a model by learning a smaller parameter updateLoRA: fine-tune a big language model on a laptop by learning a small update.Read proprietary reports without inventing hidden architectureLearn what companies' reports about GPT-4 and Gemini really disclose, and what they leave out.Project: explain AI through a historical concept mapDraw your own map connecting AI's big ideas from logic to large language models.Project: reproduce a result and audit an ablationRerun a small result yourself, remove one ingredient, and see which part really made the difference.Make a small experiment repeatableWrite small Python experiments you can rerun and trust.Follow a token through a mixture of expertsA quick preview of mixture-of-experts routing, the trick behind many of the largest models.