YOUR FOUNDATION-FIRST SYLLABUS
Language models
Begin with everyday examples of inputs, data, and learning. Then follow how a language model learns from text, writes a response, and becomes part of a chatbot.
0 of 99 core ideas demonstrated98 lessons1345 estimated minutes remaining
Already know some of this?
Go straight to an idea’s knowledge check. Passing its different checks without hints carries that evidence into this path. Checking an advanced idea does not award its prerequisites; you can still explore them.
Check what I already knowTHE IDEAS, IN ORDER
Shared credit is automaticBuild your understanding.
Chapter counts use compatible knowledge-check evidence. Reading a lesson does not mark its ideas as demonstrated.
Chapter 1 · Lessons 1–7 Meet an AI system Start with an everyday task, then follow its inputs, model, and output. 7 lessons · 0 of 7 core ideas demonstrated Up next
01
What is AI? Start with an everyday task Common Ground · 11 min
To explore02 Inputs and outputs: what goes in and what comes out Common Ground · 13 min
To explore03 Data and examples: what a computer can learn from Common Ground · 10 min
To explore04 What is a model? A small rule inside a bigger app Common Ground · 10 min
To explore05 Weights: the adjustable numbers in a model Common Ground · 12 min
To explore06 Training and inference: changing a rule or using it Common Ground · 13 min
To explore07 Follow one small AI system from start to finish Common Ground · 9 min
To exploreChapter 2 · Lessons 8–14 How computers represent a problem Explore everyday AI ideas; how AI got here; how computers represent a problem; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
08
Different ways to solve the same problem Common Ground · 11 min
To explore09 How to tell whether an AI answer is useful Common Ground · 10 min
To explore10 The questions that shaped AI Common Ground · 15 min
To explore11 How a computer stores a picture or a sentence Learning mechanisms · 16 min
To explore12 Describe a computation so someone else can follow it Learning mechanisms · 11 min
To explore13 What a representation makes easy to learn Learning mechanisms · 12 min
To explore14 Different ways to learn Common Ground · 12 min
To exploreChapter 3 · Lessons 15–21 Turning language into model inputs Explore turning language into model inputs; math preparation when it is needed. 7 lessons · 0 of 7 core ideas demonstrated
15
Words, structure, meaning, and context Learning mechanisms · 12 min
To explore16 The numbers you actually need Math preparation · 10 min
To explore17 Probability without the mystery Math preparation · 10 min
To explore18 Represent text with counts and short contexts Learning mechanisms · 14 min
To explore19 How text becomes tokens Learning mechanisms · 13 min
To explore20 Learn the units used by a language model Learning mechanisms · 15 min
To explore21 Vectors: lists that work together Math preparation · 10 min
To exploreChapter 4 · Lessons 22–28 Following attention through a transformer Explore turning language into model inputs; following attention through a transformer; learning from useful data; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
22
From token IDs to learned vectors Learning mechanisms · 14 min
To explore23 Attention: queries, keys, and values Learning mechanisms · 13 min
To explore24 A function is a rule Math preparation · 9 min
To explore25 Powers and logarithms, one step at a time Math preparation · 11 min
To explore26 From attention scores to a causal mask Learning mechanisms · 17 min
To explore27 From the world to a dataset Learning mechanisms · 11 min
To explore28 Build a tiny prediction model Learning mechanisms · 11 min
To exploreChapter 5 · Lessons 29–35 Inside a neural network: Inside an artificial neuron Explore inside a neural network; turning language into model inputs; following attention through a transformer; and supporting ideas. 7 lessons · 0 of 8 core ideas demonstrated
29
Inside an artificial neuron Learning mechanisms · 14 min
To explore30 Why layers need more than multiplication Learning mechanisms · 12 min
To explore31 How a language model takes its next step Your system, connected · 15 min
To explore32 Follow one token through a transformer Your system, connected · 13 min
To explore33 How does a model improve a prediction? Learning mechanisms · 32 min
To explore34 Inside a language model training step Learning mechanisms · 14 min
To explore35 Averages over uncertain outcomes Math preparation · 9 min
To exploreChapter 6 · Lessons 36–42 Learning from useful data Explore learning from useful data; preparing and training a language model; math preparation when it is needed. 7 lessons · 0 of 7 core ideas demonstrated
36
What can a small experiment tell us? Math preparation · 10 min
To explore37 Clean data without erasing the problem Learning mechanisms · 12 min
To explore38 Protect the examples used to judge a model Learning mechanisms · 15 min
To explore39 Assemble token batches without leaking targets Learning mechanisms · 14 min
To explore40 Choose what a pretraining prediction means Learning mechanisms · 13 min
To explore41 Where did this dataset come from? Learning mechanisms · 12 min
To explore42 Build a traceable corpus before spending training compute Learning mechanisms · 19 min
To exploreChapter 7 · Lessons 43–49 Helping a model improve Explore helping a model improve; preparing and training a language model; inside a neural network; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
43
Learn from a sample of examples at each step Learning mechanisms · 11 min
To explore44 Choose how much the model sees of each source Learning mechanisms · 13 min
To explore45 Read an equation one symbol at a time Math preparation · 10 min
To explore46 A derivative is a local change Math preparation · 13 min
To explore47 The chain rule, one step at a time Math preparation · 13 min
To explore48 How an error reaches earlier weights Learning mechanisms · 11 min
To explore49 Follow shapes through a numerical layer Math preparation · 10 min
To exploreChapter 8 · Lessons 50–56 Inside a neural network: Trace inputs through layers to a loss Explore inside a neural network; helping a model improve; math preparation when it is needed. 7 lessons · 0 of 7 core ideas demonstrated
50
Trace inputs through layers to a loss Learning mechanisms · 17 min
To explore51 Let the graph carry derivatives Learning mechanisms · 12 min
To explore52 Several inputs, several rates of change Math preparation · 9 min
To explore53 Choose a direction and a step size Math preparation · 9 min
To explore54 How momentum and Adam use earlier gradients Learning mechanisms · 15 min
To explore55 Constrain a solution or change the update Learning mechanisms · 12 min
To explore56 When the computer cannot store the exact number Math preparation · 12 min
To exploreChapter 9 · Lessons 57–63 Inside a neural network: Help information and learning signals travel Explore helping a model improve; inside a neural network; following attention through a transformer; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
57
Keep signals and gradients in a usable range Learning mechanisms · 12 min
To explore58 Help information and learning signals travel Learning mechanisms · 12 min
To explore59 Connect data, gradients, and saved model state Learning mechanisms · 15 min
To explore60 Run several learned comparisons in parallel Learning mechanisms · 10 min
To explore61 Carry a state from one step to the next Learning mechanisms · 14 min
To explore62 Map one sequence into another Learning mechanisms · 17 min
To explore63 Match information flow to the learning objective Learning mechanisms · 13 min
To exploreChapter 10 · Lessons 64–70 Testing what a model has learned Explore preparing and training a language model; learning to predict; testing what a model has learned. 7 lessons · 0 of 7 core ideas demonstrated
64
Choose the model that pretraining will fit Learning mechanisms · 15 min
To explore65 Turn a linear score into a class probability Learning mechanisms · 11 min
To explore66 Control complexity without peeking at the answer Learning mechanisms · 12 min
To explore67 Will the model work on new examples? Learning mechanisms · 14 min
To explore68 Why a simpler model can predict better Learning mechanisms · 14 min
To explore69 Evaluate the procedure that chose the model Learning mechanisms · 14 min
To explore70 More capacity changes more than one thing Learning mechanisms · 14 min
To exploreChapter 11 · Lessons 71–76 Preparing and training a language model Explore preparing and training a language model; following attention through a transformer; adapting a model after pretraining; and supporting ideas. 6 lessons · 0 of 6 core ideas demonstrated
71
Allocate compute and preserve useful checkpoints Learning mechanisms · 15 min
To explore72 Distances, neighborhoods, and transformations Math preparation · 12 min
To explore73 Give a transformer information about order Learning mechanisms · 12 min
To explore74 Continue pretraining without losing sight of the starting model Learning mechanisms · 12 min
To explore75 Track what each adaptation stage changes Learning mechanisms · 15 min
To explore76 Learn response behavior from demonstrations Learning mechanisms · 14 min
To exploreChapter 12 · Lessons 77–83 Running a model in the real world Explore running a model in the real world; learning from useful data; learning to predict; and supporting ideas. 7 lessons · 0 of 7 core ideas demonstrated
77
Turn logits into a controlled generation procedure Learning mechanisms · 18 min
To explore78 From a prompt to a stream of tokens Your system, connected · 19 min
To explore79 Who is represented by the data? Learning mechanisms · 14 min
To explore80 Decide what success means before training Learning mechanisms · 12 min
To explore81 Look inside the average score Learning mechanisms · 14 min
To explore82 Probability, likelihood, and what is unknown Learning mechanisms · 11 min
To explore83 Update a belief using new evidence Learning mechanisms · 13 min
To exploreChapter 13 · Lessons 84–91 Checking reliability and behavior Explore working with uncertainty; testing what a model has learned; how computers represent a problem; and supporting ideas. 8 lessons · 0 of 8 core ideas demonstrated
84
Value a choice and its later consequences Math preparation · 12 min
To explore85 Turn a probability into a justified action Learning mechanisms · 11 min
To explore86 Change one thing and measure what follows Learning mechanisms · 13 min
To explore87 Repeat an experiment without confusing luck with truth Learning mechanisms · 11 min
To explore88 Make the result possible to inspect Learning mechanisms · 12 min
To explore89 Measure the behavior that the task actually needs Learning mechanisms · 14 min
To explore90 Build a test whose score supports the intended claim Learning mechanisms · 12 min
To explore91 Treat evaluators as fallible measurement instruments Learning mechanisms · 17 min
To exploreChapter 14 · Lessons 92–98 Building and operating an AI app Explore making responsible system decisions; building and operating an AI app; putting the pieces together. 7 lessons · 0 of 7 core ideas demonstrated
92
Decide whose problem the system solves and who bears its errors Learning mechanisms · 12 min
To explore93 Choose a useful problem before choosing a model Learning mechanisms · 14 min
To explore94 Keep training and serving connected to the same definitions Learning mechanisms · 12 min
To explore95 Release changes with evidence and a recovery path Learning mechanisms · 14 min
To explore96 Notice when the production task changes Learning mechanisms · 14 min
To explore97 Give people useful control over model-assisted work Learning mechanisms · 13 min
To explore98 Project: build a tiny next-token model and checkpoint ledger Learning mechanisms · 90 min
To exploreOptional extensions and comparisons (42)
These lessons deepen or compare the reference design. They do not add requirements to this path’s completion.
How to use this school without getting lostOptional orientation: how paths, side lessons, and practice fit together, plus study habits that make learning stick.Practice, progress, and knowing what you understandOptional orientation: use practice, mistakes, and spaced review to find out what you really understand.Misuse in practice: jailbreaks, deepfakes, and provenanceA familiar face or voice is no longer proof of who is speaking. Learn how misuse works and the checks that still hold.Prompting in practice: instructions, examples, and step-by-step requestsGet better answers from any chat assistant today: say what you want, show an example, give it the facts, and check what comes back.Score experts and choose a small setRoute tokens through selected expert networks and compare active work with total model capacity.Represent a sequence through an evolving hidden stateThe recurrence behind S4 and Mamba, which also runs as a convolution and is a classic forecasting tool.Trace the Llama 3.1 dense decoder lifecycleFollow the published Llama 3 training recipe step by step, and learn to tell what the report says from what people guess.Compare Qwen3 dense and MoE models with reasoning modesCompare documented dense and expert designs in the Qwen3 family.Separate DeepSeek-R1, R1-Zero, and distilled studentsConnect documented reasoning post-training stages to the mechanisms you have learned.Inspect the Olmo 3 model flowInspect a documented open training recipe and its reproducibility artifacts.Final project: defend an end-to-end AI system designDefend an architecture choice against explicit task and resource constraints.Project: compare dense and sparse model lifecyclesCompare dense and expert computation with a supplied quantitative fixture.Trace an example through a batchFollow array shapes through a batch: the everyday skill of reading and debugging model code.Why a small choice can create a huge searchWhy trying every possibility explodes, and why language models can't search every sentence.What makes an accelerator useful?Why GPUs made modern AI possible, and why more chips don't always mean more speed.Be a good scientist with a small examplePractice a scientist's habits (counterexamples, exhaustive tests, held-out data) on a model small enough to check completely.Compare sequence architectures with an honest budgetTransformers vs state-space and hybrid models: compare them honestly on quality, memory, and speed.Measure adaptation gains without overlooking regressionsDid fine-tuning help, or quietly break something? Measure before and after.Keep retrieved evidence current, permitted, and resistant to manipulationStale documents, leaked permissions, and poisoned sources: how retrieval-backed chatbots fail.Evaluate the completed task and the trajectory that produced itLanguage models increasingly act. Learn how to check that an agent actually finished the job.Understand what training and deployment can reveal about dataCan a model leak its training data? Memorization, extraction, and differential privacy.Evaluate claims against evidence and recognize missing informationWhy chatbots 'hallucinate', and how to measure and reduce confident wrong answers.Protect tools and sensitive data around the modelPrompt injection and tool misuse: keep a chatbot's text from gaining powers it shouldn't have.Project: trace a grounded answer and a controlled tool actionBuild a small grounded answerer with a guarded tool action, on paper or in code.Neurons, learning rules, and the Dartmouth proposalMeet Turing's test, the Dartmouth workshop, and a checkers program that outplayed its author.Symbols, games, and early language programsSee how Shakey planned with symbols and why ELIZA fooled people.Limits, expectations, and changing supportFind out what the XOR proof really showed and why AI funding froze twice.From knowledge engineering to learning from dataCompare hand-written expert rules with filters that learn from data.Why deeper networks became practicalSee the four ingredients behind AlexNet's 2012 landslide.Games, attention, and broadly reusable modelsTrace the path from AlphaGo's move 37 to ChatGPT.Read the shape of a loss landscapePicture training as walking downhill: valleys, saddles, and why neural networks train at all.Find the first broken assumption in a training runFind the first broken step when a training run goes wrong.Turn text into structured labels and relationsParse sentences into parts of speech, names, and relations: classic NLP that still powers pipelines.One vocabulary does not divide every language equallyWhy some languages cost more tokens, and why models stumble on numbers and spelling.Look back at the information needed for this stepWhy attention was invented: the bottleneck it removed and the idea behind the whole transformer.Rearrange attention around a compressed running summaryAttention that stays fast on very long inputs, and what it gives up to get there.Adapt a model by learning a smaller parameter updateLoRA: fine-tune a big language model on a laptop by learning a small update.Read proprietary reports without inventing hidden architectureLearn what companies' reports about GPT-4 and Gemini really disclose, and what they leave out.Project: explain AI through a historical concept mapDraw your own map connecting AI's big ideas from logic to large language models.Project: reproduce a result and audit an ablationRerun a small result yourself, remove one ingredient, and see which part really made the difference.Make a small experiment repeatableWrite small Python experiments you can rerun and trust.Follow a token through a mixture of expertsA quick preview of mixture-of-experts routing, the trick behind many of the largest models.