M20.3 CONNECT THE MECHANISM
Adapt a model by learning a smaller parameter update
Full fine-tuning of a 7-billion-parameter model needs over 100 GB of memory. LoRA trains 0.06% of the numbers instead, and QLoRA squeezes the rest onto one GPU.
LESSON OVERVIEW13 min lesson
Lesson overview
Full fine-tuning of a 7-billion-parameter model needs over 100 GB of memory. LoRA trains 0.06% of the numbers instead, and QLoRA squeezes the rest onto one GPU.
What you’ll explore
- LoRA trains a small low-rank update instead of full weight matrices; QLoRA does the same on top of a 4-bit frozen base to save memory.
GO TO THE SOURCE
Original explanations, connected to the research.
LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., 2021)QLoRA: Efficient Finetuning of Quantized LLMs (Dettmers et al., 2023)Hugging Face PEFT documentation — LoRASuggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.