M27.5 CONNECT THE MECHANISM
Reduce deployment cost while measuring what changes
Store each weight in 4 bits instead of 16 and an 8B model drops from 16 GB to about 5. Learn where the rounding error goes, and how GPTQ and AWQ hide it.
LESSON OVERVIEW14 min lesson
Lesson overview
Store each weight in 4 bits instead of 16 and an 8B model drops from 16 GB to about 5. Learn where the rounding error goes, and how GPTQ and AWQ hide it.
What you’ll explore
- Quantization lowers numerical precision, pruning removes selected parameters or structures, and distillation trains a different model; each changes costs and can affect quality in workload-dependent ways.
GO TO THE SOURCE
Original explanations, connected to the research.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (Frantar et al., 2022)AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (Lin et al., 2023)LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (Dettmers et al., 2022)SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models (Xiao et al., 2022)Learning both Weights and Connections for Efficient Neural Networks (Han et al., 2015)Distilling the Knowledge in a Neural Network (Hinton, Vinyals & Dean, 2015)DistilBERT, a distilled version of BERT (Sanh et al., 2019)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.