Quantisation: the price of shrinking a model
Storing weights with fewer bits cuts memory several fold. You will measure when that price is acceptable and when it is a disaster.
4 steps
225 XP
A free account is needed
Start the lesson →
Sources
- Jacob, B. et al. 2018 · Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference · CVPR 2018
- Dettmers, T. et al. 2022 · LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale · NeurIPS 2022
- Frantar, E. et al. 2023 · GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers · ICLR 2023
- Dettmers, T. et al. 2023 · QLoRA: Efficient Finetuning of Quantized LLMs · NeurIPS 2023
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons