Context window and the KV cache
Why is the real cost of long context memory rather than computation? And how are million token windows possible at all?
2 steps
115 XP
A free account is needed
Start the lesson →
Sources
- Pope, R. et al. 2023 · Efficiently Scaling Transformer Inference · MLSys 2023
- Ainslie, J. et al. 2023 · GQA: Training Generalized Multi-Query Transformer Models · EMNLP 2023
- Kwon, W. et al. 2023 · Efficient Memory Management for LLM Serving with PagedAttention (vLLM) · SOSP 2023
- Dao, T. et al. 2022 · FlashAttention: Fast and Memory-Efficient Exact Attention · NeurIPS 2022
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons