One transformer block, end to end
You have seen attention. Now everything around it, and where 7 billion parameters actually go.
2 steps
115 XP
A free account is needed
Start the lesson →
Sources
- Vaswani, A. et al. 2017 · Attention Is All You Need · NeurIPS 2017
- He, K. et al. 2016 · Deep Residual Learning for Image Recognition (artık bağlantılar) · CVPR 2016
- Zhang, B. & Sennrich, R. 2019 · Root Mean Square Layer Normalization (RMSNorm) · NeurIPS 2019
- Shazeer, N. 2020 · GLU Variants Improve Transformer (SwiGLU) · arXiv:2002.05202
- Touvron, H. et al. 2023 · LLaMA: Open and Efficient Foundation Language Models · arXiv:2302.13971
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons