Mixture of experts: the cheap way to add parameters
The scaling laws lesson showed the price of growing. MoE exists precisely to break that price.
4 steps
220 XP
A free account is needed
Start the lesson →
Sources
- Shazeer, N. et al. 2017 · Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer · ICLR 2017
- Fedus, W., Zoph, B. & Shazeer, N. 2022 · Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity · JMLR, 23(120)
- Jacobs, R. A. et al. 1991 · Adaptive Mixtures of Local Experts · Neural Computation, 3(1)
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons