SGD, Momentum, Adam
The same loss surface, the same starting point. The only difference is how they take the step, and that difference turns into an 11 fold speedup.
1 steps
55 XP
A free account is needed
Start the lesson →
Sources
- Polyak, B. T. 1964 · Some Methods of Speeding up the Convergence of Iteration Methods · USSR Comp. Math. and Math. Physics, 4(5)
- Kingma, D. & Ba, J. 2015 · Adam: A Method for Stochastic Optimization · ICLR 2015
- Ruder, S. 2016 · An Overview of Gradient Descent Optimization Algorithms · arXiv:1609.04747
- Loshchilov, I. & Hutter, F. 2019 · Decoupled Weight Decay Regularization (AdamW) · ICLR 2019
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons