Exploding gradients: the problem is the tail, not the average
The gradient norm is not a number, it is a distribution. Most batches are fine; a handful of them wreck training in a single step.
4 steps
175 XP
A free account is needed
Start the lesson →
Sources
- Pascanu, R., Mikolov, T. & Bengio, Y. 2013 · On the Difficulty of Training Recurrent Neural Networks · ICML 2013
- Bengio, Y., Simard, P. & Frasconi, P. 1994 · Learning Long-Term Dependencies with Gradient Descent is Difficult · IEEE Trans. Neural Networks 5(2)
- Zhang, J. et al. 2020 · Why Gradient Clipping Accelerates Training · ICLR 2020
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons