ML Academy · Track 2 · Deep Learning

Exploding gradients: the problem is the tail, not the average

The gradient norm is not a number, it is a distribution. Most batches are fine; a handful of them wreck training in a single step.

4 steps 175 XP A free account is needed
Start the lesson →

Sources

ML Academy · an interactive machine learning course that runs in your browser · All lessons