At the end of this post you will get a cheat sheet of the 10 common gradient descent optimisation algorithms.
Using more readable notations, I will walk you through how the vanilla stochastic gradient descent slowly evolved into the popular Adam optimiser and others. I also came out with an ‘evolutionary map’ of the optimisers to visualise this.
The motivation for writing this post is that there is a lack of simple-to-read equations for parameter update and a compiled list of these optimisers.
Hopefully this benefits the community.