An overview of gradient descent optimization algorithms (2016)
ruder.io
ruder.io
Messing with optimizers is one of the ways to enter hyperparameter hell: it’s like legacy code but on steroids because changing it only breaks your training code stochastically. Much better to stop worrying and love AdamW.
Edit: Demon involves decaying the momentum parameter over time, which introduces a new schedule or formula for how momentum should be reduced during training. That can feel like additional complexity or a potential hyperparameter rabbit hole. Teams trying to ship products quickly often avoid adding new hyperparameters unless the gains are decisive.
In practice, clever optimisation algorithms that use the 2nd derivative won't actually form this matrix.
For non deep learning applications, Nelder-Mead saved my butt a fees times
https://docs.scipy.org/doc/scipy/reference/optimize.html#loc...
For nastier optimization problems there are lots of other options, including evolutionary algorithms and Bayesian optimization:
An MLE should be able to look up and understand the differences between optimizers but memorizing that information is extremely low priority compared with other information they might be asked.
Unless someone had a very good reason I would consider it weird to use anything other than AdamW. The compute you could save on a slightly better optimizer pale in comparison to the time you will spend debugging an opaque training bug.
As a model is trained, the gradient variance typically falls.
Those optimizers all work to reduce the variance of the updates in various ways.