Gradient Descent: The Ultimate Optimizer
arxiv.org
arxiv.org
Made me laugh. Can't believe this hasn't been done yet. Maybe I should stop looking at ML research upon a pedestal (regardless, OpenAI's PPO and TRPO papers were amazing, not to mention their Learning with Dexterity one)
The problem is that, well, to backpropagate through a hyperparameter, you would need to, say, track how it affects every iteration throughout the entire training run, rather than simply tracking one parameter through a single iteration on a single datapoint. And it's difficult enough to do gradient descent on a single hyperparameter, so it hardly helps to start talking about doing entire stacks! If you can't really do one, doing ad infinitum probably isn't going to work well either.
If you look at their experiment, they're doing a 1 hidden-layer FC NN on MNIST. (Honestly, I'm a little surprised that stacking hyperparams works even on that small a problem.)
I'm surprised by the lack of any bigger experiments (imagenet would be nice, but even cifar10 would help) given how computationally light this was. Also surprised that an 11ish author stanford paper did their experiments on 1 cpu.
Choosing hyperparameters from minor tuning or literature review or domain expertise or whatever is just a form of creating a prior.
Has been done before
Did they get stuck in a local optimum?
https://arxiv.org/abs/1803.02021
Due to the fact that larger LRs can result in worse immediate performance but better progress in low signal-to-noise ratio directions.
However, I think it's a different effect- even purely in terms of optimizing the training loss, on a quadratic (with noisy gradients), the short-horizon bias effect exists.
Relevant work:
1. What's one of the most important hyperparameters? Architecture: https://arxiv.org/abs/1806.09055
2. Dropout: https://papers.nips.cc/paper/5032-adaptive-dropout-for-train...
3. Learning activation thresholds: https://datascience.stackexchange.com/questions/18583/what-i...
edit: changed to /abs/ instead of /pdf/ link
What IS interesting, is that the problems all end up in the same place after a few levels of recursion.
ie. It's not a problem for the class of problems they are optimizing.
I've been out of that domain for awhile... but when I did a lot of optimizing / fitting in a high dimensional space with a function that had lots of local optima... I found the downhill simplex method with simulated annealing was most effective / robust.
For a local minimum to exist, that means that the function we're looking at has to have a local minimum when sliced in every dimension.
In our intuition, this is relatively common, because we usually imagine 2d manifolds embedded in 3-space (ie, "hill climbing"). And when there's only 2 dimensions, it's not hard to have a "bowl", where the function curves upward in both dimensions.
In fact, a "saddle point", where one of the dimensions curves up and the other curves down, feels "unlikely" to most peoples' intuition, because we don't often run into physical objects with that shape.
But, saddle points in higher dimensions "should be" far more common than local minima. You only need one dimension to curve downward to have a "hole" where the "water" can escape and continue flowing downhill.
Gradient descent easily gets stuck in 2 dimensions, and so we are often surprised by how well it performs in higher dimensions.
I think Don’t you, too, sometimes? As a response to that has a lot more going for it than it first looks like.
I mean, I look at it, and the choices I make in my life, and think, yeah.... I really do.
This reminds me of using GD for tuning PID (which has been done for a while).
https://www.researchgate.net/publication/287359696_PID_contr...
https://github.com/Rainymood/Gradient-Descent-The-Ultimate-O...
This inspired my consulting company / ML q+a site name MetaOptimize and it’s motto: optimizing the process of optimizing the process of...
Maybe they chopped the long tail of authors?