Schedule-Free Learning – A New Way to Train
github.com
github.com
Validation accuracy: https://i.imgur.com/8ZtX7Rd.png
Train loss: https://i.imgur.com/o5XdQ29.png
Code: https://bpa.st/NVJQ (currently only runs on my computer, but not enough time to clean it up)
Note that this is just a toy benchmark with very little hyperparameter tuning. You could probably get similar results with most optimizers and an appropriate schedule. Nevertheless, I appreciate every hyperparameter that I do not have to set manually.
In summary, this seems to be a promising optimizer. I'll add it to my list of optimizers to try for new deep learning projects.
Can you share the list of your go to optimizers outside of the Adam family?
Sure! It depends a bit on what I'm doing.
If I want to optimize someone else's model, I start with Adam, because that's most-likely what the hyperparameters have been optimized for. Once I've verified that Adam works, I'll try other optimizers.
If I have very few parameters and don't care about overfitting, I try LBFGS, which usually gets to the local optimum the fastest. Note that this will likely find a sharp local optimum. For better generalization performance, you often prefer a wide optimum, so the model still works if there is a bit of drift in the data.
If I do not want to mess around with learning rates, I use Adafactor, which is a bit slower, but usually works okay without any tuning.
If I had very little memory available, I'd use SGD, but in my opinion it's not worth the hassle of tuning learning rate, momentum, dampening and weight decay. I'd rather use a smaller model if possible.
I usually do not train with extremely large batch sizes, but if I did, I'd try the optimizers which claim to work well for large batch sizes.
All in all, it probably does not matter too much which optimizer you are using, as long as you tuned it a little bit. Same goes for the model, loss functions, activation functions and all that other fluff.
What /is/ important is that you design your problem in such a way that it is as easy as possible to solve. For example, it is very difficult to read arbitrary hand-written text from an image. If you have control over where the data comes from, it would be better to write the text character by character into a printed grid with additional optical markers for image registration. Or even better, replace it with a multiple choice list. If there are not too many exceptional cases, an "other" option for manual review could be added. Often, automating 99 % of the work is more than good enough and it is better to keep a human in the loop to handle edge cases.
Secondly, control the data capture as strictly as possible. For example, use uniform lightning, place the object to recognize at exactly the same position, exclude disruptive elements, etc.
Lastly, data is king. If your training data does not match the test data, you can train all you want and still get garbage results. Either collect enough training data to cover all test cases or, if that is not possible from the start, retrain with new data regularly. Data augmentation might help to some degree, but it is impossible to predict everything.
Additionally, the optimizer does actually appear to have a kind of momentum, despite claims directly saying the contrary, and uses it with a nesterov-like step (line 2 of 3 in the inner loop). Finally, it is 'schedule-free' because the schedule is actually hardcoded into the algorithm itself -- 1./steps_taken which is not necessarily a rare learning rate schedule. This is a decently robust but sometimes suboptimal schedule, and I find it sketchy to make claims that it is 'schedule-free'. This also cripples the optimizer by tying performance to the number of steps taken -- which is potentially a problem if you are using any batchsize+lr scaling strategies as I understand.
There is a mixture of hype and substance here, and I wish the author was more straightforward with their approach and claims. I think there is the potential for a good "bolts-included" optimizer with some of the ideas being presented here, but the amount of overhyping and deception makes me not want to trust any of the following work coming.
Unfortunately, hype is what sells best on Twitter, and some of the claims being made here appear to be at the very best deceptive, and at the very worst, untrue. I could be wrong -- these are just my personal opinions from my own experience, but I do occasionally find myself distraught about the things that tend to catch wind in the technical news cycle.
-Fern
Their past research on D-Adpatation (won ICML best paper 2023) and their follow up work Prodigy all did worse / similar than AdamW, so maybe this works on CNNs, but does not on transformers - but for CNNs we have superconvergence.
I shall wait for their paper which will come in 1-2 months.
* Aaron et al's past work on D-Adaptation won a best ICML paper, with their follow up work being Prodigy - but both on transformers did similar or worse than AdamW. https://twitter.com/danielhanchen/status/1775547139248341125
* Superconvergence + LR range finder + Fast AI's Ranger21 optimizer was the goto optimizer for CNNs, and worked fabulously well, but on transformers, the learning rate range finder sadi 1e-3 was the best, whilst 1e-5 was better. However, the 1 cycle learning rate stuck. https://github.com/huggingface/transformers/issues/16013
* A huge issue is this needs tuning??! But how about a well tuned AdamW? Eg see https://twitter.com/kellerjordan0/status/1776716388037529843 which outperformed it using a tuned SGD.
* I'm just a little bit reserved for now since the author themselves aren't providing any transformer benchmarks, nor have they compared their CNN baselines to superconvergence, which is the goto standard for fast training for CNNs. Likewise https://parameterfree.com/2023/08/30/yet-another-icml-award-... wasn't pleasant.
The issue I have with Schedule-Free is you need tuning, but a well tuned SOTA can already skyrocket past a plain tuned AdamW.
Hence why I think it might be hard to accurately compare them, likely SGD and Adam/AdamW are going to have better potential top ends but are going to get more thrashed in public comparisons vs an optimizer that seems to perform more flatly overall. Aaron works at FAIR so I am assuming that he knows this, I reached out with some concerns on my end a little bit before he published the optimizer but didn't hear back either unfortunately.
96 legitimately is pretty hard, i struggled doing it even in 2 minutes, so seeing it in 45 seconds is crazy. definitely gets exponentially harder for every fraction of a percent, so i think that's a pretty big achievement to hit :D
> (an) author here: paper will likely be coming out in O(month)
Ug. I'm adding "O(month)" to my list of bootless metaphors.
Why? (1) Because in Big-O notation, O(month) would equal O(day), which is not the intended meaning in the comment above; (2) It is non-sensical; one would never say e.g. "the run-time of an algorithm is O(seconds)" -- we write some kind of input inside the parens, not the output
Anyhow, we already have the words roughly and about; e.g. "about a month".
Feel free to call me pedantic, but words matter.
Stochastic Weight Averaging (Izmailov et al 2018) https://arxiv.org/abs/1803.05407
Latest Weight Averaging (Kaddour 2022) https://arxiv.org/abs/2209.14981
Latest Weight Averaging? (Sanyal et al 2023) https://arxiv.org/abs/2311.16294
Cyclic Learning Rates (Portes et al 2022) https://arxiv.org/abs/2206.00832
Exponential Moving Average? (Zhanghan? et al 2019) https://arxiv.org/abs/1909.01804
Here's another person in stack exchange who figured this out: https://stackoverflow.com/a/44844544
Pytorch and TG both use a default 1e-8.
Sounds like variable epsilon is optimal, that's instead of learning rate, or both together. Would be nice if this can somehow be algorithmically regulated in generic way.