How to train large deep learning models at a well founded startup*
Everything described here is absolutely not affordable by bootstrappers and startups with little funding, unless the model to train is not that deep.
How to train large deep learning models at a well founded startup*
Everything described here is absolutely not affordable by bootstrappers and startups with little funding, unless the model to train is not that deep.
Other tips not mentioned in the article:
1. Tune your hyper parameters on a subset of the data.
2. Validate new methods with smaller models on public datasets.
3. Tune models instead of training from scratch (either public models or your previously trained ones).
1. if you choose the wrong subset, you'll find a non optimum local min
2. still risk dead ends when expanding the model and lengthen the time to finding that out
3. a lot of public models are made from inaccurate datasets, so beware
Overall you have to start somewhere though, and your points are still valid.
2. Sure, but for most new hacks like mixup, randaugment and etc the results usually transfer over. Problem with deep learning is that most of the new results don't replicate so it's good to have a way to quickly validate things.
3. The lower level features are usually pretty data agnostic and transfer well to new tasks.
Unless we’re talking about the optima of test error.