(Edit: just to be clear here, I’m not saying I expect the whole field is full of dummies who missed something obvious or something like that, I don’t know much at all about machine learning so I’m sure I’m missing something).
(Edit: just to be clear here, I’m not saying I expect the whole field is full of dummies who missed something obvious or something like that, I don’t know much at all about machine learning so I’m sure I’m missing something).
I'm not saying it's impossible but if resources allow it makes a lot of sense to start with the biggest model you can still train. Especially since for whatever reason things seem to get a lot easier if you simply throw more computing power at it (kind of like how no matter how advanced your caching algorithm it's not going to be more than 2 times faster than the simplest LRU algorithm with double the amount of cache).
It sounds like there is a lot of work happening on sparse networks now, so it’ll be interesting to see how this changes in the near future.
There has been a lot of work on sparsity and discovering sparse subnetworks in trained dense networks. And intel even proposed some alternative cpu friendly architectures and torch/tf and gpus are starting to do okay with sparse matrixes so thing are changing.