DeepNet: Scaling Transformers to 1k Layers
arxiv.org
arxiv.org
The fact that layers themselves are narrower means that training and evaluation of the NN is also much faster.
Since the network will have seen far more english text than text of other languages, it suggests that performance on limited training data is more improved.
I am concerned that we are approaching an age where only massive companies and research groups with tons of GPU resources will be able to train next-gen models.
We were already there in 2020.
Huggingface is also worthy of great praise, making the models accessible and lowering the barrier to actual commercial use.
Lastly, Google deserves props for Colab. You can run models that are 80% as good as the ones which initially cost millions of usd to produce. Even the free tier is excellent, and I've pointed various people to them as a resource for developing practical hands-on experience with machine learning.
This is a wonderful ai summer we're having, appreciate it while you can!
Oh, don't be. We are already at this stage.
The argument is that otherwise, the gradient magnitude for the lower layers becomes too big. Which intuitively makes sense, because due to the residual connections, all error signals from the upper layers will end up at the lower layers.
I wonder why this is apparently not a problem for ResNet.
The training and decoding runtime and memory consumption are what is relevant. And the number of parameters is not really connected to that.
E.g. I assume this DeepNet with 3.2B params is slower and requires more memory than the M2M-100 with 12B params.
The number of layers are more connected to both runtime and memory consumption.
So regardless of the actual net topology, number of layers etc you should see a rough scaling in the number of parameters, in an ideal case with no other overhead (you can write code that scales badly or weirdly). I didn't look at this DeepNet, maybe it will run slower like you say due to some other choice of architecture..
Also yes an architecture that requires 1000x more training iterations will of course be worse at the training stage even if it has 100x fewer parameters, but maybe the resulting inference model is much faster. Lots of stuff to consider :)
You could also share e.g. all the parameters in every layer (like Universal Transformer), and thus reduce the number of params drastically, but without any change in computing time, and only negligible difference in memory consumption because the hidden activations take most of the memory, not the params.
I naively see the transformer architecture as a generalization of other common topologies like convnets, i.e. you can arrange the token stream and the attention heads (if you have enough of them) to emulate a 2D convnet if you want even though the first uses where for 1D language token streams, and so the transformer arch with proper training can converge to a convnet arch if that is the best solution. Or is this a completely backward generalization? :)
They have a section with pseudocode and a tldr explanation.