Direct initialization of transformers using larger pretrained ones
arxiv.org
arxiv.org
You can init deeper layers by copying the weights from previous layers. Not a big improvement in training time, but it does reduce the initial loss!
A good test would be to apply this technique, but then train on a totally different kind of data - for example take a text LLM, apply this technique, but then train on audio/music data and see if this technique reduces training time over random initialization.
We need something like this for "non-neural" things like transformers and normalization layers:
https://proceedings.mlr.press/v157/skorski21a/skorski21a.pdf
also pretty sure this would work because I know some text-image generators has the image network priority trained then frozen during e2e training
https://machinelearningmastery.com/weight-initialization-for...
A big breakthrough was Kaiming He's work on initialization for ReLU:
https://arxiv.org/abs/1502.01852
(Kaiming is the personal name, He is the family name. He Kaiming is the native Chinese order. Lots of citations use "Kaiming" by mistake.)
In short, I can't imagine ever having a training situation where this paper would be applicable.
> In the case of GPT-2 training, random initialization requires 64×10^9 tokens to reach a perplexity of 12, while weight subcloning accomplishes this in just 64×10^9 tokens, again demonstrating a 4× training speedup
... what?