As an aside, I'm trying to train transformers for some classification tasks on audio data. The models are "small" (like 1M-15M params at most) and I find they are very finicky to train. Below 1M parameters I find them hard to train at all. I have thrown all sorts of learning rate schedules at them and the best I can get is the network learns for a bit and then plateaus, after which I can't do anything to get them out of that minima. Training an LSTM/GRU on the same data gives me a much better loss value.
I couldn't find many papers on training transformers at that scale. The only one I was able to find was MS's TinyStories [0], but that paper didn't delve much into how they trained the models and whether they trained from scratch or distilled from a larger model.
At those scales, I find LSTMs and CNNs are a lot more stable. The few online threads I've found comparing LSTMs and Transformers had the same thing to say - Transformers need a lot more data and model size to achieve parity and exceed LSTMs/GRUs/CNNs, maybe because the inductive bias provided is hard to beat at those scales. Others can comment on what they've seen.