I don't have much help to offer, but just to echo your experience... at my group we have tried to train Transformers from scratch for various NLP tasks and we always have been hit with them being extremely brittle, and BiLSTMs working better. We only succeeded by following a pre-established recipe (e.g. training a BERT model from scratch for a new language, where the architecture, parameters and tasks are as in BERT), or of course by fine-tuning existing models, but just throwing some layers at a problem and training them from scratch... nope, won't work without arcane knowledge that doesn't seem to be written anywhere accessible. This is one of the reasons why I dislike Transformers and I root for the likes of RWKV to take the throne.