Didn't "Attention Is All you Need" bill transformers primarily as a translation model?
But the main instructions were two types - summarise x, and translate y to z.
Translation in many ways was the root of seeing that multiple tasks/instructions could fit a model and not just in the context of multiple training loss methods, like NSP vs Gap filling(a modus of training difference, not actual task itself).
And the special tokens for BERT etc trace their origins back to enabling the task above and positional embeddings(and encodings). Which largely trace their origin to translation work.