But the main instructions were two types - summarise x, and translate y to z.
Translation in many ways was the root of seeing that multiple tasks/instructions could fit a model and not just in the context of multiple training loss methods, like NSP vs Gap filling(a modus of training difference, not actual task itself).
And the special tokens for BERT etc trace their origins back to enabling the task above and positional embeddings(and encodings). Which largely trace their origin to translation work.
Every other use case is basically an accidental feature
The "Attention Is All You Need" paper frames the problem and contribution as:
"The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely."
Translation is "just an experiment" for the general architecture and they study other tasks in the paper too.
In a sense, all applications of basic research are "accidental".
Nice joke
Rather that they are in many settings overhyped and overprescribed.
Given they tripped and fell over a money printing machine and then chose to lower their API prices, it would be pretty surprising (but not impossible) if their API prices are currently subsidised.