I would argue that input scaling is not fundamental to Transformers.
Recurrent neural network size is also independent of input sequence length.
The successful removal of inductive bias is really what differentiates this from previous sequence-to-sequence neural networks.