Wow. This sounds a lot like the Resnet moment for Transformers.
The argument is that otherwise, the gradient magnitude for the lower layers becomes too big. Which intuitively makes sense, because due to the residual connections, all error signals from the upper layers will end up at the lower layers.
I wonder why this is apparently not a problem for ResNet.