The speed of transformers is memory bandwidth limited. I think that it's possible to speed them way up with different chip architectures, and I know there are lots of people working on this e.g. https://etched.ai making asics and https://untether.ai making chips will co-located memory and compute cells. So unless some architecture really starts beating transformers badly on language tasks, I think the speed problem is going to disappear as silicon architectures adapt.