Yup, the company I work for has the world's fastest LLM appliance, based on a hardware design co-locating memory and compute. It requires compilation innovation as well in a sort of hardware-software codesign. We don't see transformer models as being an impediment in terms of memory requirements or speed in the near future.
[I would be happy to say more but I just got a top-level comment flag killed, presumably because they though I was advertising, so I won't mention the company name.]