For inference, at least locally, the bottleneck is usually the memory bandwidth (and quantity, of course).
I hope that AI hype lead us to more memory and more memory bandwidth, because they are really lagging behind computer power increase from like 15 years already.