LLM inference is mainly memory bandwidth constrained so I think it's highly likely that a company will create silicon with just an insane number of memory chips and less compute. These ASICs will probably do the same thing the crypto ASICs did.
If we look back 1 decade, no one uses a GTX 950 for anything.