I hope the LLM wave will leave GPUs behind to go back to pursue more general-purpose computation rather than spending their die area on multiplying 4-bit-number matrices and such things.
Is that not literally the exect opposite of the direction asics for LLM inference is going?
My point is, that if companies develop ASICs for LLM work, then GPUs will stop being the go-to computation device for these workloads, and that will mean, hopefully, that their architectures will stop being warped so as to cater to LLM work.