Why doesn't NVIDIA also build something like Google TPU, a systolic array processor? Less programmable, but more throughput/power efficiency?
It seems there is a huge market for inference.
It seems there is a huge market for inference.
Less programmable, but more throughput/power efficiency?
I also wonder the same. It'd make sense to sell two categories of chips:Traditional GPUs like Blackwell that can do anything and have backwards compatibility.
Less programmable and more ASIC-like inference chips like Google's TPUs. Inference market is going to be multiple times bigger than training soon.