Nvidia GPUs are still at their core reliant on the PC architecture,
Inferencing on Nvidia cores will soon be like encoding a h265 stream on CPU.
I expect custom built TPUs will have progressively more and more advanced hardware acceleration where legacy aspects of the CUDA architecture will eventually limit their innovation without architecture changes (pci-e, nvme bus, cpu interrupts, reliance on system ram for index tables, etc ..) which fill their moat and level the playing field for google/Amazon/eventually apple