I do think long term there gets to be more hope for CPUs here with inference largely because memory bandwidth becomes more important than the gpu. You can see this with reports of the MI-300 series outperforming h100, largely because it has more memory bandwidth. MCR dimms give you close to 2x the exiting memory bw in intel cpus, and when coupled with AMX you may be able to exceed v100 and might touch a100 performance levels.
HBM and the general GPU architecture gives it a huge memory advantage, especially with the chip to chip interface. Even adding HBM to a CPU, you are likely to find the CPU is unable to use the memory bw effectively unless it was specifically designed to use it. Then you'd still likely have limited performance with things like UPI being a really ugly bottleneck between CPUs.
MCR DIMM is like 1/2 the memory bandwidth that is possible with HBM4, plus it requires you to buy something like 2TB of memory. It might get there, but I'd keep my money on hbm and gpus.
Obviously you can’t change an asic
Obviously an ASIC is not a general purpose machine like a cpu.
In some cases ASICs are faster than general purpouse CPUs, but usually not.
You can make an ASIC which doesn’t have the same power draw as a CPU, but provides the same performance.
It doesn’t need to be faster than the fastest software implementation, but power per performance will always favor ASIC.
Edit: an example ASIC the Pi has is the video encoder/decoder, with JPEG also supported. I think it's embedded in the GPU.
"Modern ASICs often include entire microprocessors, memory blocks including ROM, RAM, EEPROM, flash memory and other large building blocks. Such an ASIC is often termed a SoC (system-on-chip)."
What you're describing in the JPEG codec might be termed a fixed function IP block in semiconductor design terminology.
An application-specific integrated circuit (ASIC /ˈeɪsɪk/) is an integrated circuit (IC) chip customized for a particular use, rather than intended for general-purpose use, such as a chip designed to run in a digital voice recorder or a high-efficiency video codec.
Colloquially, I’d never call anything with a programmable processor an ASIC
> For example, when I run my spam.sh shell script, it only takes 420 milliseconds, which is 7x faster than my Raspberry Pi 5. That's right, when it comes to small workloads, this chip is able to finish before CUDA even gets started.
So… it depends :)