CPU DRAM can't really be used for inference efficiently -- inference mostly wants memory bandwidth, not memory capacity, and GPU DRAM has >10x more bandwidth. The fabs can switch between them but you can't switch after the fact.
Those machines with GPUs still need RAM of their own, and they generally want large caches to avoid SSD penalties. You even see this spill out in the form of costs for KV cache in <1min, 5m, 1hr rates etc.
yeah there's some KV cache offload
The majority of inference actually does happen in cpu.
This is just wrong unless you're using some confusing definition. Notice that companies trying to do lots of inference aren't looking for CPUs, they're looking for GPUs/TPUs.
Ah well, the opposite actually. Your definition, while traditional, is wrong. It was right some time ago. But it's wrong now. At least according to Intel and Nvidia in their disclosures. Intel was literally talking about their sales to "companies that do lots of inference" when they said it... so...