Memory move is the bottleneck these days, thus the expensive HBM, Nvidia's design is also memory-optimized since it's the true bottleneck chip wise and system wise.
It’s expensive, not needed for most consumer workloads and ironically, is actually often worse for latency for many patterns, even though it’s much higher bandwidth
why is that ironic? Many ways to increase bandwidth come at the cost of latency...