I thought “unified memory” was just a marketing term for the memory being extremely close to the processor?
An RTX4090 or H100 has memory extremely close to the processor but I don't think you would call it unified memory.
A huge part of optimizing code for discrete GPUs is making sure that data is streamed into GPU memory before the GPU actually needs it, because pushing or pulling data over PCIe on-demand decimates performance.
If you’re forking out for H100’s you’ll usually be putting them on a bus with much higher throughput, 200GB/s or more.