Laptops have had unified memory for a ~decade as a result, but once you go to a discreet GPU instead of integrated then there's no feasible way to share the IOMMU & memory bus, at which point it's not practical. And then up-gradable graphics or just 400W GPUs like the 3090/4090 behemoths are incompatible without extreme cost.
Sure, but to get the same memory bandwidth as a M1 Max you'd need 8-channel DDR5 memory, so your socket would look like Epyc and consume 50W+ just for data movement.
GDDR allows moving much more data than regular DDR (or even LPDDR) but both GDDR and LPDDR5X require much tighter signaling requirements that preclude socketed memory. And that means much higher watts-per-byte-transferred for socketed memory, in practice. Like dozens and dozens of watts higher for a typical configuration.
> Laptops have had unified memory for a ~decade as a result
There's unified memory, and there's unified memory.
Fusion HSA and Intel iGPUs have always required you to designate memory as either graphics or CPU side so that it knows which cache controller to run it through, and this has implications for memory and cache visibility.
ctrl-f 'garlic' and 'onion': https://www.eurogamer.net/digitalfoundry-how-the-crew-was-po...
https://forum.beyond3d.com/threads/amd-kaveri-apu-features-t...
Money shot: "One issue we had was that we had some of our shaders allocated in Garlic but the constant writing code actually had to read something from the shaders to understand what it was meant to be writing - and because that was in Garlic memory, that was a very slow read because it's not going through the CPU caches. That was one issue we had to sort out early on, making sure that everything is split into the correct memory regions otherwise that can really slow you down."
In contrast, the newer stuff like PS5/Xbox Series family and Apple Silicon family has true zero-copy where it's all just flat memory and all writes are immediately visible to all clients of the memory controller. You can interleave GPU and CPU and Neural Engine tasks without forcing explicit sync into memory to flush the cache/bus. I think most people would say that's meaningfully different.
Both are "unified" but one is more unified than the other.
You can’t have the same memory with the same latency used for both CPUs and GPUs because GPUs don’t have DIMM and CPUs do, for example.
You also can’t guarantee a fast enough SSD to be able to delay loading assets until they are needed.
Put succinctly: even if this hardware advantage were available on PC, game devs wouldn’t optimize for it.
any PC would be an example