I was curious about this topic too because M3 macs have "Unified memory" which is shared amongst their CPU/GPUs. Anyone have a link or explanation of how this works?
(400 GB/s is a lot in the form factor but the 4090 and equivalent have 1 TB/s, and H100s several times that)
Edit: Here someone asked the same question: https://www.reddit.com/r/LocalLLaMA/comments/14319ra/rtx_409...
That is ambiguous, instead I should have said that models that fit in memory on both take 3-4 times as long on the M2 as they do on the 4090.