Yep. This is a huge distinction for developers. Hodgepodge ("NUMA") systems can easily have the problem of "GPU 1 can quickly talk to GPU 2 but slowly to GPU 3", which is a productivity killer on code that is already > 10X harder to write to beginwith. GPU teams like ours (Graphistry) make simplifying assumptions on the hardware based on where things are going... and uniform memory access (within a cache hierarchy level) is one of them. In this case, we assumed saturating a single GPU is enough, and write kernels such that as single-node multi-GPU boxes become mainstream (like this), that it's in reach to use them. For initiatives like the GPU dataframe devs in the GOAI project, same thing. Today's office conversation: What will this look like for cloud multinode (e.g., rack)? :)
AND... if you like frontend js or fullstack js, and think this stuff is cool, we're actively hiring :) See https://www.graphistry.com/blog/js-gpus-ml-arrow-goai for our thoughts here.