You want memory to be close to where it's used because at the speeds of high-performance ICs, the latency caused by distance is actually significant.
Are you saying proximity here more than offsets this vs. e.g. each core having its own cache as I think they do in a "normal" CPU? And if so, is this more true of ML inference workloads than other workloads, for some reason?
https://www.realworldtech.com/includes/images/articles/snbep...
https://en.wikichip.org/wiki/intel/microarchitectures/sandy_...
e-cores do have a CCX/core cluster, but the clusters themselves go on the ringbus lol