but isn't that aspect common among this, CPUs, GPUs, ...? And it feels like the whole NOC thing would add quite some overhead to moving things around.
Are you saying proximity here more than offsets this vs. e.g. each core having its own cache as I think they do in a "normal" CPU? And if so, is this more true of ML inference workloads than other workloads, for some reason?