From what i see reading https://en.wikipedia.org/wiki/Multi-channel_memory_architect... different channels could, in theory, be used "autonomously of each other".
From what i see reading https://en.wikipedia.org/wiki/Multi-channel_memory_architect... different channels could, in theory, be used "autonomously of each other".
Memory channels are independent, however generally all cores use all channels. There are two common designs. Intel has 4 dies, 2 memory channels per core, so 2 channels are closer/lower latency than the other 6. AMD has multiple chiplets, but a single memory controller with 12 channels. So All cores have the same latency to all channels.
Generally Intel has lower latencies to 25% of the channels, but AMD has more throughput (bandwidth or random IOPs).
One thing that surprised me is that for maximum throughput you want at cache misses queued to the memory controller, at least twice the number of memory channels. These days missing in L1/L2/L3 is often approximately half the total memory latency. So on an Intel Xeon at least 16 misses (per socket), AMD at least 24 misses (per socket.).
So on Intel you could tune things (and the NUMA support helps) to prefer the local channels. Most OSs help, and C calls like numa_alloc_local() allows local control.
For memory intensive codes I have found the best scaling when there's 2 cores per channel. Of course most codes are pretty friendly.