Since when did the average developer care about how many sockets a mobo has...?
Surely you still have to carefully pin processes and reason about memory access patterns if you want maximum performance.
Since when did the average developer care about how many sockets a mobo has...?
Surely you still have to carefully pin processes and reason about memory access patterns if you want maximum performance.
Or am I out of date on NUMA systems?
Remember the Pentium D? Unfortunately, I used to own one.
This stuff rapidly starts to make my head spin. I have not studied interconnects and have never written any NUMA-aware software. I will just post this link (read the "Memory Latency" section):
https://www.anandtech.com/show/16529/amd-epyc-milan-review/4
As I understand it, the I/O die is partitioned into four quadrants. Each quadrant has two memory controllers and is attached to two compute dies. CPUs can access memory attached to the same quadrant with lower latency than going to another quadrant. This is a NUMA system that can be configured to appear as one logical NUMA node.
I believe their smaller parts with two or fewer compute dies will be UMA, but with the same non-uniform latency to L3.
>So is that 56 core Xeon that Intel was bragging about for a while there until the 64 core Epycs & Threadrippers embarrassed the hell out of it.
I believe the 64-core Epycs and Threadrippers came first. The 56-core Xeon was a purpose-built part for HPC, so it wasn't quite a marketing gimmick.
Eg simulation softwares often used in the industry (but the one I’ve on top of my head is Windows only.)
Anyway, the point the make is this: if you claim doubling performance, but only the selected few softwares as you observed would be optimized to take advantage of this extra performance, then this is mostly useless to the average consumer. So their point is made exactly with your observation in mind, that all your softwares is benefiting from it.
But actually their statement is obviously wrong for people in the business—this is still NUMA and your software should be NUMA aware to be really squeezing the last bit of performance. It just degrades more gracefully to non optimized code.
My understanding was the the dustbin was designed with one big processor because SMP/numa was a massive pain in the arse for the kernel devs at the time so it was easier to just drop it and not worry.