NUMA (Non-Uniform Memory Access): An Overview
queue.acm.org
queue.acm.org
The way I see it: You've gotta figure out your scale-out or load balance method. It might not hurt to just pretend each NUMA node is a separate server and just treat it like that.
NUMA is kind of out of fashion though. For everything it claims to do, there are better alternatives (MPI, map reduce, actors, ...).
If you architect your actor system so that all of your actors are on the same socket, all message passing (assuming messages are small enough) can happen in socket local cache. This will be much faster than having to cross socket memory or even worse going to main memory.
Conversely, if you aren't paying attention to memory locality you could see unexpected performance degradation as you add actors, because even though your system is more parallel there is more overhead in memory access.
You can get serious speed gains by reducing latency.
Similarly, writing cache-aware code (in terms of the cache hierarchy) and locking down thread affinity can have serious benefits as well, especially on OSs with rather poor schedulers like OS X and Windows, which tend to bounce threads around a bit too much.
Even then it only works if the app spends most of its time waiting for cache misses: it could reduce the runtime of a 60% stalled app by 20%, using the 100ns vs 150s latency figure in the article - assuming as the starting point that the OS heuristic memory placement fails pessimally. Less than 20% in more realistic cases.
For the vast majority of apps, there are more serious perf gains to be had for same or less effort elsewhere.
I've seen close to 250% speed increases with NUMA-aware code (and tying thread affinity down) - writing things so the code's aware of where the memory is and targets thread jobs to cores on sockets that have that memory already can significantly reduces memory traffic over the QPI links, which if you don't do (with AVX instructions trying to do 8 floats per clock with IB) systems tend to just be memory-starved, as they can't feed data to processors fast enough from main memory. Especially with Xeons, as the pre-fetchers are so aggressive.
I can believe you can see speed increases like you cite with bandwidth-bound code, but would doubt this kind of optimization is good bang for the buck for most perf critical code.
I'm talking highly-parallelisable code like in renders/raytracers, fluid simulations, image processing. All that stuff needs huge amounts of memory bandwidth. And it's generally parallisable in both threads and with SIMD with SSE/AVX, so you're literally up against the cache/memory throughput limits of the system.
Game engines aren't that parallel actually - even job queue based ones which are in theory very scalable have dependencies which means that generally there's a limit to the number of threads they can run at once. The above examples I gave can generally scale linearly, given memory bandwidth.
games was an example perf critical code that seldom benefits from numa optimization.
To say any thing for the "vast majority of apps" is a little out of bounds. In some cases on multi-socket machines a single cache miss takes as much time as ~1000 machine instructions. If you are working on calculation intensive operations that is a lot of overhead to pay for not having a good memory model.
I don't understand being confused/annoyed at the fact that he linked a video.