A few cores too many (2016)
prl.khoury.northeastern.edu
prl.khoury.northeastern.edu
Then I found out about pinning threads and processes to a cpu - which made everything go lots faster.
It is pretty annoying to benchmark nowadays unless you have access to BIOS, or at least root.
Now CPUs are still stupid fast, so you can use the poor abstractions, but if you really want to make a high core big CPU cache system work, you'll need some hardware specific stuff.
ScyllaDB vs Cassandra was a big example of this, and scylla's performance leap forced a pretty big rewrite of the java-based cassandra core from a streaming event architecture (SEDA) which would thrash the cores/caches for events, to a pinned CPU architecture.
Even though we are about, what, 15+ years into desktops having 4+ cores, and ludicrous amounts on the servers, there really isn't very good documentation, tutorials, and examples of squeezing maximum performance from these multicore systems. Platforms like Java have 20 years of multithreading code and libraries and examples, and a huge amount of them are really bad practice these days.
The big explosion of Python and Javascript aren't really helping, those languages are saddled with GIL or single process + wait architectures. That isn't going to wring the max out of a huge multicore system that pinned processes would benefit from. But then again does Java, while it is multiprocess/multithreaded, have the ability to pin a process to a CPU?
I opted out of that lawsuit in writing stating my opinion that a judgement against AMD in this case would have a chilling consequence on future architectural developments. AMD eventually settled for $12.1m.
IIRC they shared the floating point unit, not the integer unit. (A long time ago I had one with 8 "cores" but only 4 FPU.)
Ahhh 80386SX and 80387SX live again in spirit ...
I am working on a multithreaded barrier which does mass synchronization on a schedule, a rhythm.
It can send ~169 million messages a second across 10 barrier threads and 63 million event ingests from 3 external threads.
Your algorithm has to be redesigned to support this style of programming. Even message passing can be slow due to context switches. But bulk buffer processing is fast.
It’s better to follow the scatter/gather model of MPI (don’t use that) in the same process and do as much amortisation in each thread as possible before collation.
When it comes to I/O (with spinning rust), a measured rule is to use twice the cores to account for seek time and use big writes/reads rather than small ones. This can saturate the disks so do measure it.
I spent about ten years worrying about these kind of issues on enterprise storage in C++.
This avoids contention in user space. It also reduces fragmentation. You can also bound the memory usage by blocking until memory is free.
If memory serves, boost C++ has some code to help there though I did it myself.
On their system two instances work fine because they have two CPUs so two L3 caches.
Their Opterons actually have an unusual cache hierarchy by modern standards - L2 and L1i caches are also shared, in every pair of cores. L3 is shared between all cores as usual.