Gallery of Processor Cache Effects
igoro.com
igoro.com
Unfortunately, we also let our customers develop their own code. One of our best customers (Motorola) starting having problems with one of their projects. Small changes in their code would cause all kinds of bizarre crashes.
I was asked to help out and after reviewing their code, I realized that nobody had told them about the cache coherency issue. After making a few changes (just moving .bss and .data variables into cache line protected malloc's), all the problems disappeared.
The funny part was when their engineering manager found out. He calls me on the phone and starts to chew me out. He was furious, so I put him on speaker and called my managers into the lab. We stood around and just let him rant for a few minutes until he ran out of steam. A legendary episode that we chuckled about many times thereafter.
The first version of the chip was introduced at CES 2001. Pretty early for a multi-core CPU. Only 150 MHz clock rate though.
Can you elaborate? Specifically ow did you "pad all memory allocations to a cache line" and how exactly did that prevent interference?
>"After making a few changes (just moving .bss and .data variables into cache line protected malloc's), all the problems disappeared."
What is a "cache line protected malloc" exactly?
The failure mechanism is when one processor writes a memory variable, the entire cache line is written. If the variable is within 64 bytes of a variable on the other processor (and in the same cache line), the variable on the other processor gets (possibly) overwritten.
I forget exactly how the aligned malloc was coded. This was 18 years ago and gcc 2.7.2 on SPARC. For the application programmers, we just had an overloaded new operator that was aligned.
<Motorola Eng Manager ranting away in the background>
First LSI engineer: "He seems pretty upset."
Second LSI engineer: "Yeah."
First LSI engineer: "Think we should do anything?"
Second LSI engineer: "Meh, he'll get over it."
First LSI engineer: "Still, maybe we should throw him a bone."
Second LSI engineer: "<sigh> Fine. I'll add an overloaded new operator after I finish my beer."
... something like that? :)
They were just declaring static/global variables that they didn't know were forbidden.
2010: https://news.ycombinator.com/item?id=1094797
(Those are for people who like to look at old discussions. Reposts are ok after about a year on HN: https://news.ycombinator.com/newsfaq.html.)
The canonical reference for diving deeper is "What Every Programmer Should Know About Memory" [1].
[1] https://people.freebsd.org/~lstewart/articles/cpumemory.pdf
EDIT: Should have kept reading, other people apparently had this same comment and it was addressed. This was for C# rather than C++ so no wonder my intuition was off.
Things could be different by now, though, with RyuJIT and tiered compilation. That article is quite old (although processors still work the same, just with larger caches).
https://hn.algolia.com/?query=what%20every%20programmer%20sh...
I've always found it annoying that the cache offers drastic speedups for doing things 'wrong' in the sense an algorithm is faster in cache than the 'correct' way until it scales out of the cache.
Feel free to ignore 50x performance/battery gains at your own peril.
Is it? If you're excluding memory access then I would argue it's not a proper representation of the work performed. You can have an algorithm that's mathematically ideal but from an engineering perspective is the wrong choice.
Also your workload doesn't always define if you fit in/out of cache. Linear access algorithms will scale on any architecture/cache size as the pre-fetcher will step in and bridge your DRAM/network access time. It's essentially like an infinite cache.
Realtime Collision Detection[1] (which is basically datastructures for 3D space) does a fantastic job of picking algorithms that are both correct and cache friendly. Data Oriented Design, SoA/AoS and the like are all techniques that I think any Software Engineer worth their salt should be familiar with.
Regarding your first comment, see https://danluu.com/intel-cat/
They avoid my second concern by having lots of cache memory even though they technically need a small fraction of it to maintain performance. There is a trade off, so probably still lots of room for improvement if they can get finer cache control.
My fear in those scenarios is you have profiled your workload, your cluster is at 75%+ utilization, and something non standard hits and triggers a workload profile that removes any cache performance benefits, sinking your hole cluster.
Also this is about CPU caches so it doesn't have anything to do with cluster utilization. A cluster will see the CPU as being ultilized at 100% regardless of how many cache misses there are.