What every programmer should know about memory (2007) [pdf]
people.freebsd.org
people.freebsd.org
Discussion of it in slightly more modern context, worth looking at.
[0] https://gist.github.com/hellerbarde/2843375?permalink_commen...
I also like to follow Tobi Lütke. The guy is the founder/CEO of Shopify (with all which that might entail), and he still tweets amazingly interesting low level coder stuff.
I concur. There's something about concurrency and multithreading that ticks your brain in a weird way. It keeps me stimulated.
My favorite is setting up a scenario that would normally be considered as a disasterous concurrency failure. Like intentionally messing up a `volatile` state. You can't ever have those in prod, of course, but seeing them in a controlled, _expected_ way tickles my brain in a very perverse way.
You don't have to understand all of it but if you can't be bothered to get the gist of this maybe computers aren't for you?
Note that I'm implicitly making a distinction between people who write code and people who are explicitly developers.
Delving deep gives you fundamental knowledge of systems, sometimes there isn't time for that which in that case use the generic library that does stuff you never needed and ship the product. But in 10 years time when that needs to actually become performant code you need to know how to fix it or someone else will be paid to do it instead.
I think it's maybe nice to know how we interact with memory in software (memory allocation/deallocation, naybe how manual memory mgmt is done, something about garbage collection?), but that's about it.
¹This almost never happens.
float Q_rsqrt( float number )
{
long i;
float x2, y;
const float threehalfs = 1.5F;
x2 = number * 0.5F;
y = number;
i = * ( long * ) &y; // evil floating point bit level hacking
i = 0x5f3759df - ( i >> 1 ); // what the fuck?
y = * ( float * ) &i;
y = y * ( threehalfs - ( x2 * y * y ) ); // 1st iteration
// y = y * ( threehalfs - ( x2 * y * y ) ); // 2nd iteration, this can be removed
return y;
}Carmack (and id software games) are legendary in terms of their high-performance. And given the source code of these games... the source code filled with raw assembly language and other such tricks like I posted above... the source code is nearly impossible to read for the "typical" programmer.
----------
This is absolutely not a company (or engineer) who did things "simply". He did things in a extremely difficult, high-performance first engineering mindset.
I dare say that John Carmack's practice is closer to "premature optimization", built around abusing incredibly low-level tricks to improve the performance of their various video games... and strategically choosing parts of the video-game simulation that people would prefer high-performance over accuracy. (IIRC: the Fast Inverse Square Root algorithm above is only ~4 digits or 12-bits worth of accuracy).
This was a programmer (or at least, programming team) that at a high level, decided that 12ish-bits of accuracy (out of 23-maximum bits of precision) was enough for their purposes in the vast majority of their code. And then used that slight performance increase to have a smoother-feeling game than their competition.
----------
IIRC, John Carmack is a legendary programmer in terms of low-level 3d details. He is a good programmer, but his style is the complete opposite of what you're calling it.
https://github.com/id-Software/Quake-III-arena/blob/master/c...
2. This document very precisely describes the parallelization mechanism of DDR3. Read/writes to DDR3 are parallelized by channel, rank, and bank. DDR4 added bank-group as an additional layer of parallelism (above bank below rank).
3. An individual bank has performance characteristics: an individual bank is much faster when it can stay "open", thus only getting a CAS-latency, rather than a close+precharge+RAS+CAS latency cycle.
4. CPUs try to exploit this parallelism with various mechanisms. Software prefetching, hardware prefetchers, levels of cache, and so forth.
5. NUMA is more relevant today as EPYC systems are innately quad-NUMA. HugePages is also more relevant today as 96MB+ L3 per CCX have appeared on Ryzen/EPYC machines, meaning you need HugePages to access L3 without a TLB hit. TLB isn't really RAM, but its closely related to RAM and could be a bottleneck before you hit a RAM bottleneck, so its still important to learn if you're optimizing against RAM-bottlenecks.
It just appears to be an external PDF that FreeBSD committer Lawrence Stewart saved to their fileshare, written by someone else.
As example, the north bridge isn't as explicit today as the document makes it seem. A lot of functionality that used to be reserved for the north bridge is now being tucked away into the CPUs or motherboards themselves.
Still, worth reading.
I think there are better blog posts out there that explain the concepts. "Cache locality" "memory cost of context switching" are some terms to put in search engines. Maybe add "thread local storage" in there too.
The big takeaway from this paper though, at least when I first read it, and the concept has risen to prominence in the years since- backed up by benchmarks- that Big O isn't really the end of the argument when it comes to performance.
The way CPUs work is that its extremely fast to rip through contiguous datasets. Jumping around to random memory addresses is way slower. A corollary to that is that jumping between threads is very slow for the same reasons- loading and unloading random chunks of memory. There are a lot of details around the specifics of why though- different cache levels and their latencies, and what happens when there is a cache miss, etc..
This lead HFT type programs to favor linear arrays to store things, and pinning threads to CPU cores, amongst other things, but that's the gist of it.