Latency numbers every programmer should know (2012)
gist.github.com
gist.github.com
All the timings are a little faster, the caches a little bigger, and the buses a little wider, but it's still basically the same stuff with different names.
P.S. What’s I dislike most about the article, it fails to explain why L3 cache is 10-20 times slower than L1 cache, while they both made from SRAM.
Why is it? Is it because L3 is usually shared?
The best explanation I saw is this: https://fgiesen.wordpress.com/2016/08/07/why-do-cpus-have-mu...
On the date of publication and rapid changes in tech world. It might be a bit older but the generic concepts are still great to learn if you only heard the terms before but don't know the details.
https://www.akkadia.org/drepper/cpumemory.pdf
Need to supplement this excellent article with Row Hammer information. It doesn't cover clflush. I sent the author a note.
http://www.bitmover.com/lmbench/
measures a bunch of these in a portable way. It tries to give you latency & bandwidth of all the things.
Better than guessing based on wrong or old information
There's "things programmers should know" according to HN: https://hn.algolia.com/?query=%22programmers%20should%20know...
And as previously mentioned, 97 Things every programmer should know: collective wisdom from the experts, Kevlin Henney, ed., O'Reilly, 2010
http://www.worldcat.org/title/97-things-every-programmer-sho...
Most of the latter is practices, with a few technical recommendations: DRY, floating-points, IPC, linkers, BTS. Most, however, aren't.
https://books.google.com/books/about/97_Things_Every_Program...
http://www.worldcat.org/title/97-things-every-programmer-sho...
Its more about programmers needing to know what L1 cache is, the idea that some operations are faster than others, etc.
I know a lot of web dev type people who have no idea how the CPU works, what a register is, what paging or virtual memory are, etc. When you're treating compute resources like they're free and abundant (what web devs like to do these days), then of course you don't care. I just wish those devs did care, because their fancy dev machines blind them to the fact that their theoretically simple web app groans on anything other than an i7 with 8Gb ram. Their fast internet and local servers also seem to make them forget why its bad that first load requires megabytes of js. Sometimes I'd rather browse with my cheap tablet and that nonsense seriously blows.
There's a lot of gluttony in development these days. I wish every developer was required to take a basic OS or assembly course to see just how much is happening between writing js and having it actually execute. To see what it really means to program a computer and not a web browser.
I also wish more devs would take a look at the performance of MS word and the performance of Google Docs and apply a little critical thinking. On my cheap Surface 3 (4gb ram i3) Word loads instantly, sips power, and does everything I'd ever want locally. Docs takes forever, destroys the battery and is slow as molasses in January with both Edge and Chrome.
Yes, every programmer does need to know these numbers and why the numbers are what they are.
I have personally encountered people that dismiss the whole "lets stuff everything in /usr" issue with claiming that everyone (or at least those they care about) are using lights out management anyways.
You may be a programmer in a language that doesn't permit you to influence L1 cache performance, for instance, but you better understand the mechanisms involved and how that applies to your language and computational model 5 layers up.
Given that any simple bit or int arithmetic might be 50x faster than accessing a bloated field in a bloated struct, they'll never be able to write performant software, yet understand why more code and more lines are faster.
I think programmers or system designers that miss things like new secondary storage technologies are doing it wrong because the orders of magnitude can really affect development time.
That knowledge alone can speed up queries orders of magnitude and prevent the need to move to a NoSQL solution.
Ubiquity of SSDs brings to a close a time when getting data from another computer is faster than getting it from your own disk drive. A ratio that had previously been true, IIRC, not since a time in the eighties (see the Sprite system, process migration).
Is this from the future?
The SATA III bottleneck is still common today, but Mac laptops, at least, have sported SSDs that can do ~1GByte/sec for a couple of revisions now. I think they're PCI Express too.