And the computation is much less trivial.
Still, an interesting and important observation - page sizes change every decade or two, and almost exclusively upwards, so this is worth pursuing if you need to squeeze more speed from your heap.
And the computation is much less trivial.
Still, an interesting and important observation - page sizes change every decade or two, and almost exclusively upwards, so this is worth pursuing if you need to squeeze more speed from your heap.
The last time I had to optimize a b-tree based k/v store getting this right literally could boost performance in test cases by 40%.
Everyone needs to "squeeze more speed from their heap". CPU-bound problems are quite rare in modern software - despite the legions of software engineers who will stare at a CPU profile that says 100% utilisation, and declare "it's CPU bound, nothing to be done".
(this is a pet peeve of mine, in case you couldn't tell)
I think the point of the article is that
> memory latency
Is actually a hierarchy of its own, and non-linear depending on access patterns.
Instruction fetch latency, branch mis-predict latency, cache hit/miss latency (L1 through L3), cache pressure, register pressure, TLB hit/miss, VM pressure, etc. all become significant at scale.
Just knowing big-O/big-Ω characteristics aren't enough.
I’m interested to know what you think that CPU is fully engaged doing if it’s not making progress toward a solution?