How Misaligning Data Can Increase Performance 12x (2013)
danluu.com
danluu.com
Optimization for performance has always been about cache lines (64 bytes), not pages (4 KiB). Of course you're going to get terrible results when you're wasting huge amounts of memory.
We can do OoO execution, why can't we do OoO prefetching w.r.t. page faults? (I.e. try to fetch into a cache, if there would be a page fault don't. If there's something that would cause that fetch to page fault between it being prefetched and it logically being fetched, invalidate the cace.)
OoO engine were able to issue address calculation from loads (if address register is ready) ahead and thusly when real load instruction execute the data is already in cache or closer to cache.
It is really easy to do, actually. It can even simplify page fault handling in CPU.
That being said if you're really looking to optimize for TLB misses you should be using huge pages if your OS and processor support them.
Is gate count really that tight in L1, that they can't throw a few XOR taps in front of the cache bus? Or is it simply to make cache collisions more predictable?
Relatedly, in the L3, there is a hash function used to distribute different address to different regions of the L3. The cost of doing this is less significant in two ways: the L3 access latency is already much much higher (as elaborated on above) and the hash calculation can be done in parallel with other required logic (e.g. in parallel with L2 access).