I believe there's still a throughput penalty for cacheline-crossing loads as well, though it's quite modest compared to pre-Nehalem (where each cacheline crossing cost you 20 cycles!)
I believe there's still a throughput penalty for cacheline-crossing loads as well, though it's quite modest compared to pre-Nehalem (where each cacheline crossing cost you 20 cycles!)
Looking now, Intel's guide says that Skylake reduces the cross page load penalty from 100 cycles to 5 cycles. I tried with and without crossing page boundaries, although this wasn't my intent. In one version I reloaded loaded 16KB of floats, and in the other I reloaded the same 4 floats. I saw minimal difference between these two on both Haswell and Skylake. But since the penalty is increased latency, it's possible that the "load and throw away the result" doesn't illustrate this issue.
I'm thinking of the virtual address translation costs having impact on the run times of common algorithms, e.g., as demonstrated in the following work by Jurkiewicz & Mehlhorn: http://arxiv.org/abs/1212.0703
The recent research I'm aware of is, e.g., Generalized Large-page Utilization Enhancements (GLUE) mechanism, proposed in "Large Pages and Lightweight Memory Management in Virtualized Environments" (from this year's Micro): slides: https://dl.dropboxusercontent.com/u/36554102/BPC-1.pdf ; paper: http://paul.rutgers.edu/~binhpham/phamMICRO15.pdf
Admittedly, it focuses specifically on one aspect (the Double Address Translation on Virtual Machines issue in the Jurkiewicz & Mehlhorn context).
What I'm wondering about is: Has there been any progress on that on the "practical implementation" side, in the recent/coming Intel (or other, for that matter) CPUs?
The other interesting thing is that page walks themselves are actually not very expensive: something on the order of 10 cycles. They only become painfully expensive when the page table is too large to fit in cache, and spills into memory, necessitating a load from memory just to get the page table. So improvements in memory (and cache) latency will have strong positive effect.
Interesting about parallel misses handling, thanks!
One worry is that this tends to compound other effects -- say, non-prefetch-friendly access combined with TLB misses resulting in increasingly expensive slowdowns (as in the continuous-vs.-random array access example in the paper).