I haven't read these papers (thanks for the links!) but there has been one major recent improvement. Starting with Broadwell (and I presume continuing with Skylake), the CPU can now handle two page misses in parallel: http://www.anandtech.com/show/8355/intel-broadwell-architect...
The other interesting thing is that page walks themselves are actually not very expensive: something on the order of 10 cycles. They only become painfully expensive when the page table is too large to fit in cache, and spills into memory, necessitating a load from memory just to get the page table. So improvements in memory (and cache) latency will have strong positive effect.