An instructive thing here is that a lot of stuff has not improved since ~2004 or so, and working around those things that have not improved (memory latency from ram all the way down to l1 cache really) requires fine control of memory layout and minimizing cache pollution, which is difficult to do with all of our popular garbage collected languages, even harder with languages that don't offer memory layout controls, and jits and interpreters add further difficulty.
To get the most out of modern hardware you need to:
* minimize memory usage/hopping to fully leverage the CPU caches
* control data layout in memory to leverage the good throughput you can get when you access data sequentially
* be able to fully utilize multiple cores without too much overhead and with minimal risk of error
For programs to run faster on new hardware, you need to be able to do at least some of those things.