A single CPU core actually can execute more than one instruction at a time, by leveraging out of order execution. The trick for leveraging out of order execution is avoiding having data dependencies locally. By swapping the iteration order, they allow the CPU core to continue with the next iteration before the previous has finished. Why? Because there is no data dependency anymore!
I haven't profiled that code, but I guess that now the bottleneck would be the sum. But it doesn't matter, as accesing a register is the fastest operation. Accesing the memory cache is slower, and accessing RAM is even slower.
Algorithms and methodologies don't require being well-known or being well-maintained to be valid and useful comparisons.
[1] https://dgraph.io/blog/refs/bp_wrapper.pdf
[2] https://highscalability.com/design-of-a-modern-cache/
[3] https://highscalability.com/design-of-a-modern-cachepart-deu...