Well, if programs are created to process simple instructions on lots of data there are huge speedups to be had on current x86 processors. By doing this memory latency is no longer a bottleneck and SIMD instructions can be used much more often. If unnecessary heap allocations have already been taken out, structuring a program like this can result in very substantial speedups.