It's just vectorizing calculations so that it's faster than pure iterative calculation. 64 bit version doesn't get this probably because that optimizer isn't SSE2 aware yet (just a guess I actually don't know) and can't do SIMD arithmetic with 2 64 bit floats
Harder to vectorize 64-bit arithmetic?
was going to say this. It probably only does SSE and not SSE2. Therefore vectorization only happens for 32 bit ints.