I've heard Intel's compiler makes optimisations that best favour the specifics of Intel CPUs. All other compilers are likely to make CPU-neutral optimisations. Do you think the results would be much different if run on an AMD chip?
Now I'm not sure how much this'd actually affect. The autovectorization in Intel's compiler is weak and, at least on Win64, sse2 is allowed in normal code without CPU dispatching. It does affect library functions like math and memcpy, but that'll only matter if your program spends a ton of time in them.