What can we conclude from all this? The software world is complex enough that anything you can say will almost undoubtedly have counter-examples. While we haven’t proved it here, one of your take-aways should be the following:
It is almost always better to vectorize than not to vectorize on Intel SIMD capable hardware.
If you are vectorizing (“LOOP WAS VECTORIZED”), then it is almost always better (faster) to vectorize with SoA “stride-1” data layout than the AoS data layout simply because your chances of the compiler doing something wonderful for you are much better. This is because there are fewer instructions needed (less work) in the SoA case. It is possible that SoA may not be dramatically faster than AoS for a given code depending on how a program exercises the memory subsystem, but it is unlikely that you will see many examples where AoS code would execute significantly faster than SoA code. So when in doubt or unless you can prove otherwise, prefer SoA to AoS, specially on Intel Xeon Phi. This tends to go against how many of us were trained to arrange data structures, but for quadruple-speed single precision execution and double-speed double precision execution over the alternative, you will want to consider it.
tl;dr:
So when in doubt or unless you can prove otherwise, prefer SoA to AoS
FYI Xeon Phi supports AVX512, i.e. 512 bit SIMD registers (8 doubles or 16 floats). In the consumer CPU world there is AVX2 which has 256 bit registers (4 doubles or 8 floats). AVX also has a 'fused multiply with add'(FMA) instruction which is useful for matrix multiplication amongst other things.