[1] The x86-64 ABI actually requires SSE2 extensions to work correctly.
[2] SSE2 added double-precision and vectorized integer support to the SSE registers, the former allowing you to replace x87 FPU usage for floating point (unless you need long double, which is extremely rare). The newer SSE sets generally add only specialized operations that are unlikely to show up in autovectorization anyways, and the wider instruction set of AVX is less useful for performance in the "it might be useful" autovectorization scenario. They are useful for specific known hot regions of code, and in distributing binaries, variants of these are constructed for different levels of feature sets and are dynamically selected based on the user's actual hardware.
You'd tell the compiler "target a machine with <basic instruction set>", so it's not allowed to use advanced features like FMA because it can't assume they're supported. With this library, you'd check at run-time if FMA is supported, and then change functions accordingly.
https://gcc.gnu.org/wiki/FunctionMultiVersioning
https://lwn.net/Articles/691932/
https://llvm.org/devmtg/2014-10/Slides/Christopher-Function%...
> Just in case anyone is tempted, please do not write code that assumes the host on which it is compiled will always be the host on which it will be executed.
Scroll down to get to the table that includes Physical Processor, Intel AVX, Intel AVX2, etc.
T2 instances do run on a number of different processor models, therefore it is listed as "Intel Xeon Family." M4 instances run on Intel Xeon E5-2676 v3 (Haswell) or Intel Xeon E5-2686 v4 (Broadwell), though m4.10xlarge only runs on Haswell and m4.16xlarge only runs on Broadwell. Pay close attention to the * in the table for those instances that run on either Haswell or Broadwell.
Generally you will find that the recent generations of EC2 instances for C and R have identical CPUs within a generation.