Certain hypervisors have the ability to disable features on the virtual CPU to enable live migration between different generation physical CPUs, in which case a binary that depends on a disabled virtual CPU feature (a.g., AVX-512) will simply crash (or otherwise fail) when it executes an unsupported instruction.
Other than that, I'm drawing a blank.
Hypervisor performance will vary, but I can't envision any scenario where a binary optimized for the processor's architecture would perform worse than one without any optimizations when running on a VM vs bare metal.
Most compilers assume that emitting the code in certain modes (SSE/AVX etc.) have particular cost. That cost may drastically change depending on how the implementation of the hypervisors handles the registers in question.