These days, it’s not only compilers that are optimizing for the particulars of the CPU, but CPUs optimizing for compiled code. My favorite examples of this are the inconsistent and sometimes downright crappy performance of “rep” string instructions and the “loop” instruction. Both seem like ideal easy things to write for handwritten assembly, yet the performance of both constructs can be quite awful on certain modern CPUs compared to the naive loops that compilers output. Much of this can be blamed on the fact that compilers rarely use either construct, so chipmakers had no reason to make either of them efficient (and, indeed, seem to have actually pessimized them on some chips!).