> Obviously compilers do use byte registers for byte operations because those are the only way to get them on x86, ...
OK, but what does it even mean that they avoid AL then, if they use them for 8-bit math? Under what other scenarios would they even want to use them?
FWIW, compilers could use 32-bit arithmetic ops almost all the time even when bytes are used in the source, because excect for right shifts and division, information flows only to more significant bits, so 32-bit bit ops give you the right results in the low bits.
In fact, I thought they did this more often – but when I looked into it recently they seem to prefer the low 8 bit registers.
> but they don't use the high bytes.
Not often, but of course the use case for the "second least significant byte" is very narrow in the first place [2]! If you just want to deal with byte values, you'll use the low bytes, after all.
The high 8-bit values are reserved for special cases, e.g., extracting that particular byte, and there compilers do use them (at least gcc, clang and icc on the first try, but not msvc):
https://godbolt.org/z/m5MbsM
The performance characteristics for the high bytes are quite different from the low bytes on modern, x86 too: see [1]!
> And neither Clang nor GCC emits inc or dec
Yes, if you don't set the march they will compile for an archaic blend of chips at least one of which may have the old inc/dec behavior.
Try setting a "modern" march flag like Haswell and fixing the issue where i wasn't actually used except as the loop induction variabale (allowing loop transformations away from simple +1) and both clang and gcc use inc:
https://godbolt.org/z/gtkd-v
---
[1] https://stackoverflow.com/a/45660140/149138
[2] It is possible one of the original motivations for the high bytes was to double the number of 8 byte registers in the 16 bit era. I.e., 8x byte registers was considered significantly better than 4x, and you couldn't use the low byte of di,si,sp and bp for other reasons.
In 64-bit, however, we now have 8 additional registers each with a low byte variant (r8b-r15b), so there are plenty of low byte registers for most work, so the remaining use of the high byte registers mostly seems to be taking advantage of their ability to quickly access the 2nd byte of a larger value.