That said, there's really no way to correctly pick the "perfect" integer size for every platform - There's just too many variables. IMO I much prefer to just use 'int' when I know the values are going to be small - smaller then a 16-bit value - and I don't care about the actual width. But obviously it's a hot topic for some - As long as you follow the standard in regards to it's size then it doesn't really matter what integer value you choose to use. I prefer saving the fixed-width integer types for situations when I know I need such a size.
If the loop variable is a int_fast16_t the loop increment step compiles to
addq $1, -16(%rbp)
If it is int16_t it compiles to movzwl -4(%rbp), %eax
addl $1, %eax
movw %ax, -4(%rbp)
The 64 bit width version is shorter, and so conceivably faster. Performance seems the same though, but that's probably because this program isn't a proper benchmark.Without profiling there's no way to tell on modern processors.
Comparatively, while one is three instructions and the other is one, that single instruction is still doing everything the other three are doing. The only reason they aren't both one instruction seems to be that `addw` must not being a thing (Why, I don't know). Since the I/O done by both will be nearly identical, and the addition's should be very comparable in speed, it's not crazy to say they should perform basically the same. If you compiled with -O2 I bet you'd see almost identical code, since it would probably remove the memory access.