I'm not conviced. Maybe it really is that the wider type is faster, and int_fast16_t should be 64 bits. I just tried it out on a 64 bit system. The program is just a for loop.
If the loop variable is a int_fast16_t the loop increment step compiles to
addq $1, -16(%rbp)
If it is int16_t it compiles to movzwl -4(%rbp), %eax
addl $1, %eax
movw %ax, -4(%rbp)
The 64 bit width version is shorter, and so conceivably faster. Performance seems the same though, but that's probably because this program isn't a proper benchmark.