Fast. But inc rcx is slow.
> Or the even shorter "mov cl, 1"?
Fast.
> I suspect this is a width restriction in the bypass/forwarding network...
I think we just found the explanation on Twitter: https://x.com/corsix/status/1874965887108976858
Alder Lake adds support for mov/add/sub with small immediates to the register renamer. So `add rcx, 1` gets handled by the renamer, potentially with zero latency.
Unfortunately shifts are slow when the shift count is renamed in this way. This leads to fun things like `mov rcx, 1024` being fast while `mov rcx, 1023` is slow. I'll update the blog post in a bit.