Shift and add are both one cycle on all modern CPUs. I think this might be different on the Pentium 4, but we x86 programmers pretend that the Pentium 4 didn't exist.
With regard to the other bitwise operations (and, or, xor), I don't know any CPU in existence where these are slower than add.*
*Except for the Pentium 4, where the add/sub unit is double-pumped and nothing else on the chip is, but that's a bizarre exception.
It's common for variable-length shifts to be slow on modern CPUs once you move outside of the x86 space. On the PS3's PPU and the Xbox 360's PowerPC cores they are very slow and to be avoided.
> I think this might be different on the Pentium 4
What was particularly slow on the P4 were variable-length shifts and rotates. With an immediate operand they still had 1 cycle throughput. That sounds good until you realize that the P4 had a double-pumped ALU design (as you also mentioned) and it had two ALUs. Whereas shifts and rotates could only execute on the general integer unit.
They're hardly modern CPUs -- even outside the x86 space, slow variable-shifts are pretty rare. They're horrifically crippled, with the PS3's PPU not being intended to be used for any serious processing (supplemented by the SPEs), and the Xbox 360's CPU just being plain bad. Fast variable-shifts are critical for bitstream reading -- the end result of this crippling is that despite having 3 highly-clocked cores, and being supplemented by GPU processing, the Xbox 360 still can't play Blu-ray video.
They're annoying to program but why aren't they modern? You're just redefining terms to try to win an argument. They were intentionally scraped down to lower power consumption and manufacturing costs. That doesn't make them not modern.
I haven't done any Itanium assembly since McKinley (Itanium 2), but their shifts had 3 extra cycles of latency. Making a fast shifter has only become a harder problem since then, because the ratio of wire:gate delay has grown; they could have a single cycle shifter in the current implementation, but it would cost them even more than it did back then.
x86 is unique among modern architectures because it has so much legacy baggage. Shifts are only fast on x86 because when Intel implanted slow shifts on the P4, it killed performance in a lot of apps that were compiled with compilers that were accustomed to having fast shifts, so that they would substitute shifts for multiplies whenever possible.
That is not true whatsoever.
http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc....
On the classic ARM processors like ARM9TDMI, I think post-ALU immediate shifts were free and register shifts were 1 extra cycle.
A shifted operand is frequently required to be available as input one cycle earlier than a non-shifted one, simply because the shift happens in an earlier pipeline stage than the main ALU ops, which is presumably what the figure in the table is meant to reflect. If the operand is ready, there is no additional delay. In ARM9 documentation, this behaviour is frequently referred to as the instruction having one or more "early" operands.
I verified this just now on an actual A9 using the cycle counter. A sequence of independent adds with shifted inputs executes at two instructions per cycle. If each add is made to depend on the previous, two cycles per instruction are needed as suggested by the table. Short chains of dependent instructions are executed out of order masking the added latency.
No ARM processor has ever had a "post-ALU" shift.
> No ARM processor has ever had a "post-ALU" shift.
Typo. I meant pre-ALU.