This is exactly the case.
Just a few minutes ago I was implementing a parse tree data structure my thesis advisor commonly uses. It encodes the number of types of children of each node as a pair of bits, ordered from least-significant to most-significant. If the nth child is a leaf node, it's encoded as `10`, and if it's an internal node, it's `11`. I implement the "number of children" function as a loop, shifting this pattern right by 2 until the value is zero.
Naturally, x86 has a useful opcode for this: BSR, bit-scan right. It tells you the index of the first set (1) bit scanning from most-significant to least-significant. It returns the index from 0=LSB, 31=MSB. In theory, this takes my O(n) solution and makes it O(1). I just BSR the pattern, shift right 1, add 1. I know it's a micro-optimization anyway, but I want to have some fun on labor day, so I give it a shot.
In practice, since the number of children is always 8 or fewer (this is for Ruby's grammar), the loop is always very short. Even while running 10,000,000 iterations to get the total time up to a few seconds, and even while using `rand` to try to trick the branch predictor (to make the looping case less predictable), I couldn't get a discernable difference to show up between O(1) and O(n). The O(n) case was definitely executing at least 3-4 times as many instructions in the test I was running, so I can only conclude BSR is deoptimized.
Edit: DarkShikari says that the deoptimization may be because I'm on an Athlon 64 processor and that this is not the case on intel processors.