What happens when you shift a register by more than the register size?
devblogs.microsoft.com
devblogs.microsoft.com
How the CPU would handle the theoretical assembly instruction is usually of little importance when demons are flying out of your nose. Here's Regehr on some of the standard ways to protect yourself from a malicious or overzealous compiler when you need to have a function that does variable sized rotations: https://blog.regehr.org/archives/1063
It depends. In some cases, really, no they do not. Compilers have been caught deleting security checks when they noticed that the only way to reach the error case was to trigger UB… and since "UB doesn't exist", the error case is considered dead, and the whole test is deleted. Sometimes this leads to remote code execution vulnerabilities, that if exploited could actually result in your hard drive being encrypted, or a keylogger being installed.
UB is really, really scary.
And another commonly mentioned one: https://godbolt.org/z/6bf17W1Ee
Have you got a link to a CVE of an RCE that was caused by UB being exploited by an optimizer?
Im with you that UB is dangerous, but let's be honest in 98% of situations it's correct to assume that integer overflow can't happen, and that if you pass a pointer into a function and don't check it that the call site has checked it.
If you take the time to actually read Annex J of the C language you'll see that most of what is intended by "UB" is just "The language or platform could be buggy or the spec could be weird".
For example, a compiler can assume that a variable sized shift is in range, then use value range propagation to eliminate tests, which can be very confusing if you expect the shift to be implicitly & 63.
You don't need to rely on the trust and common sense of anyone in those cases, because oversized shifts aren't undefined behavior in Javascript or Rust. If C (or C++, or any other language) wants the same benefit, they can just say that it's not undefined behavior, and thereby require implementations to define the behavior somehow. And "backwards compatibility" isn't an excuse here: UB means that, currently, anything can happen, so having the standard specify that one particular thing happens is an entirely backwards-compatible change.
They used to, but unfortunately increasingly not so much now.
(Since the most popular ones are open-source, we should theoretically have the power to change that.)
I have to disagree - I do a lot of hacky low-level stuff, and C/C++ compilers love to "break" my code. I mean they have a right to, because I know it's technically UB, but I end up playing a game with the compiler where I "trick" it so it doesn't notice the UB and delete my whole function silently.
Actually, undefined means "undefined by the specification". It doesn't matter what the hardware does. The specification allows the hardware to do whatever it wants. It's perfectly valid to wipe your hard drive if you trigger undefined behavior.
A conforming compiler is allowed to completely remove undefined behavior like this:
if((n << 64) == 0) {
println("Hi!\n");
}
if((n << 64) != 0) {
println("Hi!\n");
}
If it were implementation-defined, it would have to compile it to the equivalent of println("Hi!\n");Separately, I don't think their point was that they were speaking as a layman, but rather that the loose definition they gave is historically accurate, easier to reason about, and what a working programmer not versed in the nuances of UB would expect. The fact that UB has its current definition is indicative of "undefined [getting] lost somewhere."
[0] https://kristerw.blogspot.com/2017/09/why-undefined-behavior...
How useful it is from a programming perspective is moot because this is the only possible definition if you want
a) A language that is powerful enough it can violate the invariants the compiler establishes and upholds in the runtime environment and then (necessarily) assumes are being upheld, and
b) any meaningful optimizations at all.
What's not reasonable in C is the extent to which stuff lazily gets shoved under the UB umbrella, but the concept itself is not avoidable if you give the programmer enough leeway to conjure up arbitrary pointers out of thin air (but even much less would necessitate UB).
What's vague to the extent of being totally useless from a specification perspective is the idea that you can "reason about" something that is undefined (and undefinable).
Sure, historically, the original authors had some vague concept of a "portable assembler" in mind. That was a mistake that got rectified since people most definitely wanted b). We have a distinction between "implementation-defined" and "undefined."
No, that's called implementation defined, isn't it?
This is what ‘undefined behavior’ turned out to be, so it's clear to me that the consequences of the definition were not understood at the time.
Foundationally C is unsound. For its purpose this made real sense - who cares about soundness, we need to ship Unix ASAP? But with unsound foundations nothing built upon them can be sound itself. ANSI and eventually ISO C finds a way to write the unsound language down formally but doing so doesn't fix it, Mother Nature isn't fooled by words.
In places where WG14 declined to write formal language excusing the unsound nonsense (such as provenance rules), there's still unsound nonsense, it's just not written in the ISO document. Again, Mother Nature doesn't care, programs with either kind of unsound nonsense malfunction, the fact that one kind is written down on paper and the other isn't makes no practical difference.
While that part of the specification predates MMX, they may very well have had the possibility in mind that shifts might be accomplished by multiple instructions with different behavior (if we're gonna give them this much credit). But one way or the other, it's useful now for auto-vectorization.
Also, a lot of it could easily be promoted to implementation dependent (this is what you are describing) at the 90's without any problem.
But, of course, the modern interpretation is complete bullshit.
Even worse, undefined behavior doesn’t just mean “what happens next isn’t defined by the specification”, it is defined, and it’s “anything the compiler wants to happen or that might make the compiler’s life simpler or emit more efficient code”. It's not that there's no rules about what might be emitted, it's that the standard does explicitly say what happens next. And what is specified is an unlimited pass to reorder or transform any input source code that performs this undefined operation, since it was nonsensical anyway.
The “undefined” part is the thing that you did, it's the input not the output. The compiler’s output is defined, and it’s “the compiler can do whatever it wants". And that's actually probably the only thing that could make the situation worse!
It can return a NOP;RETURN; and continue executing code instead of throwing an exception, or generate any other output for this function that makes its life easier, without let or hindrance by ordering or visibility, or by any relation to the input source code, in a reign of terror that makes a smashing blog post. It is valid to emit fdisk because you aliased a variable in an unreasonable portion of the code. It is valid to return 1 or null or anything else immediately and without notice, regardless of how this affects your expected invariants about data safety or code behavior. Etc.
this is, of course, a wildly terrible and terrifying decision to make and Ritchie was absolutely right about that. It's mostly only because compilers have stayed somewhat sane about the code they do emit that this hasn't all come crashing down, but, they get smarter every year.
Because an aggressive compiler will probably do constant optimization.
And the software guy who optimizes 1<<32 might have a different notion of what to do than guy designing the hardware would do.
For example, the SUNW compiler sometime around the year 2000. As me how I know.
For example, bsr/bsf instructions (count leading or trailing bits) have undefined behavior if the argument is zero. What I observed is that the CPU leaves the destination register unchanged in this case (but it can do something else according to the spec, for example - set it to zero).
The compiler exposes these instructions for convenience as intrinsics __builtin_clz/__builtin_ctz, which consequently, have undefined behavior if the argument is zero. But you can use them directly as an optimization if you know in advance that the argument is non-zero: https://github.com/ClickHouse/ClickHouse/blob/c0a43df749c827...
What is unusual is that the 32-bit flavours of these instructions leave the upper 32 bits unchanged, whereas other 32-bit integer instructions write zeroes there even if the lower bits in the destination are unchanged. Perhaps Intel intends to make these instructions zero-extend in the future ... or does there exist some processor (from Intel or someone else) that zero-extends?
The article claims UB only at greater but not at equal:
> Bonus chatter: The wide variety of behavior when shifting by more than the register size is one of the reasons why the C and C++ languages leave undefined what happens when you shift by more than the bit width of the shifted type.
If think your "equal to or greater" is correct, and the article is wrong. The reason might be that the instructions are defined up to including equal for many hardware.
I remember checking one case that included equal as valid for only one of left/right shift and invalid for the other (left/right encoded in 5 bit signed).
So some horrible bug involving shifting registers ate a portion of his life, and we get a cool blog post.
uint32_t foo(uint32_t x, int s) {
return x << (s % 32);
}
foo(unsigned int, int):
shlx eax, edi, esi
ret foo(0x12345678, -1)
? I think that, in modern C -1 % 32 == -1
(% computes the remainder, not the modulo) it would try to compute x << -1
and AFAIK that’s undefined behavior.Even if it is well-defined, it possibly is not what you want.
If you were to call
foo(0x12345678, 32)
I think the most logical result would be zero, but the code you give returns 0x12345678.x86-64's 'bzhi' instruction (does 'a & ((1<<b) - 1)') uses mod 256 (and thus, if mod 32/64 was specified in C, compilers couldn't optimize to this).
ARM NEON's SIMD shifts merge both left and right shift in one instruction (reads low 8 bits as a signed int, positive being left shift, and negative - right shift)
This reads a bit ambiguously, though the conclusion would be same.
If C specification had that requirement, yeah, they have to do mod 32/64 every time integers with unknown bounds are used for `b`. But then Intel should have noticed this first and define LZHI to use mod 32/64 according to the operand, or make two versions of LZHI.
If C code had that requirement, for example, by always using `a & ((1 << (b & 31)) - 1)` to avoid an UB, compilers indeed would not be able to optimize it into a single BZHI. But they can use AND+BZHI, and I think BZHI has only a half of throughput than AND in almost all supported processors, so an excess AND probably can't do much harm.
Unsigned: should produce a zero
Signed: should sign extend the leftmost bit, so it will either be 0 or -1
Shifting left should result in a 0.
Machine-specific behavior starts at register size, not beyond!
That's why in C, the behavior is undefined.
Some processors also have ror and rol instructions that accomplish shifting in the shifted out bits. 8086 also had rotate with carry flag: rcr and rcl. Aids implementing sign-extended shift right and arbitrary precision math.
That's still the case today, no?
http://www.tigernt.com/onlineDoc/68000.pdf
I suppose the point is the behavior undefined in C/C++, so you have to look at the chip level doc to see...or test it yourself.
Wasn't this also the case on Pentium 4 (with different max value than 255), when they removed the barrel shifter?
On all Intel and AMD CPUs starting with 80186 and 80286 (which have been launched simultaneously) only the lowest 5 (or 6 for 64-bit operations) bits of the shift value are used. All the other bits are ignored.
It does not matter how the shift operation is implemented, with a barrel shifter or without it. The implementation influences only the execution time of the instruction, not its effects.
This difference in the behavior of the shift and rotate operations was used by many programs to identify whether they were executed on an 8086/8088 or on an 80186/80188 (or a later) CPU.
This was important, because 80186 had introduced some new useful instructions, and the standard means of detecting CPU features with the CPUID instruction has been introduced only more than a decade later, in Intel Pentium (1993).
The detection of the CPU was based on the fact that shifting left a 16-bit register by 32 will not change it on an 80186 or later CPU, but it will produce the result zero on an 8086/8088.