Incorrect transformation: (llvm.maximum undef, %x) – undef
bugs.llvm.org
bugs.llvm.org
The bug report claims
> The domain for (llvm.maximum undef, %x) depends on %x,
But this claim is not true. Here is why, when the LLVM docs state:
> [llvm.maximum returns] the maximum of the two argument, propagating NAN [...]
they are referring to IEEE's 754 maximum semantics, which state that if one floating-point argument to the function is a NAN, the function returns a NAN.
However, IEEE assumes that floating point values are composed of bits, which can be 0 or 1.
That assumption does not hold in LLVM, where a bit can take any of three values: 0, 1, or undef (well there is also poison, but lets leave that aside).
According to IEEE, llvm.maximum(%x, NAN) must return NAN _if_ x is not undef. But if x is undef, then IEEE does not hold.
The LLVM docs do not say what this function should return in that case, so the behavior is IMO currently undefined, and any optimization for this is correct, so we might as well pick the one that enables most optimizations, and that's returning undef (e.g. these cases only legitimally happen in dead code, e.g., expanded from macros, and you want the compiler to better remove this dead code).
There is also the issue of consistency. If we were to replace llvm.maximum(%x, NAN) with llvm.select(%x >= NAN, %x, NAN) then you would think that any number compared with NAN compares to false, and this always returns NAN, but that does not hold for undef, where you get llvm.select(undef, undef, NAN) returning undef or just being plain UB.
So IMO the assumption motivating this fix is wrong. It assumes that undef is a floating-point value, but it isn't. It is something else, and the floating-point rules of NAN comparisons with undef do not apply to it. This fix will only remove optimizations, and delay broken code from standing out, because if your program tries to execute `llvm.maximum(undef, %x)`, your program is broken.
[0]: https://llvm.org/docs/LangRef.html#llvm-maximum-intrinsic
double x;
double y = std::nan("");
double z = std::fmax(x, y);
is undefined behaviour. So the current transformation in LLVM looks correct to my C++-coloured eyes. (Incidentally, std::fmax(x, nan) is supposed to return x, but even if it was defined to return nan then this would still be undefined behaviour, because passing an uninitialised variable to a function is undefined behaviour even if the parameter is ultimately ignored.)Could someone explain a practical situation where the listed transformation causes incorrect compilation of some code in a programming language (maybe not the right term but I mean as opposed to LLVM IR)?
An actually interesting question might be: which of the correct transformations is the more useful?
To get undefined behavior, you need to do something like use the unreachable instruction, or call a noreturn function that does in fact return. You may also have "poison values", which trigger undefined behavior on certain uses.
fmax in C/C++ specifies that the non-NaN argument is returned if one argument is NaN. The llvm.maxnum intrinsic does likewise, although it has more clarity on what it does with signaling versus non-signaling NaN. The llvm.maximum intrinsic specifies that NaN is returned if one argument is NaN.
Edit: Here's a longer discussion of undefined behavior in LLVM:
Undefined behavior in LLVM is more complicated than C/C++, and this is one of the areas where assuming C/C++ semantics are going to confuse you more than help. Unlike C/C++, there are several gradations of undefined behavior, and what is UB in C/C++ may be weaker in LLVM.
Strongest is undefined behavior. Like C/C++, the optimizer will optimize under the assumption that it does not occur, and can even draw inferences that it will not occur. Undefined behavior largely arises out of cases where you add annotations that are incorrect (e.g., a readnone function reads memory), but it is also retained in a few basic situations: dereferencing null pointer, branching on undef, divide by 0, unaligned memory accesses [1].
Somewhat weaker is the poison value. This is effectively delayed undefined behavior. Poison originates from violating nuw/nsw flags (basically, C/C++'s signed integer overflow rule), as well as out-of-bounds pointer arithmetic, and maybe a few more cases I don't remember off-hand. Poison propagates through most instructions, and becomes undefined behavior if you use it to control a branch or access memory.
Weaker still is the undef value. This value can take on any bit pattern, and--crucially and confusingly--it need not take on a consistent bit pattern. So in this code:
%x = undef
%y = sub i32 %x, %x
The value of %y is undef, because you can "read" different values of %x in the two arguments of %y. It is undefined behavior to dereference an undef pointer, or branch on an undef value.Weakest is the frozen values. The LLVM freeze instruction converts its input to an undefined, but stable value. If we were to consider the equivalent of the above program:
%x = freeze i32 undef
%y = sub i32 %x, %x
Then we can say that %y = 0 here, since the two values of %x must read the same value. Frozen values are still undefined behavior value to access memory, but they are no longer undefined behavior if used as a branch condition.[1] Note that, unlike C/C++, LLVM does allow underaligned accesses, so you can specify a load of a 4-byte value that is 1-byte aligned, and this is not undefined behavior.
And having answered that, when looking at that C++ code, would I expect it to be well defined even though this bug (if it is a bug) makes it undefined?
Here, LLVM's intrinsic names are following the IEEE 754 standard names. C and C++ only has the old definition provided in their standard library, although there is a proposal for C2x to add the new comparisons under the names "fmaximum" and "fminimum" (see http://www.open-std.org/jtc1/sc22/wg14/www/docs/n2489.pdf).
NaNs are one of the bigger parts of IEEE754 that make me wish posits could take over...
I personally think perhaps having two kinds of float/double, one for arithmetics, with full IEEE754-compliant semantics, and another for general-purpose code, where NaNs would be just normal values that you can compare with themselves and get "true", would make sense.
Also, we are seeing the first papers trying to do numerical analysis of posit computations and... it is not better than IEEE-754, just different. Probably a better default for 16 bits arithmetic but highly debatable otherwise.