1. To give compilers some leeway when optimizing stuff (for examples, see http://blog.llvm.org/2011/05/what-every-c-programmer-should-...)
2. To make C usable for those worried about erroneous overflows in their code.
3. To make it easier to write C compilers for CPUs that trap on integer overflow.
4. To allow for performant C compilers on CPUs that use one's complement arithmetic.
I'm curious about this. See below.
> 2. To make C usable for those worried about erroneous overflows in their code.
I think you're saying that the rule allows compilers to implement -fwrapv if they choose, which sounds reasonable.
> 3. To make it easier to write C compilers for CPUs that trap on integer overflow. > 4. To allow for performant C compilers on CPUs that use one's complement arithmetic.
I wonder how much these matter nowadays.
For optimization, the blog post you cite describes the following examples:
- "X+1 > X" to true
- "X*2/2" to "X"
- "<= loops" and "int" induction variables
On the first two, presumably these usually only come up after macro expansion and inlining, however they're still suspicious. If a function is scaling its return value and its caller is de-scaling it, it's usually a sign that the API isn't designed quite right. I'd be curious to know how often these come up.Code using "int" induction variables to step through arrays on 64-bit targets is often sloppy. Such code won't handle very large arrays properly, due to the limited range of "int", which is a bug that may not be quickly noticed.
And for the "<= loop" itself:
for (i = 0; i <= N; ++i) { ... }
it's really unlikely that the code is actually intended to be an infinite loop in the case where N happens to be INT_MAX. Code like this would usually be clearer written as "i < N + 1" to emphasize that it really does intend to iterate N+1 times rather than just N times, and it just so happens that this form makes optimizers happier as well.Instead of having compiler writers sit around and think up clever ways to repurpose anachronistic language rules, I might prefer to have them focus instead on ways they can help me write better code instead :-).
The other vision for C is a portable systems programming language. When trying to write portable code, undefined or implementation-defined behavior is a big problem. I'm glad that compiler writers employ a take-no-prisoners approach here. People shouldn't be relying on undefined behavior when trying to write portable code, so compiler writers shouldn't be bound to implement consistent behavior in those cases. That's especially the case if code ever needs to be compiled with a different compiler, which might decide to implement different consistent behavior.
It's also worth noting that in some cases, the exploitation of undefined behavior doesn't always happen in a single place in a compiler. Instead, it can be a combination of applying a few different rules in different optimization passes that produces surprising results.
With that being said, it sure would be nice if the compiler writers figured out how to give more warnings when undefined behavior is detected. If x + 1 > x is optimized to false, tell me! If a dereference of a pointer that could be null leads to potentially dead code being eliminated, tell me about that too! It's the silently surprising behavior that causes the most problems.
The problem with undefined behavior is not people ignoring portability. It's that it's actually really easy to accidentally misuse it. For example, Regehr's group has found quite a few such bugs in widely ported code written by smart people [0].
It's fairly non-trivial to tell a user "if you switch these 5 loop nests, we may be able to do something cool!"
INT_MAX + 1 is never 0 for unsigned int. (unsigned)INT_MAX + 1 is equal to INT_MAX + 1. You're thinking of UINT_MAX + 1, which is always 0.