The inflamed backlash should tell you just how damaging it is to impose silent failure on meticulously written, previously fine programs.
The inflamed backlash should tell you just how damaging it is to impose silent failure on meticulously written, previously fine programs.
With relatively few exceptions, if your program hits undefined behavior, then your program was already doing something pretty wrong to begin with. Signed overflow is a poignant example: in how many contexts is INT_MAX + 1 overflowing to INT_MIN actually sane semantics? Unless you're immediately attempting to check the result to see if it overflowed (which is extremely rare in the code I see), this overflow is almost certain to be unexpected, and a program which would have permitted this was not "previously fine" nor "meticulously written."
I feel compelled right now to point out that software development is a field where it is routine to tell users that it's their fault for expecting our products to work (that's what the big all-caps block of every software license and EULA says, translated into simple English).
The code was written in Pascal and assembly, though, so it was safe from a C compiler.
As such my vision of low level coding isn't tainted by the ways of C.
It's not that rare - I know that postgres got bit by particularly that issue, and several other other projects as well. Particularly painful because that obviously can cause security issues.
int a;
// lots of code
int b = a + 1;
// check for overflow
if (b <= a)
abort();
the compiler is allowed to remove the check because in the absence of UB a + 1 > a, therefore the conditional is always false.In that case, something like: [...] if (a > INT_MAX - 1) abort(); int b = a + 1; [...]
Much easier said than done:
https://github.com/postgres/postgres/blob/master/src/include...
Yeah, a standard library function that does this would be good. But many people would just use * instead of this function, and so the problem would partially remain.
In c++ it would be a little simpler due to templates, so the types of the arguments to the function can be derived. But the type of the result can still confuse programmers. Although maybe it's not so bad, because an overflow that happens during multiplication (or other math operations) is undefined-behavior, but an overflow that happens during assignment of the multiplication result to a variable that can't hold it is only implementation-defined behavior, not undefined behavior.
The author claims: “We have the absurd situation that C, specifically constructed to write the UNIX kernel, cannot be used to write operating systems. In fact, Linux and other operating systems are written in an unstable dialect of C that is produced by using a number of special flags that turn off compiler transformations based on undefined behavior (with no guarantees about future “optimizations”). The Postgres database also needs some of these flags as does the libsodium encryption library and even the machine learning tensor-flow package.”
So basically the programers of the most used C programs consider the C standard so broken that they force the compiler to deviate from it. (Or they are not able to do things right?)
Standard is just common denominator agreed by actual competitors. (Standard takes into account the chipset where incrementing max_int produces a beep instead of min_int). If user wants to use more advantages of the product (dbms/compiler/hardware), one must sacrifice portability and use cnonstandard extensions.
Well, I learned of the change in compiler behavior some years back because I had written loop code with a sanity check which depended on signed integer overflow wrapping, along with a test case to prove that the sanity check worked, and that test case started failing:
not ok 2 - catch overflow in token position calculation
Failed test 'catch overflow in token position calculation'
at t/152-inversion.t line 70.
To the extent I can be, I'm done with C. I leave it to people who think that silently optimizing away previously functional sanity checks is an acceptable engineering tradeoff, and who disparage those of us who have been bitten.To be fair, as several people and TFA have pointed out, this isn't a problem with C, but with defective/malicous C compilers. Admittedly, that's not much help if you can't find a compiler that isn't defective/malicous, though, so I can only wish you the best of luck.
Do you want compilers to stop adding optimizations while staying within the bounds defined by the spec? That they somehow guess that a given piece of code that may trigger UB is too important for them to optimize it based on the assumption that the developer knew what she was doing and ensured that it wouldn't?
The case of a compiler update breaking the test sounds like desirable to me. It pinpoints a critical piece of code that was relying on a specific implementation's behavior that is not specified. This could have been triggered by a compiler or architecture change. If this is something that is a hassle to fix immediately, you can temporarily downgrade to the version of the compiler used earlier or disable some optimizations.
I fail to see why one would consider a compiler evolving while conforming to language specification defective or malicious (however, with that definition I fear that finding one that isn't may indeed be difficult).
It would be fine if I could opt into new semantics — something like Rust's "editions" would resolve my objections about these compiler optimizations. But that doesn't seem to be on offer in the C ecosystem.
C culture and to certain extent Objective-C and C++ ones are tainted by microptimizaitons while typing, where the compilers are the worst examples.
Unfortunely UNIX and C go together, so those that want to keep UNIX like platforms around bettter fix C somehow.
Yet Azure Sphere, despite its security message, uses C only SDK.
Meanwhile Visual Studio now supports C11 and C17.
Market pressure, their customers weren't willing to buy into it, and the new Microsoft also wants all those POSIX FOSS packages written in C running on Windows.
Hopefully the newer compiled languages will kill C++.
Quoting from another of my comments:
> > [What are you objecting to?]
> Inferring any propositional statement about the program (eg "this pointer is not null") from the fact that its negation would imply undefined behaviour.
That is what the problem is. Undefined behaviour is a licence to implement operations without regard for unusual corner cases, not to infer the absence of said corner cases from those operations and then apply that 'knowledge' elsewhere.
> I fail to see why one would consider a compiler evolving while conforming to language specification defective or malicious
It's https://en.wikipedia.org/wiki/Malicious_compliance in near-textbook form.
I'm not completely sure I understand that correctly, but do you mean that statements that are constant unless considering a possible (and "credible"?) implementation of UB shouldn't be fair-game for the compiler to optimize out? EDIT: I think I see a more tricky case that may be one of those you're referring to. Dereferencing a pointer further in the code shouldn't be a valid justification for optimizing out previous tests of it being null. I can relate with that but I suspect that it would prevent many classes of branch pruning.
I get what you suggest while describing malicious compliance but I can imagine that it could be a false impression resulting from trade-offs that favor optimization opportunities to "out-of-spec but canonical/natural" implementations.
Not further. Anywhere. If the implementation wishes to rewrite pointer dereferences from `use(*p)` to:
if(!p) abort();
use(*p);
it may do so (undefined behaviour!), but if it chooses not to do so, it may not later pretend that it did, and remove a explict `if(!p)` that the programmer wrote. Given: use(*p);
if(!p) return NOPE;
this is fine: if(!p) abort();
use(*p);
//if(!p) return NOPE; // unreachable because of if, not because of use
but not: use(*p);
//if(!p) return NOPE; // CVE-20XX-XXXXX
The implementation can optimize based on information (like "p is not null") that is actually true (even if that's because it made it true), but not based on information it assumed was true on the basis that it counterfactually could have made it true (but didn't).> it would prevent many classes of branch pruning.
Yes, that's the general idea.
It dereferences a null pointer, invoking undefined behaviour, then calls the function `use` with the resulting value.
> what should the compiler emit?
Probably something to the effect of:
ld r0 [sp+.p] # if p is not already in a register
ld r0 [r0] # *p
jsr use
but it would be fine to emit something like: ld r0 [sp+.p] # if p is not already in a register
jz r0 .panic
ld r0 [r0] # *p
jsr use
because the jz can only be taken when undefined behaviour happens.Conforming to the spec is not a virtue. We want the compiler to be reasonable, regardless of whether the spec is. When the spec is malicious, conforming to the spec is malicious behavior.
For example, the Java spec says that the << operator (bit shift left) will accept any right operand, but performs a modulo-32 operation on the right operand before doing the shift. So `a << 4` is `a` shifted 4 places left, `a << 20` is `a` shifted 20 places left, and `a << 36` is `a` shifted 4 places left again.
This is absurd, and I'm comfortable calling it a bug in the spec. `a << 40` needs to have 0 in the lowest 40 bits. It does not need to have random values in bits 8-31.
This behavior is documented, but that doesn't make things better, it makes them worse.
But the philosophy that says "if it's documented, then it's OK" doesn't even allow for the concept of a bug in the spec.
It had code like:
int *p;
// lots of code
if (p != NULL)
return 1;
// use p
Then a later refactor wrongly added a single line before the if statement: int *p;
// lots of code
int a = *p;
if (p != NULL)
return 1;
// use p
This meant the null check was optimized away, since de referencing a null pointer is undefined behavior, so the if-statement can be assumed to be always false.
This then lead to actual errors (perhaps even an exploit, I do not recall) arising from the removed null-check.I think in general the "sanity check" cases are the worst. It is hard to determine whether an expression causes undefined behavior if you cannot try and evaluate it. Perhaps a compiler intrinsic that checks (at runtime) whether an expression causes undefined behavior could be useful here. Though I can imagine such an intrinsic being essentially impossible to implement.
In user code GCC will still happily remove the null check because in user code this is an actual bug in the user's code.
struct foo { ...; bar_t bar[NBAR]; };
struct foo* p = ...;
bar_t* q = &p->bar[0]; // add rq, rp, #foo_bar_offs
// other declarations
if(!p) return NOPE; // optimized out
// use p and q
Not even any dereferencing, just pointer arithmetic.There's some disagreement on whether you can call a program "fine" that breaks after switching to a newer version, or a different compiler.
I see a lot of programmers out there that unfortunately use the behavior of their code on whatever compiler they're using at the moment as a proxy for what the language actually guarantees.
a = 5 //* junk here */ 2 ;
I'm sure there are ioccc or underhanded C contest entries doing this to make code work differently on compilers based on if // is starting a comment line or not.
Sure it is an ugly way of writing stuff and you'd be hard pressed to find lots of real world traps like this, but when/if you did have code that "suddenly" miscompiles you might actually think your old code with an old compiler did work, and a new compiler for "the same" language breaks your program. I don't think everyone code base should need full rewrites ever time a new compiler comes out.