Dangerous Optimizations and the Loss of Causality in C and C++ (2010) [pdf]
pubweb.eng.utah.edu
pubweb.eng.utah.edu
if (cond) {
A[1] = X;
} else {
A[0] = X;
}
> A total license implementation could determine that, in the absence of any undefined behavior, the condition cond must have value 0 or 1.> If undefined behavior occurred somewhere in the integer arithmetic of cond, then cond could end up evaluating to a value other than 0 or 1
I don't even follow this very first example. Neither C nor C++ have any requirement that their condition statements be given 0 or 1.
In C:
> the first substatement is executed if the expression compares unequal to 0.
In C++:
> The value of a condition that is an expression is the value of the expression, contextually converted to bool
Where conversion to bool is defined as:
> A zero value, null pointer value, or null member pointer value is converted to false; any other value is converted to true
I'm sure such a thing could occur if cond had a boolean type but contained uninitialized data. That would be similar to the situation talked about in this link: https://markshroyer.com/2012/06/c-both-true-and-false/
I think what they mean is that if the compiler has analyzed the code, not shown in the example, that comes before the if statement, and concluded that unless something undefined happens, cond can only be 0 or 1, then it can optimize out the condition.
They are not saying that the code actually shown is enough to conclude that cond must be 0 or 1.
> could determine that, in the absence of any undefined behavior
"could determine that" based on the code example shown
vs
"could determine that" based on static analysis performed on some preceding code
It would have been a lot easier to wrap my head around if it were an example where cond could be 0 or 4 or something along those lines. It would really underscore the compiler's desire to reuse the cond as the index.
I might not code daily in C++ anymore, but it would be nice that the software I rely on, specially in devices I am not able to upgrade, wouldn't suffer from such issues.
Butcher knifes are deadly sharp, yet those conscious ones do use proper gloves when handling them.
But alas, one was too expensive to become mainstream and the other died when Compaq acquired DEC.
The big difference is that you as developer will notice that you are doing something that might eventually blow up, because it is very explicit.
For example something like doing a reinterpret_cast would be LOOPHOLE call in Modula-3 and is only allowed inside an UNSAFE MODULE implementation.
As for the rest, yes C and C++ have enjoyed 30 years of compiler research, but as Fran Allen puts it, C has brought compiler optimization research back to the pre-history.
If you read IBM's research papers on PL.8, the systems programming language used on the RISC project, it already had an architecture similar to what LLVM has. Where the multiple compiler phases were taking care of the optimizations. Then when they decided to go commercial, they rather adopted UNIX as platform.
So maybe we would be somewhere else had AT&T been allowed to sell UNIX at the same price as VMS, z/OS and others.
As for the LLVM, we aren't going to throw that code away, but as you also assert, it is possible to use safer languages while taking advantage of the optimization work that went into LLVM.
That is how I see C and C++'s future, down into a thin infrastructure layer, with userspace mostly written in something else.
In C code, any cast at all, of any kind, should be viewed with the utmost suspicion.
That's why unions have been chosen as the correct way to alias pointers. Because it is unlikely that a programmer accidentally used the wrong type or that code changes made a type cast obsolete.
And in C all bets are open.
As for the future, I'm slightly ambivalent. On the one hand, common sense would lead me to agree with you that, with the exception of some numerical code (where I'm hopeful for further improvements in JIT compilation), most userspace code should be written in languages other than C/C++, retaining most of the efficiency of the underlying infrastructure. On the other hand, there are languages out there (like D, for instance) that essentially fix all the glaring issues of C and fail to get significant adoption. Weird as it may seem from the confines of the HN bubble, usage of C has actually grown over recent years and betting against C has historically turned out poorly.
It will be impossible task to kill C on UNIX clones and embedded due to culture and existing tools.
However the world of application development is much bigger than just UNIX systems programming or embedded development.
The reason is backwards compatibility. C++ has a long history, and mistakes were made in the past including mistakes carried over from C:
https://www.digitalmars.com/articles/b44.html
and C++ tries to fold improvements in without breaking compatibility.
That's actually not completely true.
Take the restrict keyword for instance. Super painful to use in C/C++ and very dangerous. In Rust since the language has constrains around single mutable pointers they can turn it on globally and you get it "for free". If I recall correctly it's going live in one of the upcoming stable releases.
Sometimes constraints can actually let you get better performance than a wild-west of pointers everywhere.
If I remember correctly, it was originally enabled but had to be disabled because LLVM had some bugs around how it was handled. These arose because `restrict` is comparatively rarely used in C / C++ code.
There are two caveats to that statement. The first is JIT compilation, for obvious reasons. The second is that some domain specific languages with very restrictive semantics could do memory layout optimizations that are off the table by definition in any general purpose language (transform AoS patterns to SoA, etc.).
Having had to chase down stray aliased pointers in something marked restrict I can tell you it's not trivial to root-cause.
In C++ there is no restrict keyword. So either you're using and abusing compiler-specific behavior that you decided to do so without any test at all or you're confusing programming languages.
Either way, you're trying to pin the blame on C++ for something that has nothing to do with C++.
That speaks of the complexity of C++.
When you say: "you can select what you're comfortable with", I am guessing that you are not using c++ in production environment.
That is not a questioning, but to point out that if that's true, you generally have no experience to support such claims.
People just want to "optimize" the code.
(Also, different optimizations can have surprising effects when combined that isn't clear at all from the individual levers.)
On a large codebase this is very hard to do, and introduces an additional layer of complexity for very little reward.
There is no "invalid C++, but valid C++-noX-noY-noZ". It's one or the other. The former is a bug in the code; the latter is a bug in the compiler.
(x * 2000) / 1000 can be reduced to x * 2 but only if we make some assumptions about x (about if x can overflow)
The reason why this is such a problem in C languages is because most of the time one is working in a mixed abstraction environment. The underlying model for many parts of the language is assembly. In order to get away from these issues and to produce more optional code we should move away from mixing models of abstraction.
IMO the wierdist thing to read in C++ though is X / A A. Its roughly the same as X -= X % A. They values may differ if X or A are negative.
Good practice as per my experience is to check MCDC and code coverage at unit test level
Those are reserved for people who value speed over correctness.
First time I heard it was about 1993.
#include <stdio.h>
int main() {
int a, b, c;
printf("Please enter two numbers: ");
scanf("%d %d", &a, &b);
c = a + b;
printf("The sum of the numbers you entered is %d\n", c);
return 0;
}
This is the problem. There is no non-trivial C code that doesn't use some undefined behavior somewhere. And it works just fine, right now. But who says it will on the next version of the compiler?That is not good. It's hard to reason about such a program, in other words can you trust results of such a program? How do you know whether your input caused Ub at some point or not?
The burden of sanitizing the input is on the user (either programmer or the data provider).
Checking this is usually a performance tradeoff, so it was decided not to be done by default.
(And you're missing a few error checks.)
in this precise case, (in C++ because I can't be arsed to search the modifier for int64), fairly easily :
#include <fmt.h>
#include <cstdio>
int main() {
int a{}, b{};
int64_t c{};
fmt::print("Please enter two numbers: ");
scanf("%d %d", &a, &b);
c = a + b;
fmt::printf("The sum of the numbers you entered is {}\n", c);
return 0;
}
in the case where a, b, were also of the largest int type available, you could do : #include <stdio.h>
int main() {
int a, b, c;
printf("Please enter two numbers: ");
scanf("%d %d", &a, &b);
if(__builtin_add_overflow(a, b, &c))
{
printf("The sum of the numbers you entered is %d\n", c);
}
else
{
printf("You are in a twisty maze of compiler features, all different");
exec("/usr/bin/0ad");
}
return 0;
}
but you could also use the more portable macros provided in emacs (100 % independent of any other code) : https://github.com/jwiegley/emacs-release/blob/adfd5933358fd...https://blog.regehr.org/archives/1139
HN discussion:
https://news.ycombinator.com/item?id=7665254
Your first suggestion is similar to his checked_add_1(). The Emacs macros are similar to his checked_add_4(). Presumably builtins are optimal, performance wise, for a given compiler.
FWIW, I tried this example:
#include <stdio.h>
#include <limits.h>
#include <stdint.h>
int main () {
int64_t x;
int a = INT_MAX, b = 1;
x = a + b;
printf( "%lu %lu %d %lld\n", sizeof(x), sizeof(a), a, x);
return 0;
}
After compiling gcc -std=c11 -fsanitize=undefined -Wall -Wpedantic -Wextra (getting no warnings), on running I get c-undef.c:8:9: runtime error: signed integer overflow: 2147483647 + 1 cannot be represented in type 'int'
8 4 2147483647 -2147483648
Also with -O0 and -O3.At least according to the answer in [1], signed (but not unsigned) integer overflow is undefined behavior (unless C and C++ have diverged on this matter, but I get the same result when the same source code is compiled as C++17.)
This may seem to be pedantic (if it is not just wrong), but the point is that you have to be unremittingly pedantic to avoid undefined behavior.
[1]https://stackoverflow.com/questions/16188263/is-signed-integ...
that's 100% correct, nice catch.
"Portable C" doesn't exist.
in C++ you'd just add
static_assert(sizeof(int) < sizeof(int64_t), "fix your types");
to be safe about thisIt's a performance hit though.
I don't think you'd want a banking app written with assumption that you'll never have more than MAXINT amount of dollars, and adding few more will roll you back to 0 or even put you in debt... silently... because well it works this way, and designer didint think you'll hit the limit. If it ever happened you'd like a siren to go off.
Yes, and it could also decide to wipe your hard drive because of the latitude given to it by the C standard. Many compilers have an option to enable some sort of special behavior when a signed integer overflows, but such extensions are non-standard.
Reference: N1570 (C11) 7.21.6.2 paragraph 10: "If this object does not have an appropriate type, or if the result of the conversion cannot be represented in the object, the behavior is undefined."
strtol() avoids this problem.
So yes, they kept the verified code small, so that they could focus on all of the tooling. I'm unconvinced that formal methods intrinsically can't scale.
Almost nobody does because formal code writing forces a constraint / logic / contract thinking which is alien to many alleged programmers.
It could be taught. Afterwards, compare performance and quality of the result. Could be a big competitive advantage...
The fact that after four decades it has not proven to be so is evidence that either it is very difficult to put into use or not that effective.
Also, I'd add that when you say
> We already have a good automated test culture, but they can only prove everything we thought of works, we have lots of issues with situations we didn't think about
in a lot of ways formal verification only helps you with situations you know about, specifically problem classes that you know about. For instance sel4 was susceptible to Meltdown.
... no, C explicitely has no well defined semantics since it's undefined behaviour. You may believe that C do due to habit but that's not the case.
when you program in C you program against the C abstract machine, not against a particular hardware.
> C has perfecly well defined semantics for (eg) null pointer dereference: read or write memory location zero, and consequently probably crash.
C does have "perfectly well-defined semantics" for null pointer dereference, but it's undefined behavior. Sure, the null pointer happens to be 0 on the vast majority of architectures most programmers work with, but apparently non-zero null pointers are still used these days (at least in the admgcn LLVM target, from what I understand after a quick glance at the linked page), so it's not even a "all sane modern machines" thing [0].
In any case, I'd love to see some good benchmarks showing the effect of the more "controversial" optimizations. Compiler writers say they're important, people who don't like these optimizations tend to say that they don't make a significant difference. I don't really feel comfortable taking a side until I at least see what these optimizations enable (or fail to enable). I lean towards agreeing with the compiler writers, but perhaps that's because I haven't been bitten by one of these bugs yet...
[0]: https://reviews.llvm.org/D26196Personally, I mostly use C/C++ for fairly high performance numerical code and happen to benefit greatly from all the "unsafe" stuff, including optimizations only enabled by compilers taking advantage of undefined behavior. I'm therefore naturally inclined to strongly oppose any attempts to eliminate undefined behavior from the language standards. At the same time, however, I fully recognize that most people would probably benefit from a safer set of compiler defaults (as well as actually reading the compiler manuals once in a while) or even using languages with more restrictive semantics. Ultimately, there is no free lunch and performance doesn't come without its costs.
"C is for speed only"... Do you not realize there is an entire embedded world that runs on C?
https://www.ptc.com/en/products/developer-tools/objectada
https://www.ghs.com/products/AdaMULTI_IDE.html
http://www.astrobe.com/default.htm
Just a couple of examples from many more, after all I don't want to overflow you with information.
I call that production deployment.
As for myself, feel free to believe whatever makes you feel better.