When you see a Heisenbug in C, change your compiler's optimization level (2010)
esr.ibiblio.org
esr.ibiblio.org
Yes, when you see something like this, changing the optimization level may indeed influence whether or not the bug is visible, and is a useful tool for figuring out what went wrong. But these days it's pretty rare that the optimizer is at fault. Usually what you'll find is that the optimizer is indeed making a valid optimization, but your code is relying on unspecified C or C++ behavior that just happens to work when you're at -O0.
Your code is still the thing that's technically incorrect.
C and C++ are specified as the behavior of a straight-line sequence of code run in isolation. Basically, any concurrent C/C++ program treads into the unspecified behavior area. Thusly it's not the optimizers fault, but C/C++ certainly aren't helping out at all.
Even the mechanisms that are common place to make a concurrent C/C++ program run correctly basically wade into unspecified behavior, for instance using asm("lock esp"), asm("mfence"), etc.
TL;DR: concurrency is unspecified in C.
>TL;DR: concurrency is unspecified in _C99_.
Since it is specified in C11's memory model.
One of the many reasons I'm eagerly the 1.0 release of Rust.
Does the same argument apply to binary code ran on a preemptive operating systems? The actual execution order of your code is still unspecified, as the OS an interrupt it whenever it wants.
Race conditions are a lot like unsanitized input. They don't cause problems by themselves, but if you make incorrect assumptions it's easy to write incorrect code.
Few programmers have a deep enough understanding of the C++ spec to realize things like how the optimizer is free to ignore NULL checks after a memory location has been accessed. They just see a `pointer == NULL` passing, then code failing on a null pointer.
And the problem is that they shouldn't need that level of understanding to write C++. The optimizer should be doing things like warning about unnecessary null checks that could degrade performance, it would notify the developer that there's a problem.
But that would cause issues for people working with large amounts of legacy code, so it's not done.
void f(struct some_struct* p) {
int x = p->some_field;
/* ... */
if (p != NULL) {
/* this block might be executed even if p = NULL */
}
}
Because reading `p->some_field` is already undefined behavior unless `p != NULL`, the compiler is free to assume that `p != NULL` is always true, and might avoid the check.If the memory access doesn't crash the program for whatever reason (maybe it got reordered somewhere else or eliminated as dead code or whatever, I dunno), then if you call that function with a NULL pointer, you fall into undefined behavior that might manifest as that check that you put right there being skipped.
No; if `p` is NULL, this function has undefined behavior. Full-stop. It is a 100% meaningless function as soon as `p` is NULL, because the "NULL check" happens after the pointer is dereferenced.
So the issue isn't that the compiler can make incorrect optimizations -- the compiler makes optimizations that are entirely correct, assuming that the code that you wrote isn't meaningless.
In fact, most of the compilers I've used actually don't do this, and are (in my view) all the better for it.
See also this rant on gcc's strict aliasing, borne of the same philosophy: http://robertoconcerto.blogspot.co.uk/2010/10/strict-aliasin...
That's what undefined behavior means. Semantically speaking, the C language assigns no meaning to that function if the input pointer is NULL, and is therefore "wrong" by any reasonable definition of the word if it is NULL -- so the compiler is free to make an array of optimizations based on the fact that the input pointer is not NULL.
``behavior, upon use of a nonportable or erroneous program construct or of erroneous data, for which this International Standard imposes no requirements.
``NOTE Possible undefined behavior ranges from ignoring the situation completely with unpredictable results, to behaving during translation or program execution in a documented manner characteristic of the environment (with or without the issuance of a diagnostic message), to terminating a translation or execution (with the issuance of a diagnostic message).'' (italics mine)
This sounds like a long way from "meaningless" in my book. To my reading, the purpose of undefined behaviour appears to be to avoid unduly constraining implementations by not mandating behaviour that could be inefficient, costly or impossible to provide.
You (or anybody else!) may disagree on how far this inch given could or should be taken. But I think the fact the standard explicitly suggests that undefined behaviour could do something reasonable is evidence that programs producing undefined behaviour do not necessarily have to be considered meaningless.
(As a concrete example I have worked on one system where NULL was a pointer to address 0, and where address 0 was readable. Not only that, but in fact address 0 actually contained useful information, and some system macros used it. It was some kind of process information block and so there was a whole family of macros that looked like "#define getpid() (((uint32_t * )0)[0])", "#define getppid() (((uint32_t * )0)[1])", that sort of thing. I'd say this is rather odd, but the standard would appear to allow it. (However, perhaps needless to say, gcc was not the system compiler.))
(See also, the approved manner for using objc_msgSend, since time immemorial.)
> behavior, upon use of a nonportable or erroneous program construct or of erroneous data, for which this International Standard imposes no requirements.
NO REQUIREMENTS -- so, semantically, programs that employ undefined behavior are completely meaningless.
With regard to the second quotation, the part you italicized is nice, but the part you didn't is just as important:
> Possible undefined behavior ranges from ignoring the situation completely with unpredictable results
The entire quotation basically says "when you write code with undefined behavior, anything can happen; the results can be unpredictable, or they can appear to be sensible. An error could also be triggered." But that's the point: no behavior is specified. There's no restriction to what might happen.
Take the function above that has a NULL dereference when the input pointer is NULL. That function could be compiled in such a way that it writes an ASCII penguin to stdout if the input pointer is NULL; it's totally within its rights to do that. Your mental model of how C programs work is entirely inaccurate if you expect undefined behavior in C to do something that you deem sensible.
(It may be OK for anything to then happen, but as a simple question of quality - and common decency ;) - an implementation should strive to ensure that the result is not terribly surprising to anybody familiar with the system in question. And I'm not really sure that what gcc does in the face of undefined behaviour, conformant though it may be, passes that test.)
1. Undefined: anything is permitted, the standard imposes no requirement whatsoever. A typical example is what happens when a null pointer is dereferenced.
2. Unspecified: anything from a constrained set is permitted. Examples include the order of evaluation of function arguments (all must be evaluated once though any order is allowed).
3. Implementation-defined: the implementation is free to choose the behaviour (possibly from a given set), but must document its choice. An example is the representation of signed integers.
> the compiler is free to assume that `p != NULL` is always true
Only until the first point in the "..." part where it cannot prove that p has not been modified.
http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
Please read that article, and the rest in the series. Undefined behavior is far more pervasive than you think.
This and the Apple SSL bug makes me think that optimising compilers should be far more explicit (i.e. emitting messages or even warnings) about what they're doing than they are now, because it seems far too much about the optimisation process is being hidden and not transparent enough. Not only unreachable code, messages like "result of computation x is never used", "if-condition assumed to be {false, true}", "while-loop condition always false - body removed", etc. would be extremely useful for detecting and fixing these problems.
1. Typically yes, and that is exactly why the compiler can remove the check after it, it's already undefined behavior so it can assume the optimistic case (and remove the check).
2. But it is not always true that it would crash. p->some_field might happen to be in a valid memory location to be accessed. This doesn't happen normally because low memory addresses (0-1024, say) tend to not be accessible by userspace programs. But I am not aware of a spec ensuring that. An OS could in theory let your program map memory at address 4, in which case p->some_field would succeed if p is NULL and the offset if 4.
That would introduce non-determinism in concurrent executions that might have nothing to do with the semantics of the program.
For example, I'm under the impression that the most recent C++ standard basically says "if you have no data races then all will be well" even though they may be benign data races like dual assignments of the same value.
I saw a large codebase completely break when turning on intermodule inlining. The cause: more opportunities for the compiler to find pointers that aren't allowed to alias.
That is 99.9% rubbish. By which I mean in 99.9% of cases, when code works at -O0 and not at higher optimization levels your code has a bug, almost certainly involving invoking undefined behaviour.
Certainly it is possible to find optimiser bugs, I have done it myself, but it is many times more common for your code to have bugs. Now, you could argue that optimizers shouldn't make such extensive use of undefined behaviour, which leads to these kinds of bugs, but it is perfectly valid for a C compiler to do so.
The only time I've seen optimization bugs is in the emscripten compiler, but I'm sure those will be ironed out eventually (and the current version seems pretty stable).
The most successful way of finding what's causing these "heisenbugs" is to NOT change the binary, but just load it in the debugger as-is, make a guess at where it's failing, and set a few breakpoints/watchpoints. Either the compiler is not generating the code that you expect, or your expectations of the language were wrong, and both cases are fairly clear from looking at the generated code.
Rushing to blame a bug in the compiler is, much more often than not, a way to waste a lot of time while debugging. Most commonly, we are the cause of our own bugs.
/// I am a comment! \\\
a = 5;
The SGI OCC compiler accepted it just fine, as did most of the other compilers. I wasn't until we compiled on HP, or perhaps AIX, that we had a compiler which elided the "\" + newline to give /// I am a comment! \\a = 5;
and caused the downstream code to fail.As for heisenbugs specifically, I agree with you. I just wanted to tell a story. :)
unsigned char initdata[] = {
0x01, // some value, e.g. 513
0x02, // as 16bit, little endian
0x40, // ASCII @
0x5c, // ASCII \
0x42, // ASCII *
0x09, // number 9
(...)
}; //\
this is a comment
/\
/ this is also a comment
GCC has an option (-Wcomment, enabled by -Wall) to warn about this and other likely mistakes involving comments (e.g. nested /* */ comments).If the generated code does things that can't be represented in the debug info, there are plenty of better solutions, with generating useless debug info being an obvious one. Useless debug info is annoying, but far from the end of the world, not least because just about any experienced programmer will be used to it by now...
In fact the code part of the result should be identical in both cases, and ideally at runtime it should be impossible to tell the difference between -g and not (since presumably the loader won't load/map the debug info even if it's embedded).
At least that's been my experience in almost 20 years of programming a lot of C and C++.
I've found my share of compiler bugs, particularly in embedded compilers, but 99%+ of the time the Heisenbugs are mine: stray pointers, stack overflows, buggy concurrency, unexpected interrupt timing, etc.
When you see a Heisenbug in C, it's probably NOT:
1. The hardware. 2. The compiler. 3. The OS.
It's not you (compiler), it's me.
Also known as "select isn't broken".
If you have made a 'fix' but you don't know what problem it solves and how it does so, you probably haven't fixed anything, and you can expect further trouble from the root cause.
I can't express the joy when the source of the bug ended up not being my fault.
That bug chase in particular made me extremely weary of people who alter systems to fix bugs without understanding exactly why it made the issue go away. It might not have fixed it but just moved it.
Compiler bugs tend to be more deterministic, where some code is being generated incorrectly and it will do the wrong thing every time you execute that code. Changing the optimization settings may cause the bug to go away, but it's still deterministic given the same compiled binary. One exception would be a bug where a compiler is reordering instructions across a memory barrier, but those are even rarer than normal compiler bugs, since compilers are generally paranoid about reordering anything with barrier intrinsics or inline assembly.
Which is a good idea even if you don't observe any bugs.
#pragma optimize(push)
#pragma optimize("",off)
void the_function()
{
}
#pragma optimize(pop)
Usually some kind of bisection, or lucky guess is done to figure out that it's really compiler bug.I don't know about newer versions, but VS 2010 and priors have a nasty trait of spitting out wrong binary code now and then even when all optimizations off, edit-and-continue and incremental linking disabled. You would just make a little change, build, launch and get a crash in the middle of nowhere with the most bizarre stack trace. Then rebuild it and it will magically go back to normal. For a project of 100k lines this is a weekly, sometimes daily, routine. Very trust inspiring.
Granted, when the debug symbols end up in the final ELF binary which is executed by an OS, it can still affect execution in subtle ways such as load times.
We were using a proprietary C compiler tool chain provided by our vendor and did not get much help from them either.
Finally, We had to sit down, get the assembly from disassembler and went through whole 800 lines of it. And we found the bug, sitting quietly in one of the pipelines.
I am not a compiler guy, but that day I understood the beauty of compiler optimization.
For example, on one system I used there was a ~2 second difference between my desktop machine, the build machine, and the NFS host for directory I was working in. This would sometimes lead to:
make: warning: Clock skew detected. Your build may be incomplete.
If I did a save to a .C file just after the .o file was written, then it might still have a timestamp which was older than the .o, so not included in a rebuild.This required occasional manual deletes of the .o file to make sure that it was building correctly, and we probably did a full rebuild of the entire code every once in a while to reduce these sorts of problems. (This was back in the c-front days, with 15 minute compilation times.)
A decent percentage of the time someone is experiencing a "heisenbug" it's because there are some stale object files sitting around, and the build state has gotten inconsistent (makefile bugs, bad tracking of dependencies, or what have you).
You go to build debug, everything works, you shrug, make clean, remake, and then you're golden again.
Obviously that's really not the way to go, though, because you risk mistaking an uncommon crash for "just had to rebuild, I guess" - but if you're comfortable reading assembly or what-have-you, you can usually confirm the "stale object files" thing easily -
and then fix your build ;)
1. lupus 2. a tumor 3. a compiler bug
I'm not at all surprised to see ESR being enough of a scrub to go to 'compiler bug!' as his first guess.