Volatiles Are Miscompiled, and What to Do about It [pdf]
cs.utah.edu
cs.utah.edu
Imagine you work on a production quality C compiler. The pedestrian pieces (e.g., lexer, parser, AST) have been stable forever. The piece that you're very likely going to be working on is the optimizer (or maybe the code generator) to make use of new techniques or new instructions. While you're going about your business, you're probably not thinking about that arcane corner of the C language spec that discusses volatile. What's more, when you finally complete your feature, all of the compiler's test suites pass because there's insufficient coverage for volatile.
The biggest contribution of this paper, besides the fact that it identified this issue across various compilers, is the notion of access summary testing and advocating for it to be included as part of the test suite for C compilers.
Really, how often when you encounter a bug do you think it's a compiler bug? Never. It can't be. Compiler writers are infalliable.
You'll first think it's your program, or maybe you misunderstood how volatile works, so you'll read the spec again. You'll write a poodle to isolate the problem. That will reboot each time too. Then you'll think it's some odd race condition related to volatile. But you're just doing a load and a store. The reboot happens every time, OK, that's promising. Then maybe, MAYBE, if you're awesome, you'll think to look at the generated assembly. And when you realize you have a no-op, you'll start to think if you maybe inadvertently specified something wrong in your -O parameters. Because how could the compiler be wrong? It's never wrong.
Code generation bugs are the worst.
Blaming the compiler (or the hardware or the OS or...) is only a problem if you do so without investigating to see if your blame is well placed. Once you can investigate properly, then it's just another possibility in your bag of tricks.
If I'm that low level I'll probably take a look at the generated assembly, even if I think it is my fault. I don't think that's too uncommon either. It helps filter through the abstractions.
But then, I've found bugs in compilers and standard-libraries for embedded a few times already, they are much less battle-tested than your regular x86_64-linux-gnu-gcc. So at some point, I normally switch into "trust no one" mode and start reading disassembler outputs in the vincinity of the crash-site ;-)
In this case, I'd probably suspect the watchdog and add some print statements around the get/set. Then it would work. Then I'd remove the print statements and it would break again. Nothing says "compiler bug" like print statements magically fixing the problem. (Could be a timing thing too, of course. This must be why drivers log so much useless information.)
In particular, it says "For every read from or write to a volatile variable that would be performed by a straightforward interpreter for C, exactly one load from or store to the memory location(s) allocated to the variable must be performed."
This is wrong. It later kind of gets it right for C, explaining about sequence points, but it entirely misses that implementations are free to combine and eliminate multiple volatile accesses within the same sequence point.
Now certainly, most of what was reported were genuine bugs (and John reports a lot of correctness issues). But it does/did nobody favors to start with an incorrect definition.
❝An object that has volatile-qualified type may be modified in ways unknown to the implementation or have other unknown side effects. Therefore any expression referring to such an object shall be evaluated strictly according to the rules of the abstract machine...❞ (§6.7.3)
For combining multiple volatile accesses within the same statement, I think I cannot find an answer in the standard.
Sounds like "volatile" variables don't really provide good semantics for most uses even without considering compiler bugs, so it's better to just use explicit load and store macros or functions.
The "buffer_ready" in the paper is a very good example that I have seen many times in the real world. If anyone can share the "better solution" that avoids "volatile" (and works on a bare-bone ARM microcontroller for example), I would love to see it.
http://en.wikipedia.org/wiki/Memory_barrier
https://www.kernel.org/doc/Documentation/memory-barriers.txt
Memory barrier instructions:
http://infocenter.arm.com/help/topic/com.arm.doc.faqs/ka1404...
GCC example:
http://stackoverflow.com/questions/6751605/data-memory-barri...
As far as I understand, volatile doesn't give you the needed "barrier" semantics (what is not allowed to happen before or after, on the deep hardware level) if the code with "volatile" works, it can be just an accident.
Discussions about barriers (or volatile) are only meaningful in the context of related accesses, so explicitly mentioning this isn't really necessary.
volatile is an antipattern. Don't use it. Don't encourage other people to use it. It is not a thing which should be used, by you. To use it would be wrong, because not using it is correct.
Use atomic instructions.
Embedded programmers are paranoid. They check such things in the debugger (examine assembly code). Not the big issue the OP seems to think.
"Although the symptoms of this compiler bug—spurious periodic reboots due to failure to reset the watchdog timer—may be relatively benign, the situation could be worse, for example, if the hardware register were used to lower control rods, cancel a missile launch, or open the pod bay doors."
So that's what the problem was with those pod bay doors.
i'm a little surprised at how much worse clang fared.
Also, it seems there are quite a few compiler bugs being found, even today. This looks like a very productive field of study, though the same could likely be said for software correctness in general. http://blog.regehr.org/archives/1061
This particular paper's software (its modern equivalent) is at https://github.com/csmith-project/voltest