A glimpse of undefined behavior in C
blog.chris-cole.net
blog.chris-cole.net
Really, that's one of the reason why Java still works well at big companies. The code is usually so verbose and 'un-smart' that it's quite hard for programmers with minority complex to obfuscate it when they're trying to prove that they're smart. It's a shame however that if the language is primitive then you can always create a big framework to obfuscate stuff. :)
Hear, hear. I am considered the "language lawyer" of my embedded group and often get asked questions about C minutiæ. It isn't uncommon that my answer is "I don't know how that works, because I would never write something that requires an answer to that".
Modern compilers are blind to shorthands etc.; the only useful knowledge about "C" (really C compiler & hardware) minutiæ relates to (a) how to convince the compiler to optimize certain high-level constructs (e.g. loop unrolling), (b) how to convince the compiler to emit certain low-level constructs (e.g. SIMD instructions), and (c) how the hardware behaves (e.g. the cache model). Generally anything else – you can rewrite it so you don't have to think too hard about how C works.
EDIT: Understanding type promotion is an exception to this. The integer type hierarchy is unfortunately (a) deeply baked into C and (b) mostly brain-dead – C mostly conflates physical integer width with modular arithmetic and provides no non-modular integer types, often leading to subtle software bugs (see last week's story about binary search).
EDIT: So is understanding const-ness. At least you can ignore this if you don't get it (or if C doesn't, as is the case with certain nested const types).
More info here: http://blogs.msdn.com/b/vcblog/archive/2007/06/04/update-on-... and here http://en.cppreference.com/w/cpp/language/eval_order
int a[] = {10,20,30};
int r = 1 * a[i++] + 2 * a[i++] + 3 * a[i++];
with int a[] = {1,2,3};
int r = 1 * a[i++] + 10 * a[i++] + 100 * a[i++];
Then if your program outputs r = 111, it's obvious that it's doing a[0] + 10*a[0] + 100*a[0]
and if it outputs 321, it's obvious that it's doing a[0] + 10*a[1] + 100*a[2]
No disassembly required.What exactly is not undefined behavior? The post describes a statement where a variable is modified multiple times between sequence points. What does the standard say about this?
N1256, 6.5 says
Between the previous and next sequence point an object
shall have its stored value modified at most once by the
valuation of an expression.
J.2 Undefined behavior The behavior is undefined in the following circumstances:
[..]
* Between two sequence points, an object is modified more
than once, or is modified and the prior value is read
other than to determine the value to be stored (6.5).
What's with the know-it-all voice?In some cases the behaviors is well defined, if there are sequence points between the assignment, but I would rather keep my code simple.
This article goes about it the wrong way.
Classic example: signed integer overflow. It worked for decades. Then one day it didn't.
If you want to know about this particular example: https://news.ycombinator.com/item?id=6824514 (not personally confirmed)
Compiler users generally want something that "just works" and doesn't do anything unexpected, but in the case of a low-level language like C, doing away with undefined behavior essentially would mean pessimistically avoiding many optimisations on the 99%+ of straightforward, reasonable code out there in favor of not doing anything surprising on the remaining fraction of dubious code that depends on certain things happening in scenarios where behavior is undefined according to the C standard. There are languages that make that choice, but C isn't one of them.
I have fixed feelings about the integer overflow issue because it's so easy to trigger, unlike triple post increment fake examples. And it usually results in a security problem. For very little benefit, IMO.
Clang says:
zsh% clang -o sequencepoints sequencepoints.c
sequencepoints.c:7:18: warning: multiple unsequenced modifications to 'i' [-Wunsequenced]
int r = 1 * a[i++] + 2 * a[i++] + 3 * a[i++];
^ ~~
1 warning generated.
And prints: zsh% ./sequencepoints
140will give "0 0 1" (gcc 4.2.1), the increment "shouldn't" happen until after the ; if you're going with post statement. I think your rule would expect "0 0 0" with a being 2 after the printf.
BUT! you get a warning, so that's nice.
int a = 0, b, c;
b = a++, c = a++;
But, there is no defined order in which to evaluate function arguments."0 0 1" is what I would expect from your code, and identifier a should end up as 2.
/* x++ */
inline int postincrement(int *x) {
int temp = *x;
*x = *x + 1;
return temp;
}
/* ++x */
inline int preincrement(int *x) {
*x = *x + 1;
return *x;
}
But in typical usage, where the expression value is not used, it doesn't make any difference: for (int i = 0; i < 10; ++i) {
If one looks at the assembly, it's clearly equivalent to: for (int i = 0; i < 10; i++) {
Incidently, the latter seems more readable.Compilers are sometimes smart enough to remove the unnecessary operation in C++ (e.g. switch to ++x themselves), but I always use ++i for the same reason people simplify their usage of C in between sequence points --- it's an easy transformation, and why tempt fate?
To her credit, she didn't penalize me for disagreeing with her and I still got the offer.
Here's a good article that addresses the question "Why have undefined behavior?"
http://blog.regehr.org/archives/213
This one linking to the above is also worth reading:
http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
I always imagined undefined behaviour was there to avoid tying implementations' hands, by having the standard not mandate things that vary in practice. But it seems that people are assuming it's there to give compiler writers carte blanche to do whatever they like, and then point at the standard as justification. Given how much stuff could potentially be added to C to improve it, I don't know why people are spending all this time trying to figure out all the ways in which the letter of the law allows them to confuse the programmer.
See also somebody else's rant about strict aliasing: http://robertoconcerto.blogspot.co.uk/2010/10/strict-aliasin...
the answer will be different in different compilers. I use this as an interview question ever since, not to get the right answer but to understand the candidates thought process in solving the problem and his/her understanding of operator precedence :-)
Yo always do multiple lines, with comments on each one.
It sounds ridiculous, but this simple thing made some kind of bugs impossible: those that you have in front of you but you can't see in a million years. The atomic operation in code is the line, you can't debug a complex line(1).
It is painful forcing people to do that, people use to hate being told what to do, but at the same time they love the outcome so much. In the end everybody loves it.
1.With assembly you are debugging an instance of your code. You are not debugging what will be created with any compiler, any os or architecture.
I mean suppose I split the first part of the statement like this:
int r = a[ i ];
++i;
These lines are self-explanatory. It's pretty hard to argue they need a comment? Especially when used in a function that already should have a name/comment explaining what it does?The ++ operator should be deprecated except when it's the only operation on that line.
(int i=0;i<10;i++)
would count as three lines for purposes of that rule, and the i++
as one. I could have been more precise, but I thought it would be understood.The point is, to be on the safe side, don't use ++/-- in any line (in the sense of ;{}-delimited statement) that is also doing something else.
>Don't forget that =, += , ... can all be abused in the same way.
I didn't, and people should use the same safeguards around them, i.e. don't mix them within lines that do other things, "cleverness" or "C golf" be damned.
foo foos[255];
int ct = 0;
if (<condition for adding first element>)
foos[ct++] = foo1;
if (<condition for adding second element>)
foos[ct++] = foo2;
...
for (int i = 0; i < ct; i++) {
/* do something with foos[i] */
}
IMHO, this is more concise and readable than separating the increment and assignment into two lines. (prefix ops) => ( statement ) =>(postfix ops)
Which made code like *a++ = *b--;
Change the pointers after the copy as opposed to having them change before the copy.It is interesting that this has become compiler defined.
That's something that took me way too long to learn and I feel that novice programmers should embrace that more. We all like to feel like Einstein sometimes, but a big code base is not the place for solutions only we can understand... temporarily.
You can never rely on your compiler's implementation of undefined behavior.
I wish I could somehow get my coworker to understand this. He's mostly "Yeah, well, I know it's undefined, but there's really no way this could be anything else than <foo>. I know what the CPU does there."
http://en.wikipedia.org/wiki/De_facto
:)