int i = 0;
f(i++, i++);
More examples here: https://www.cl.cam.ac.uk/teaching/1415/CandC++/lecture10.pdf int i = 0;
f(i++, i++);
More examples here: https://www.cl.cam.ac.uk/teaching/1415/CandC++/lecture10.pdf f(1, 0);
Because function arguments are pushed in reverse order in C and funnily enough, GCC actually does that even though that's crazy. But clang catches it as an "unsequenced modification".I don't think that's a good example. The C standards make no guarantees about unsequenced modifications so relying on them is the programmer not using C, it's the programmer using a language derivative implemented by their compiler.
Why do people say "C remains central to our computing infrastructure but still lacks a clear and complete semantics" when the standards are very clear about what it allows? If you use something it doesn't allow you can't blame the spec for that.
No it's not. It's undefined behavior meaning, by definition, it's not C. It's not defined in C. If I had an list of all definitions, would an undefined item be in my list? No.
When you are using Undefined Behavior, you are not using C. You're using a C-extension that a compiler designer has implemented. This was done because there is no sane operation to do in these UB cases for one reason or another. Either it would not be portable or it wouldn't make sense.
If there was a standard that implemented standards for UB cases it would not be portable.
To me, that's like saying the tax code doesn't contain any loopholes, because any loopholes that it contains are, by definition, not actually part of the tax code.
While amusing, I don't think this is a fruitful way to look at the problem. :)
The C99 standard defines undefined behaviour as
"behavior, upon use of a nonportable or erroneous program construct or of erroneous data, for which this International Standard imposes no requirements"
Undefined constructs are not necessarily erroneous (and therefore, in a way, "not C"). Many of them are perfectly legal constructs for which the standard you're targeting simply imposes no specific requirements. In fact, C standard documents list several (as in, I think, hundreds of) examples, many of them perfectly correct in terms of syntax. The standard just doesn't dictate what the compiler should do with them.
If you really like playing the words game, you could, at best, claim that undefined behaviour isn't C99 or C89 or whatever your compiler is targeting at a specific time, but it's very much C.
> It's not defined in C. If I had an list of all definitions, would an undefined item be in my list? No.
And yet Annex J of the C99 standard has a list of undefined behaviours.
And that list itself is just a reference, because the C standard is otherwise full of cases where it does not dictate requirements -- effectively, you know, defining a bunch of undefined items. E.g. 6.2.4, 2 in C99, where object lifetimes are defined, states that "If an object is referred to outside of its lifetime, the behaviour is undefined".
Besides, come on, look at code like this:
int foo = INT_MAX;
foo++;
How exactly is this "not C"? Can you think of any perfectly standards-abiding compiler that will not compile this? Does it use any keywords, identifiers or grammar that a C compiler does not understand?
Yes you are using C when you're using undefined behavior. It's one of the cornerstones of C's "design." The standard is poorly defined enough that different people building C compilers will come up with their own interpretation of what specific, legal, C statements mean. These are statements used in real projects that also compile with real compilers. The massive problems that resulted led to static analyzers and compiler checks for undefined behavior. Plus, people recommending you avoid it wherever possible for predictability, safety, and security. It is legal C, though.
Whereas, in languages like Pascal or Ada, you could express similar operations with their semantics making it totally clear what you could expect. They'll either not compile or include runtime checks to avoid similar problems wherever possible. C's designers just didn't give a shit: they were using what modifications of unsafe BCPL they could to squeeze max performance & memory utilization out of a PDP-11. Wirth's Modula-2 at least gave a clean language with many checks & sanity by default that you could selectively turn off if absolutely necessary. Not C, though.
So, it's legal, real C designed without safety, clear semantics, and so on. Different projects implement "real C" differently due to unclear semantics. Different, broken executions result.
#define f(a, b) printf("%d %d\n", a, b);
This would change absolutely nothing. A macro is just a piece of code that is expanded. It still follows all of the rules (and disallowances) of C. #define f(a,b) \
f_var1 = a; \
f_var2 = b
Equally you could use macros to make it a whole lot worse #define f(a,b) \
f_arr[a][b] += f_arr[b][a]For all the criticism Haskell gets over the purported difficulty on describing sequential computation, at least the order of monadic operations is clearly captured by the requirements to be a monad. There is no such concept for C, except that each compiler mostly seems to perform operations sequentially.
Source: http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1256.pdf page 439 (!) Note that Annex C is merely informative, but it contains references back to the actual standard where these things are defined.
Edit: I think a better advantage of Haskell is that Haskell is much more free to not define sequence points. Haskell effectively defaults to not sequencing expressions unless you use a monad or there is some obvious dependency between those expressions like the output of one function used as the input to another.
All that implicit parallelism is something that C compilers have to work hard to derive, but in Haskell it's explicit in the source.
Can we agree that people are learning a programming language that they call "C", and writing text files in that language and running programs that they call "C compilers" such as gcc on those text files to produce binary executables?
You are correct to note that the programming language I'm describing does not perfectly conform to any of the documents that are ever called a "C standard". But that programming language is central to our computing infrastructure, and whether you or C standard authors like it or not that programming language is widely called "C".
That language's semantics are unclear and incomplete---even though both gcc and clang are commonly called "C compilers", they have differing behavior on `f(i++, i++)` as you found; reportedly, there's code on which gcc behaves differently depending on the optimization level option.
I don't agree with this. Yes, people often write code that unintentionally rely on UB/implementation-specific behavior; but are there important projects that do this intentionally? In any case, compiler authors are going to follow the C standards, not you hypothetical "C-in-practice" standards, so what good does it do to for formalize or clearly specify the latter?
The Linux kernel.
Expecting integer overflow to be two's complement is so common that successors like Java, Go, and Rust all standardized on it, used by hashing algorithms and more: https://github.com/rust-lang/rfcs/pull/560
> In any case, compiler authors are going to follow the C standards, not you hypothetical "C-in-practice" standards
"The optimizer does go to some effort to "do the right thing" when it is obvious what the programmer meant (such as code that does "* (int *)P" when P is a pointer to float)." --blogpost by the second most widely used C compiler, clang/llvm: http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
UB is not a flaw. It is an efficient engineering method. Kids will be alright, that's hard thing to learn, but once you get it, it is simple, clear and complete. If kids want to stay in a sandbox, well, we have a lot of them.
Those who blame UB probably never turned their attention to electronics, physics or chemistry. That is a mess. Every time you pass wrong arguments or do something out of spec, it can explode and really hurt your foot.
Edit: typo
But if you're saying all of C's UB is good, and none if it is a flaw, I think you'll find yourself in the minority there. Do you consider integer overflow to be an "idiotic case that should never happen and poisons code simplicity"? Relying on it to be two's complement is so common that successors like Java and Go and Rust all standardized that.
If you'll concede that some of C's UB is relied upon in real, important, non-idiotic code, then it follows that it is indeed important for there to be efforts like Cerberus that develop unambiguous semantics for more of C than is currently covered by C standards.
As UB is an engineering method, we need to look at it from practical point of view ONLY. In my limited knowledge, platforms with really random overflow-then-back (x+n-m) do not exist. Note it has nothing to do with n's-complement, it is just a sad consequence of (x+n)==undefined, so compiler is free to assume (x+n <= INT_MAX) and throw out some "dead" code, which is not dead, ymmw. This specific UB is about what you get in (INT_MAX + positive) use cases, and these are very strange, because irl 127+3 is 130, not -126. It doesn't even help in famous m=(l+r)/2.
Though I can see social rationale behind standardization you mentioned. Some behaviors too suddenly appear undefined to newcomers who do not bother to read tfm or faq at least, and even those who program for 5-7 years still cannot explain what restrict keyword means. Maybe I'm wrong or somewhat elitist in it; today's software development is different. We have lots of programmers who hardly go further than learning "a new syntax" and we are doomed to use them as a productive force. Then these UB-clearing moves have perfect sense, that's why I mentioned sandbox languages.
I will note that "from [a] practical point of view ONLY", it is incredibly impractical for a compiler to gratuitously break real code that relies on behavior that deviates from spec; if that reliance is common, the practical thing to do is to support that behavior, which they do: "The optimizer does go to some effort to "do the right thing" when it is obvious what the programmer meant (such as code that does "* (int *)P" when P is a pointer to float)." --blogpost by the second most widely used C compiler, clang/llvm: http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
Which is why projects like Cerberus are valuable.
Actually, pushing function arguments right-to-left means that the first argument ends up on top of the stack - and that makes vararg-style functions possible.
Think about it - a called vararg method must be able find the fixed arguments to figure out how many arguments it actually gets.
About the potential order of evaluation, it's not limited to C, but at least it's clearly and dilligently explained even in tiny standards like Scheme.
I had always thought it was evaluated left->right (which it was in the first scheme implementation), but it is in fact standardised as "The order of evaluation is unspecified". The reason I hadn't noticed before was that most my code is rather functional. Now I was reading a char from a file handle, and peeking it as the other argument.
It failed because read-char is destructive (it changes the position of the file pointer), and since the order was the other way around in the scheme I was porting to, all the procedures relying on my "read-until" method started failing left and right.
It just took me ~20 years to figure it out...
It's not good style to exploit that, but you can.
The worst part is that I admitted to writing code that rely on order of evaluation on Hackernews.
To my defense it is _very_ old code that I converted from an even older turbo Pascal program I wrote as a teenager
$ gcc -Wall test.c
test.c: In function ‘main’:
test.c:12:13: warning: operation on ‘a’ may be undefined [-Wsequence-point]
test(a++, a++);
~^~
$ clang test.c
test.c:12:8: warning: multiple unsequenced modifications to 'a' [-Wunsequenced]
test(a++, a++);
^ ~~
1 warning generated. #include <stdio.h>
int postincrement(int *x) { int y = *x; *x = y + 1; return y; }
int main(int argc, char **argv) {
int i = 0;
printf("%d, %d\n", postincrement(&i), postincrement(&i));
return 0;
}
The point is that function arguments introduce sequence points, but the specification for i++ is weaker than what you get for a function call. C is full of gotchas like this...