C Questions and Answers
kukuruku.co
kukuruku.co
That people can teach themselves what to expect the compiler to do isn't all that surprising, and it also isn't surprising that a "modern" compiler does stuff an "old" C programmer might think is ridiculous.
I've engaged in this particular argument a few times only to throw up my hands in frustration over some exquisitely twisted line of reasoning that gave the compiler hacker person a fraction of a percent improvement[1] by exploiting this kind of situation. As long as I have a compiler flag that turns it all off its tolerable. But sheesh, sometimes I think these are C programmers who don't have the guts to become Rust programmers. You want to start fresh dudes, clean slate. Embrace it.
[1] "But Chuck, over the millions of machines out there its like an entire computer's worth of CPU cycles you can use for something else!"
Actually no. I think there was maybe one case where you could argue it was the optimizer that is producing different results than you'd expect for unoptimized output, but in most cases you have a problem that exists because of poor assumptions on the part of programmers/error prone language definitions.
I think the compiler behaviour is meant to be illustrative of the main point, not the main point itself, which I think is simply that the actual C language is more complicated than people appreciate. Most of the examples have nothing to do with optimization whatsoever; they're purely focused on correctness.
http://blog.metaobject.com/2014/04/cc-osmartass.html
...and attempts at turning C into something a bit less programmer-hostile:
https://news.ycombinator.com/item?id=8233484
My point of view is that compilers should be optimising at the level of machine instructions, not by attempting to second-guess the programmer and remove code that it thinks invokes UB. I've looked at tons of compiler output over the years, and there's plenty of opportunity for optimisation in instruction selection and register allocation... C should be a "do what I say, not what I mean" type of language.
[0] http://libcello.org/ [1] https://github.com/eatonphil/viola
My boss who is a C developer had only two answers wrong, he didn't remember you could tentatively declare a global var, and he was fooled by the comma operator inside the array index.
I suppose another perspective is that there are no shortage of programmers who are suffering from the Stockholm Syndrome, and having them make excuses for existing language's shortcomings are one reason it is harder to get critical mass behind less borked languages ;-)
In a similar vein to the original article, some may like to play along with: http://www.gowrikumar.com/c/index.php
...anyway, I don't see the problem exposing people to potential pitfalls. It's almost like people arguing that people shouldn't read "Expert C Programming: Deep C Secrets":
https://www.google.com/search?q=expert+c+programming+deep+c+...
...also I'll just throw this out here:
http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
But in fact it's quite a reasonable (the only reasonable?) approach to optimisation, given a function that might invoke undefined behaviour on certain arguments, to emit code that is optimised for work on arguments that don't.
That doesn't seem to me to be tortured logic. The compiler ought to make that optimisation, always. It might be perfectly clear to the programmer that the function in question can never be sent a null pointer, but only by reasoning about the program on a level the compiler can't. It's only a minor side benefit that this can allow a certain amount of reasoning about the code paths that might be taken when you DO invoke undefined behaviour. That usually won't be much use, but might help one identify the kind of error one has made.
I expect most of us here have puzzled over some confusing output from a C program and tried to work out, from the output, whether we made an allocation error or overflowed a buffer or were off-by-one on some bounds. It's a wonderful language in some ways, but the pitfalls are there. Which is the point the author is trying to make.
Even more insidious is the fact that memcpy of zero bytes from or to null is undefined behavior.
[1]: http://stackoverflow.com/questions/26906621/does-struct-name...
Rarely did the quizzes provide great insight. Rather, they confirmed the benefits of keeping your code idiomatic.
I don't disagree on the usefulness of teasing your brain with these things once in a while. However, I think the best way to ensure you don't hit bugs caused by such things is to avoid the situation altogether.
#define Sum(a,b) a+b
This of course fails in any context where the precedence of '+' doesn't match your intent e.g. Sum(1,5)7 becomes 1+57Do you remember to always define your expression macros with parentheses? Any time you didn't do that, you have a bug.
Yes - that's a necessary part of being idiomatic. Any macro that's defined without brackets sticks out like a sore thumb, and won't pass code review, even if it's (initially) used in a place where they wouldn't be necessary.
Oh, yes! No parentheses around elements of a macro definition is an obvious sound of trouble and looks very wrong on my retina.
I do agree, though, that if I somehow forgot to do that (tired? nervous?), that would be a bug that's difficult to spot. Point taken :). That's what code reviews are for, but there isn't always time or availability for one, sadly.
5 is also missing a check for a NULL pointer. ;)
Most of there rules are unfortunately ignored, as obscure information. The worst offender I see in wild code is 5.
The second most ignored is not checking values before computation: 10, 11, 12.
No.7 is very interesting, rarely violated, most programmers don't even know that is a thing or just assume the processor won't trap on an unaligned read.
Not necessarily. It could be up to the caller to ensure that the pointer isn't NULL.
These examples illustrate well-intended and useful features of C, not flaws. I will explain why for each in a comment below:
You mean Algol, Mesa, and few others from the same vintage or even older?
Or rather Macro Assemblers like MASM, TASM that provided higher level macros for structured programming?
Algol is higher level than C.
Sure, but it is also safer and already had helped implementing a few operating systems before C's authors could imagine coming up with C.
My understanding of history is compiled Algol wasn't nearly as fast as compiled C, which was needed for operating systems and performance-critical code.
Edit: Forgot to mention that up to the early 90's, C compilers generated pretty crappy code vs what any average Assembly coder could write. And was only relevant for those fortune to have UNIX at their company or university.
What do you think was the reason C took off while Algol use diminished? Was the growth of Unix a significant reason? Do you think C's adoption was misguided?
I think C's took off because C fit the sweet spot of ability to produce fast code while still being cross-platform and human-readable.
As those workstations became a success in the US market, its use spread outside US and the need to have developers that could write software for them increased. This meant knowing C.
All the other operating systems at the time didn't offer C compilers. The few that did, it was just another language to choose from, most of the time only a subset of K&R C.
This is how the distinction between libc and POSIX APIs came to be. The original libc is mostly what could be implemented in other OSs without depending directly from the UNIX API semantics.
If the likes of Sun and SGI hadn't succeeded, most probably C would be a footnote just like Algol.
I have been writing software since 1986 and 1992 was the first time I cared to learn C, just to quickly ditch it for C++ on the year thereafter.
#2. Treating dereferencing a NULL pointer as undefined behavior means that the compiler is not required to generate additional instructions such as asserts or crashes to guard against potentially dangerous side effects. C compiler assumes that the programmer is in control of his/her code. In this example, it can be assumed that a careful C programmer has already guaranteed that the pointer will not be null when it is dereferenced. This C feature is an optimization to avoid generating redundant or unnecessary asserts or handling code.
3. C allows the programmer to handing pointers, allowing for such low-level optimizations that may not be possible in higher-level languages. A careful C programmer may have taken steps outside the function to handle the situation where yp==zp, or may otherwise be unconcerned about a particular case, for performance reasons.
4. A correct implementation of IEEE 754.
5. Since C is designed to efficiently compile to any computer architecture, it needs to be aware of the distinction between the arithmetic width of an instruction set vs the width of addresses. Ints are optimized to default to the natural arithmetic width of an instruction set (so that compilation doesn't produce unnecessary packing/unpacking instructions whenever they are accessed) but is guaranteed to be atleast 16 bits wide. However, since data structure can be as large as addressable memory, it is necessary for size_t to be the width of addresses.
6. Allowing size_t to be unsigned allows all bits of a size_t variable to be utilized for expressing size.
7. Undefined behavior is, again, an optimization feature, allowing each compiler to implement as it sees fit.
8. Comma operator is useful when first operand has desirable side effects, such as compactly representing parallel assignment or side effects in for loops.
9. C allows unsigned integers to wrap around 0 and UINT_MAX. This feature can be utilized as an optimization, for example as a free (no additional instruction) deliberate modulus operation. This is usually how unsigned integers behave in assembly.
10 & 11 & 12. Some ISA's, like MIPS, treat overflow of signed numbers as an exception. Others simply treat the result as a valid two's-compliment value. Since C is machine independent, C's official specification for overlow of signed numbers must be compatible for all ISA's. Simply treating the result as undefined does the trick, and means the compiler doesn't have to make guarantees or version for each ISA.
One disadvantage of making int 64 bits, even if that's the natural size, is that if char is 8 bits, then short has to be either 16 or 32 bits (or 64) -- which means that you can't have predefined types covering all the common sizes (8, 16, 32, 64).
That's not quite true, since C99 introduced extended integer types -- but I don't know of any C compiler that has provided them.
(The intN_t and uintN_t types in <stdint.h> don't solve this; they still have to be defined in terms of existing types.)
6. size_t is required to be unsigned.
9. C requires wraparound behavior for unsigned integers.
10, 11, 12. It's not just the result of an overflowing signed integer arithmetic operation that's undefined, it's the behavior. `INT_MAX + 1` can yield `INT_MIN`, or it can yield 42, or it can crash your program and reformat your hard drive (at least in principle).
Author needs to stop his anti-intellectual everyone-is-as-ignorant-as-me bullshit. I see that a lot re programming to justify a lot of silly positions. If you only know Javascript, that's great, I rather like having shiny things in my browser (I like it too much, even). That doesn't mean that C programming is obsolete; some of us know C. For example, I got #5 and #9 wrong, and the rest I got right including the general idea of the justifications. (10/12 is pretty good for someone who grew up in the Java era, but I want to get better.)
What I do think he's implying, if not outright claiming, is that there are many people who grossly overestimate their knowledge of the language. That doesn't mean there aren't a lot of very talented and knowledgeable C programmers— but there are a lot of people out there who think they're hot shit because they've done all the exercises in K&R, but who have no real familiarity with the formal semantics of standard C.
He never said that or implied it.
The majority of them can be answered correctly by someone who understands computer architecture and programming languages. E.g. #3 is about pointers which do not necessarily have only to do with C. Even JS developers implicitly deal with pointers:
var v = {} // v is a pointerIn C I can perform arithmetic on a pointer.
`v` is a symbol which references an object. It might be acceptable to refer to `v` as a reference if we're being sloppy.
JavaScript has a DataView[1] which could be used to implement what I think is a pointer, but I do not think that most JS developers have used it.
[1]: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
> In computer science, a pointer is a programming language object, whose value refers to (or "points to") another value stored elsewhere in the computer memory using its address. A pointer references a location in memory, and obtaining the value stored at that location is known as dereferencing the pointer.[1]
We are not being sloppy, in fact we are being extremely correct and precise. Pointers and dereferencing pointers might be implicit and automatic in Javascript but that does not change that a pointer to a value is being used, instead of the value itself.
Pointer arithmetic is not the same thing as pointers, it is merely something you might be able to do with pointers if the language you are using supports it. Rust calls its "symbols that reference objects" (what on earth?) pointers, even though you are unable to do pointer arithmetic on them (unless you drop to unsafe). C++ Smart Pointers are called as such even though you can't do arithmetic on them.
[1]: http://en.wikipedia.org/wiki/Pointer_%28computer_programming...
> In computer science, a pointer is a programming language object, whose value refers to (or "points to") another value stored elsewhere in the computer memory using its address.
You have your answer here. In JavaScript ("programming language"), v ("object") has a value of JavaScript structure called object. Internal representation of v sure does use a pointer, but from JavaScript perspective, it surely is NOT a pointer. If you get the value of v in JavaScript, you don't get memory address ("value [that] refers to (or "points to") another value stored elsewhere in the computer memory using its address") - you get the object itself. That's what references do, not pointers.
I got 12/12 here. I would say I know C.
bool is_zero(int x) {
return x == -x;
}
bool is_zero(float x) {
return x == -x;
}
and ask them which is wrong and for what value. Most of the time the instinctive response is that it must be the float code (because floats are evil, duh).This works even in languages with defined overflow for integers.
int i = 10;
Q. Is this code correct?A. Yes.
2.
extern void bar(void);
void foo(int *x)
{
if(x == NULL)
{
return;
}
int y = *x;
bar();
return;
}
Q. It turns out if you check the validity of your variables before you use them it prevents you from having to understand undefined behavior.A. Is there a question here?
3. There was a function:
#define ZP_COUNT 10
void func_original(int *xp, int *yp, int *zp)
{
int i;
for(i = 0; i < ZP_COUNT; i++)
{
*zp++ = *xp + *yp;
}
}
I optimized it this way:...because nobody had any idea what it was doing and so it wasn't used anywhere.
4.
double f(double x)
{
assert(x != 0.);
return 1. / x;
}
Q. Is it possible for this function to return inf?A. If you're at the point where you're asking that question you should have been using a decimal library a long time ago.
int my_strlen(const char *x)
{
int res = 0;
while(*x)
{
res++;
x++;
}
return res;
}
Q: The provided above function should return the length of the null-terminated line. Find a bug.A. They didn't use `strlen()`.
if(condition); {
some stuff
}
Note the semicolon after the if-statement.The tests became games of "find the semicolon or = in place of ==". Ugh.
As someone on HN once said (in jest), it's easy to write bug-free C, you just need to never make a mistake ever and spend a million hours auditing it.
When "int i = 10;" is encountered, the tentative definition behaves effectively as a declaration. If the compiler were to reach the end of the translation unit and the variable i was never defined elsewhere, however, "int i;" serves as a definition.
EDIT: To answer my own question, the submitter is clearly the author based on his submission history. So yes, this was an act of self-censorship to try to hide the author's disgusting attitudes.
for(i = length - 1; i < SIZE_MAX ; i--)
but it is completely valid.
for(i = 0; i < length; i++)
.. do something with (length - i - 1) ..
I've been much happier since I started writing all loops as incrementing loops. for (i = 0; i <= length; i++)
.. boy I sure hope length is not a max value ..
It's not about the incrementing vs. decrementing so much as the equality bounded loop on a value that could be a min/max value. Gets you every time.The only people I would not want to work with are ignoramuses and ironically, smart-asses. So you.
Am I the only one who tried compiling :
#include<stdio.h> #include<stdlib.h> int main(int argc, char *argv[]) { int i; int i=10; printf("i=%d\n", i); return 0; }
and got the redeclaration error I expected?
int i;
int i=10;
#include<stdio.h>
int main(int argc, char *argv[]) {
printf("i=%d\n", i);
return 0;
}Something like:
/* file.h */
int global_i;
and you have: /* file.c */
int global_i = 0;
Then if file.c includes file.h both will be in the same file during compilation.Assuming of course the compiler doesn't already check for all of these...
He succeeded in convincing me that he does not know C!
>and how they correlate with multiple C standards
There's nothing going on here specific to any particular revision of ISO C (that I can see).
Did he make sure the behavior he mentioned is uniform across all ISO standards? What is his source for coming up with a certain answer to a certain piece of code? any of the standard documents? K&R book? gcc output?
As an example:
In the answer to point 2, he claims:
> ... the compiler thinks that x cannot be a null pointer ...
First of all, this gives a strong indication that he's analyzing a compiler output, a compiler that he didn't reveal in the article.
But even if we ignore that, and he truly is going by a rule that is uniform across all of K&R, C89, C99, and so on, could you or him point me to any page in the C99 standard document where it is explicitly stated along the lines that the compiler "should assume" the pointer to be not NULL after an undefined dereferencing in line (1) and hence ignore (2) and (3)? (Based on my experience, I have a very strong hunch that a standard would not enforce assumptions as a result of an undefined operation.)
If you could, you/he "may" have a point ("may", because I still have 11 other points to critique). If not, he and you clearly have no idea what you guys are talking about!
I don't know.
>What is his source for coming up with a certain answer to a certain piece of code?
His conclusions are consistent with my understanding of the C standard. His justifications for his conclusions refer to specific rules regarding program behaviour, which suggests he's using the standard(s).
>First of all, this gives a strong indication that he's analyzing a compiler output, a compiler that he didn't reveal in the article.
He is using the output of some compiler to illustrate the potential consequences of making the given mistake. It doesn't really matter which; the point is that it is not legal to perform lvalue-to-rvalue conversion on the result of indirecting through a null pointer.
>could you or him point me to any page in the C99 standard document where it is explicitly stated along the lines that the compiler "should assume" the pointer to be not NULL after an undefined dereferencing in line (1) and hence ignore (2) and (3)?
No, because the C99 standard document does not say that. What it does say is effectively that the compiler MAY assume the pointer not to be null; more generally, the compiler is allowed to assume that the program never exhibits undefined behaviour.
and therefore his claim:
> ... Turns out, bar() is invoked even when x is the null pointer ...
is incorrect (bar may or may not be invoked) and misleading for C newbies, without any mention of the subtleties of standards and implementations. Hence my original comment.
In this case, the code is invalid because it invokes undefined behavior, and the compiler is allowed to do literally anything. The author of a portable C program is not allowed to rely on any particular behavior. I don't have any of the ISO C standards docs handy, but I'm quite certain that they all agree here. The (probably hypothetical) compiler in use here is apparently trying to apply several heuristics that are useful in other situations but fail here because there is no right answer.
In general, the C standards avoid requiring a specific behavior where choosing to require a specific behavior could hurt portability. Most C compilers make use of their leeway to interpret things in a manner that helps make bad code run and many compilers make promises beyond those required by the standard to aid in writing non-portable code.
and yet he makes the claim that "bar() is invoked" which is not only incorrect (bar may or may not be invoked), but also misleading for C newbies who are actually trying to learn something by reading this article. Hence my original comment.
In fact the compiler can't make any assumptions about a pointer whose value has been hard coded.
uint8_t b = (uint8_t) 0x30; // 0x30 is the address of PORT D on an AVR Atmega8
You won't find a line that states it explicitly, but the standard does allow it.
If the pointer is null, de-referencing it invokes undefined behavior, so the program is allowed to do literally anything. That includes doing whatever it would have done if the pointer wasn't null. So the compiler is allowed to assume that the pointer is not null.
As I replied to others, saying "bar() is invoked" is not only not obvious, it's incorrect. bar() may or may not be invoked.
Odd of you to use 'mythical' when one of the most common compilers in existence does this.
% cc -std=c11 -o dingens dingens.c dingens.c:6:9: error: redefinition of 'i' int i = 10; ^ dingens.c:5:9: note: previous definition is here int i; ^ 1 error generated. % cat dingens.c
#include <stdio.h>
int main() { int i; int i = 10;
printf("hello, world\n");
return 0;
}%
From his webpage: "Reminding you that it’s a separate source file and not a part of the function body or compound statement"