So you think you know C? (2016)
wordsandbuttons.online
wordsandbuttons.online
But knowing every engineering detail is not the same thing as knowing how to program in C effectively. It's like being the engineer who designs a Grand Prix car. It does not mean you can drive it faster around the track than anyone else. Not even close.
For example, the C preprocessor is surprisingly complicated. I had to scrap it and rewrite it completely 3 times. If you try to make use of all those oddities, my advice is don't waste your time. Over time I removed all the C preprocessor tricks from my own code and just wrote ordinary C in its place. Much better.
You can limit yourself by choice to certain areas of the language that you know inside out (you know the asm they produce, etc.) and use it to solve problems. Don't worry about every single corner case. As you said, you can easily forget those things, especially if it's not your day-to-day job.
The same idea and principle can be found in "JavaScript - The Good Parts".
Sometimes the technology is objectively the wrong choice i.e. compiling javascript to native code (Not a JIT).
No, I won't name names.
It's not finished but let's just say I've managed to make yours and Andrei's thoughtful design into a monster
What counts is how easy to use in practice. You can get along just fine in C++ without knowing the exact aliasing rules from C or or how to specialize a template. What matters is that the extra features of C++ makes actual programming simpler, not harder. (for example destructor (RAII), standard library, classes, ...)
Nope, you can't. These are exactly the things that introduce undefined behavior (i.e. total breakage) if you aren't very careful about what you're doing at all times. Don't take my word for it, check out what the C++ designers themselves state about the issue in the C++ Core Guidelines. C/C++ is far from simple, and thinking that you can just make things up as you go along is a serious mistake.
So despite C++ being more complex than C, if you limit yourself to some practical subset, it is actually easier than C.
And it doesn't work in practice.
We use both C and C++ embedded as well as C++/Qt for the desktop control system, and we have no issue keeping to a sane subset of C++.
There are clearly defined rules and code review does the remaining enforcement. And it's not even hard or time consuming as everyone is pretty much aligned and every small issue can be easily resolved with a quick chat.
(In fact I try to avoid writing C or C++ whenever possible these days; undefined behavior in the language is too pernicious and unfixable without breaking compatibility. I think both languages are approaching obsolescence.)
But yes, Rust, or even in userspace, as newer and/or more microkernel-ish OS's allow for. Apple is doing work to allow drivers to be written in Swift...
Some device classes (not to mix with OOP ones) can only be programmed with C++, while others can be developed in any compiled language able to link to the OS APIs.
There is a WWDC session on it.
Rust does not have these same features to date.
Astrobe has been selling development kits for ARM based boards for years now.
The features required are still unstable (as in not in the stable version of the compiler), but it’s getting there, and doing it fast.
In any case, firmware behaves a lot like banking software, language change will take time
Also, most toolchains of the sector are way more complex to maintain and way more bloated.
Microsoft says the future is a mix of constrained C++ (Core Guidelines), Rust and AOT C#.
Google says the future is C++ and Java (as of Treble) on Android, with Go, C++ and Rust on ChromeOS and Fuchsia.
ARM says C++ on mbed.
GenodeOS says C++ and Ada.
Newton, Symbian, Bada and BeOS used C++.
C is married with UNIX, they were born to each other, other OSes have long followed other paths.
Apparently they are also adding support for desugaring Java 10 language features (yep 10 not 12) that don't rely on new JVM bytecodes (as per Google IO talk about state of Android tooling).
At least for now.
One advantage to being an older programmer is I don't feel any need to show off any more. I try to make it so obvious that anyone would look at it and think that's so simple, anyone could do it.
It's surprisingly hard to write simple code. Any idjit can come up with Rube Goldberg code.
I don’t think I’ve ever written one correctly on first try.
Because this is crazy to me! Some languages only use for loops (Go).
Or maybe I’m misunderstanding the use case?
Functional graph programming is also still a bit of an open problem. There are awkward scenarios in both cases.
Not sure how you deal with that with for loops either. Increment the iteration var in the body of the loop? (Seems scary to me, but like I said, I’ve got terrible intuition with them)
var ie1 = foo.GetEnumerator();
var ie2 = bar.GetEnumerator();
while(true)
{
var has1 = ie1.MoveNext();
var has2 = ie2.MoveNext();
if (!has1 && !has2)
break;
if (has1)
// do something with ie1.Current
if (has2)
// do something with ie2.Current
}Please don't take offense, but this is odd to me. I honestly severely doubt I am some sort of programming super genius, but I have never had any issues setting looping logic correctly. (I took four years to teach myself programming & CS and now I've been at my first professional dev job for ~6 months.) None of my colleagues seem to have such issues either. What are you experiencing trouble with most? Off-by-one?
But yeah, syntax (what order the arguments go in) and off-by-one issues are the majority, I think. Plus figuring out what my initial accumulator needs to be.
Idk, map/filter and friends are just a much more direct mapping of how I think of programming.
It's usually much easier to take extra time to make sure it's right than to debug it later.
#include <stdbool.h>
#include <stdio.h>
typedef long T;
bool find(T *array, size_t dim, T t) {
int i;
for (i = 0; i <= dim; i++);
{
int v = array[i];
if (v == t)
return true;
}
}
There are 5 errors in that example.Are you thinking more of JS or another functional scripting language than C/C++? C doesn’t have map, so it’s not an option, and in C++ it’s called something else.
In JS, I so wish that I could switch to functional constructs like map permanently, but map and foreach are much slower than loops, an order of magnitude or more for tight loops. I’m still forced to use loops in performance critical code, even if I consider map a better choice.
Coming back to C++ after having been in JS land for 5 years, C++ feels constantly difficult to use, and all the names for the functional primitives don’t seem to make intuitive sense like they do in JS.
I just learned this about JS's map/foreach a few weeks ago. I didn't think it was possible to be more disappointed by JS than I already was, but somehow I managed it.
C++ only when Java or .NET need it's assistance, or integration with OS APIs that require C++ (NDK, WinUI, DX).
Then C only when there is no alternative (customer wants it, we only do C here, required lib is C only e.g. SDL, ...).
It is an herculean project, but maybe some day LLVM could be rewritten into something else. After all it isn't the first compiler stack, just the one that got most famous.
I don't get it. If my macros produce standards-compliant code and they make my code easier to read and understand, why shouldn't I use them?
My goal as a developer is to write clean performant bug-free code... not to make life easy for compiler developers.
The problem is they don't make code easier to read and understand. Worse, the unhygienic nature of C macros makes it hard to contain them.
I haven't seen your code, so I'm speaking based on what I've seen of mine and others' code. If you dial it back, the person who has to deal with your code after you leave will appreciate it.
More generally speaking, if you're doing metaprogramming with the C macro system, you've outgrown the language and should consider a more powerful one.
It worked something like: DEBUG(msg); Except, it was defined in such a way that you actually had to have two closing parens, like DEBUG(msg)); It looked syntactically invalid, but whatever the macro did required it.
The entire code base was littered with WTFs like that...
debug printf("I got here\n");
and the printf only gets compiled in when compiling with -debug. (Any statement can be used after the printf.) Even better, semantic checks for debug statements are relaxed - for example, purity is not checked for them.Meaning you can embed debug printf's in functions marked 'pure', instead of having to use a monad.
For me writing embedded code is the ultimate test for your programming skills, since a lot of C toolchains for embedded devices (as in bare metal embedded) are unstable, only almost compliant and full of weird behaviors and hacks.
As long as you invoke use it properly, e.g.
ELASTICARRAY_DECL(PTRLIST, ptrlist, void *);
there's no way it will create non-compliant code.
Exaggerating here, but rule of thumb is that it takes twice to debug as it takes to write it, so how long do you want to take to do maintenance fixes?
Ok, so... macros which let me write code faster should help with debugging time as well? ;-)
So it depends how complex they are.
--Brian Kernighan
tia
But if you want a competitive compiler, better pencil out 10 years.
after taking the test, wouldn't this be:
at one point I knew everything there was to know about (my implementation of) C
also: one time long ago I tried to use the c-preprocessor to preprocess a data file. ha ha ha ha ha. (conclusion: don't do that)
Say for example, in question 5, the statement "return i++ + ++i;" is undefined, because the value of i is read and modified twice in a single sequence point (and of course, the order of addition is unspecified), which is not allowed in C. So the answer is "Undefined." (The explanation given in the page not accurate enough in terms of C)
And for question 1, the code is valid, but the result is not strictly defined. It depends on the implementation. So the answer is "Implementation defined."
The usage of "main()" hurts me, the strictly conforming way is to write it as "int main(void)" (or similar)
I feel like the questionnaire piss off people who really knows C.
The code is undefined, but "I don't know." is still the correct answer for what happens to the variable.
Well, then a better choice would be "I can't know."
Ans: A compiler.
2) Probably that's not needed? A normal optimizing compiler just inlines the function somewhere new — and boom? Then again, can that really happen with practical contemporary compilers and this exact statement?
I don't know what the chances are that such a thing could ever translate into good assembly being emitted in one run and bad assembly being emitted in the next.
This pretty much is the definition of "undefined behaviour" in the context of a standardized language specification.
That's only when using Boehm GC[0] in kernel device drivers that self-modify. Or any MSVC binary.
:-D
The original problem definition, as specified by @pksadiq, read thusly:
> Say for example, in question 5, the statement "return i++ + ++i;" is undefined ...
This inspired a response by @Filligree of:
> Compile it, look at the assembly. You can know. The answer will vary from place to place, but it isn't non-existent.
Given the original constraint of an undefined statement result, and the suggested activity to address same, I posited that the recommended action is an exemplar of observing the product of undefined behaviour.
You then contributed:
> You mean “unspecified behavior”.
As per c-faq.com[0], there are three categories identified relating to this topic:
1 - implementation-defined: The implementation must pick some behavior; it may not fail to compile the program.
2 - unspecified: Like implementation-defined, except that the choice need not be documented.
3 - undefined: Anything at all can happen; the Standard imposes no requirements.
Whereas you imply a standards-conformant implementation of "return i++ + ++i;" is unspecified (category #2), it is, in fact, undefined (category #3). The support for this assertion is as follows.
As per the same site, Question 3.8[1] includes:
> Between the previous and next sequence point an object shall have its stored value modified at most once by the evaluation of an expression. Furthermore, the prior value shall be accessed only to determine the value to be stored.
And further states:
> ... if an object is written to within a full expression, any and all accesses to it within the same expression must be directly involved in the computation of the value to be written. This rule effectively constrains legal expressions to those in which the accesses demonstrably precede the modification.
And concludes with an example stating:
> ... the Standard declares that it is undefined, and that portable programs simply must not use such constructs.
Therefore, the original expression presented by @pksadiq is in fact an exemplar of an undefined expression as defined by category #3 shown above. Since both it and the message to which I originally responded satisfy same, I stand by my response given to @Filligree as having had informally defined the standard C concept of "undefined behaviour."
You're misreading things. amluto's assertion is that "The answer will vary from place to place, but it isn't non-existent." is a description of category #2. That assertion is basically correct, depending on how exactly you define "place to place".
An informal definition of category #3 is "The answer can vary from place to place, or not exist at all." ideally followed by "It might crash or run unrelated code or even prevent the preceding code from running." It's flat-out wrong to say a value "isn't non-existent" when it comes to source code exhibiting undefined behavior.
"I dont know" was absolutely the correct answer.
In this case, a non-C programmer should answer "I don't know" to all of them. A person with a passing familiarity should answer similarly. A seasoned pro would be forced to answer the same. Making it a rather useless tool for distinguishing people who think they know C but are honest when faced with their limitations or those who truly know it and know the answer is undetermined, which is supposed to be the point of the exercise.
I don't think that assumption is justified. Someone could say the point of the exercise is to illustrate that C is confusing.
Never once does the author mention that C is confusing, use the word confusing, or otherwise indicate that general idea. If you're getting that impression, it's your own reading into it. I'm not even saying you'd be incorrect, but that's not the author's intent, which was the basis of my comment.
Only through extreme levels of compiler code inspection, as it can vary based on optimization heuristics.
> Quite the contrary, they indicate their intent is to demonstrate that certain segments of people who believe they know C don't in fact understand its intricacies.
Demonstrating that people don't know C is subtly different from an intent of testing whether people know C. The point being made is about C itself.
>"So you think you know C?"
That goes along with the interpretation that the point is to illustrate C is confusing. It would go along with something like "You think you know it, you think it's simple, well actually you don't know it, it's confusing."
>intended identify to test takers whether or not they really understand the intricacies of C, and to think critically about that source of their knowledge
Yes, its intent is to indicate to test takers that a lot of them don't really understand the intricacies of C, which demonstrates that C is more confusing than they originally thought.
>Never once does the author mention that C is confusing, use the word confusing, or otherwise indicate that general idea.
Here are some quotes that indicate the idea that C is confusing:
>C is not that simple.
>It’s only reasonable that the type of short int and an expression with the largest integer being short int would be the same. But the reasonable doesn’t mean right for C.
>Actually, it’s much more complicated than that. Take a peek at the standard, you’ll enjoy it.
>The third one is all about dark corners.
>The test is clearly provocative and may even be a little offensive.
Then the author says that he did C for 15 years and thought he knew it, but then realized he didn't. That indicates to me either that the author is saying that he's not smart, or that C is confusing. The second appears to be the point the author is actually making.
My interpretation is also based on what the author states, fairly explicitly. And I don't think there's anything that explicitly contradicts my interpretation.
There is a world of difference between "don't know" and "can't know", as the first implies a shortcoming on the side of the developer while second one states that the question is patently meaningless to someone who does master the language.
Which is both not true (because the compiler manual usually won't define undefined behaviour) and irrelevant (because the questions were about C, not about a compiler).
You're being pedantic about something silly, but you're also wrong in your pedantry.
Actually, as undefined behavior should not be used at all then the correct answer should be "nevermimd these examples, they are all bug-ridden".
This isn’t a recent standardization; it’s been an explicitly specified feature of C pretty much since the very beginning of the language. See page 7 of the prehistoric https://www.bell-labs.com/usr/dmr/www/cman.pdf
In your comment you are jumping from "I don't know", which is the first step, to wanting to explain why.
This is an article about the C language and the starting point with all the examples is that you do not know for a fact what the result will be.
Using "I don't know" as a substitute for "I know that the standard clearly covers this, and it says that the result depends on the implementation" does seem to be designed to piss off people who know C. If they really wanted to get the point across that you don't know what is and isn't implementation defined or undefined, they shouldn't be using vague questions to mislead people; they should just plainly ask questions which people don't know the answer to.
I hate this kind of questioning where you 100% know the subject matter the quiz is asking about, but the question and possible choices is so vague you have to try to interpret what you suspect the person who wrote the quiz wants the answer to be. I once had an exam which was full of that kind of multiple choice question, and guessed the exam author's intentions wrong on most of them.
My point isn't that there's no logically consistent answer. My point is that there are _two_ logically consistent answers, and which one is correct depends on the unknowable state of mind of the quiz author. On the other hand, if the options included "It's implementation defined" or "it's undefined", the author would have made their expectations clear, and the quiz would actually test people's knowledge of C rather than people's ability to try to reason about what sort of answer the author expects.
I hated this test. I’ve spent 12 years working on C targeting various flavors of arm and x86.
Just because the behavior is undefined when compiled without warnings and run on a Soviet water integrator doesn’t mean the language is undefined for the 99.995% of the industry uses.
Behavior of c89 or later with -Wall -Werror on modern clang, gcc, icc, visual studio, is well understood on arm, x86, mips, risc, ppc, Cortex-m and just about every other hardware architecture.
But, C is a pia, and I’ve been using rust instead :)
What rubs me about these sorts of articles is they make some presumption about the importance and nessecisity of writing truely portable C, as if the "C Standard" were in and of itself a terribly useful tool. This is in contrast to where I live most of the time which is "GCC as an assembler macro language" (for a popular exposition on this subject see https://raphlinus.github.io/programming/rust/2018/08/17/unde...). And yeah, reading through the problem set I was critiquing it in context of my shop's standards, where we might be packing and padding, using cacheline alignment, static assertions about sizeof things, specific integer types, etc. So these sorts of articles just come off as a little pendantic to folks like me. I don't doubt they're useful for some folks, and I guess it's interesting to come up from the depths of non-standard GNU extensions and march= flags to see what I take for granted.
So, today, using the compiler installed on your system right now, sizeof(int) = 32. Great. That means nothing, and changes nothing about whether your code is correct. You should not write code relying on it. Just like you should not measure the output of the questions on this test, and declare that you know what the answers are.
While I feel the tone of your comparison was intended to be a bit hyberbolic, the reality is a bulk of modern C development occurs in a context similar to the one you describe. Further the thought, utterly foreign to the vast majority of software developers, that the physical machine may not be some utterly abstract and constantly mutating target which there is no hope of understanding is, imo, one of the great dying arts of software engineering - a death perpetuated by the same sort of folks who think CS education should be carried on in Java.
I contend that, these days, most C is written to target a particular compiler, physical machine, and/or device.
My first job was a 4GL targeting customers running DOS on the 80286, complete with runtime linking. 100% of that work has been abandoned due to incompatibility. It contributed nothing to the profession beyond what I personally learned.
The author said he never did a full scale rewrite. He slowly migrated code from one platform to the next.
Today, Apple’s code runs on both ARM and x86 and with Marzipan, as will developers code. True most will be in Objective C, but some low level code is still in C.
Not necessarily, I took it to mean that engineering is holistic and things like compiler behavior in the face of undefined parts of the standard are important to account for.
"So standards are not some kind of holy book that has to be revered. Standards too need to be questioned."
The way I see it, a lot of compiler writers are basically taking the standard as gospel and ignoring everything else "because the standard doesn't say we can't" --- and that's a huge problem, because behaviour that the standard doesn't define often has a far more common-sense meaning that programmers expect. IMHO the onus should really be on the authors of compilers to find that reasonable meaning. In fact, the standard even suggests that one possible undefined behaviour is something like "behave in a manner characteristic of the environment" (can't remember nor be bothered looking up the standard.)
I would be rather disappointed if they didn't, honestly.
1) The standard says I must do this, so I must do it.
2) The standard doesn't say I must not do this (but does allow me to either do it or not do it), so it's totally OK if I do it.
I think you're thinking of cases covered by statement 1, and I think pretty much everyone agrees that compiler writers should behave that way for the standard to mean anything.
The issues arise in cases covered by statement 2. Just because the standard allows a behavior doesn't mean that the behavior is a good one. And yes, code relying on you not having the behavior is not following the standard, and that's something the authors of that code should consider addressing. But on the other hand, the standard may allow a lot of behaviors that only make sense in some situations but not others (totally true of the C standard, depending on the underlying hardware) and as a compiler writer you should think carefully about what behaviors you actually want to implement.
AS a concrete example, you _could_ write a C compiler targeting x86-64 which has sizeof(uint64_t) == 1, sizeof(unsigned int) == 1, sizeof(unsigned long) == 2, and sizeof(unsigned long long) == 2 (so 64-bit char, 64-bit short, 64-bit int, 128-bit long, 128-bit long long). Would this be a good idea? Probably not, unless you are trying to use it as a way to test for bugs in code that you will want to run on an architecture where those sizes would actually make sense...
GCC and Clang do give you the option to avoid optimizations based on undefined behavior: compile at -O0. We think of the low-level nature of C as being good for optimization, but in many cases the C language as people expect it to work is at odds with fast code.
It's fascinating to actually dive into the specific instances of undefined behavior exploitation that get the most complaints. In each such case, there is virtually always a good reason for it. For example, treating signed overflow of integers as UB is important to avoid polluting perfectly ordinary loops with movsx instructions everywhere on x86-64. It's easy to see why compiler developers added these optimizations: someone filed a bug saying "hey, why is my loop full of movsx", and the developers fixed the problem.
Edit: Should be movsx instead of movzx, sorry.
"fixed" by breaking other expectations. Regardless of what the spec says, that's still a stupid way to do things. There's a child comment below which examines this case in detail; and the real solution is to make the analysis better, not use UB as a catch-all excuse.
One immediate redflag I have noticed is using "int", "char", "short" as if they have a definite size. They don't. C standard only guarantees a minimum size. For example, many PDPs are 36-bit. Assuming the size of a variable is a common practice nowadays, but at least one should use uint8_t, int32_t, etc. from stdint.h.
But I was still tricked, it should be obvious in hindsight, 12 years of schooling led me to think: If the author was asking these questions, at least one or two questions must be answerable (even if it's technically incorrect, but you'd better to guess the original intention of a question). So I still tried to guess and got two wrong answers... Get to be careful next time...
This test.. it reverses that entirely.
I knew "int" was sort of platform-dependent (was 16-bits generally when I was learning to code, later 32-bit became more typical) — so combined with that niggle and all the "I don't knows", I (correctly) reevaluated by first couple of answers.
Still, didn't realize the last one was compiler-dependent.
https://news.ycombinator.com/item?id=12900279
Similarly titled in 2012: https://news.ycombinator.com/item?id=4657317
I've worked on 16 bit C code, 32 and now 64 bit code. So I knew that the behavior was implementation and optimization dependent. :)
Ignorance is bliss in C.
In the end this test proved to be a really valuable, because the "I don't know" drove the point home, specially for smart folks who don't like to answer any test, ever, with "I don't know".
In the C89 days, you'd use 'short' in aggregates (structs and array) for values you knew wouldn't exceed 16 bits so didn't want to potentially waste space; 'long' in situations where you knew 16 bits wouldn't be enough; and 'int' the rest of the time (where 16 bits was enough, and there weren't any storage benefits to outweigh the performance benefit of using the native word size).
Anyhow, for a computer with 16 bit wide data bus, having 16 bit ints might be justified by performance (and/or reducing memory usage.)
Ada, maybe? I don't know enough about it to comment. You definitely don't want to use Haskell for that sort of work load, though, at least not directly. Laziness-by-default is precisely the sort of hard-to-reason-about logic you don't want in that sort of application.
That said, if I I had no alternative but to try and tackle this problem, I would seriously consider a strategy where I would write a Haskell program that would generate the actual program (potentially in ASM directly) for me.
In these situations, you likely know your hardware and know your compiler, so you can actually provide an answer for 4 of the questions. The last one is a situation where someone should tell you not to get cute in the code review.
I wrote C in telecom and finance and in both places we enforced a rule: when you define a structure, put a comment after each element that says what you think the structure offset should be, and at the end of the structure #define a constant that says what you think the size of the structure should be. In a code review, if anyone noticed something that didn't look right, you could talk about it. In testing, you could also check that sizeof(foo_s) == FOO_S_SIZE and fail if it wasn't.
In some of our code, we would test the size of various types and structures on startup and immediately exit if they weren't what we expected. We'd print type sizes to logs to help debugging if there was ever a problem. We were supporting a single code base that ran on big endian, little endian, X86, Itanium, SPARC, ARM. Compilers change, but automated tests of type and structure sizes catch things immediately.
It may sound like a lot of work, but it actually isn't at all. It also helps a lot with long-term maintainability.
This is one of the things that C++ has actually improved a lot recently: doing this with static_assert is much nicer in terms of catching problems early... And yes, it's great for long-term maintainability.
But that reminds me of the joke: there are 10 kinds of people: those who know binary, those who don't, and those who didn't know the '10' was written in base 3.
Reminds me of the dumb exams some teachers would set to trick you when in school to make themselves feel superior.
I appreciate the apology here, and I can totally understand the concern about the spec in a safety critical environment.
Still, all questions on this test except the first are clearly examples of things you should never ever do in production code, which might undermine the message a bit? Yes, you can write bad code, and that’s true in every language I’ve ever used.
I’m guessing it would be hard to find a modern compiler on Windows, Mac or Linux that produced padding other than rounding up to nearest 4 bytes?
Sizeof(a+b) is obviously a weird thing to do.
char a = ‘ ‘ * 13 produces an overflow warning in gcc.
(((((i >= i) << i) >> i) <= i)) I hope nobody really did that.
return i++ + ++i; Not doing exactly this was drilled into us in CS 101. Still, I’d be interested to hear about a compiler that doesn’t return 1, since many people rely on the fact that ++i is pre-increment and i++ is post-increment. I don’t doubt one is out there, I’m curious to know which.
There probably weren’t much better choices 20 years ago... what would be the best choices today for a branch new nuclear power plant?
gcc (Ubuntu 8.3.0-6ubuntu1~18.10.1) 8.3.0 returns 2 (executes from left to right... at least the first time I ran it).
If you want to see some real tricky C code, see some of the articles here: https://locklessinc.com/articles/
uint32_t foo = 1; foo <<= 32;
Here, I figured that foo would always be 0. Wrong. It was always 0 with GCC, but this is undefined in the spec and code like this can have a different value in clang. I actually had to make a security update to my little open source project because of this (although the code I wrote did not manifest the bug in an insecure way, even with clang).
Indeed, testing your assumptions, because even a defined standard may result in a differing implementation of it. Especially in critical applications, testing the expectations gives some sense of a defined behavior.
This quizz is amusing as a mental exercise and a parable, but in reality all of these cases had to be fleshed out on a real platform, with real compiler and ... specified expectations of the behavior.
None of the cases in fact communicate a clear intent, except maybe #1 to figure out the padded size, still it's somewhat open-ended. Perhaps returning a specific condition (return sizeof(struct ...)==5; ) would show a clear intent. Not that it would change the right answer, just such a case may indeed be true on a specific platform, compile flags erc.
Result:
1:8 2:0 3:-96 4:1 5:2
Sure, there's a bunch of warnings emitted... fortunately.
And time_t[0] is not one.
;-)
0 - https://pubs.opengroup.org/onlinepubs/9699919799/functions/t...
But often we encounter UB in code that's already shipped. So it's good to have an intuition about what machine code was actually emitted, for example when deciding if a crash report is due to this particular UB, or not.
int a=1, b=2, i, j;
i = a += 2, a + b;
j = (a += 2, a + b);
Whats the value of a, b, i, j? Hint: i and j are different.Which begs the question, why does C have those features in the first place? The only C code where it's reasonably common seems to be in crypto algorithms.
Anyway, I am inclined to agree that these are misfeatures. I’d almost certainly ask for it to be changed in code review.
And then you find "it's a trap".
I think the questionnaire is not honest enought, a better answer for D should have been: "we need more information" or "there are programming inconsistences"...
I thought that when selecting "I don't know" I was telling "I don't know what's happening with the code and the inner datails".
But the thing that would be better to do, in my opinion, may be like having LLVM with macros (including standard macros for dealing with differences of systems, and user-defined macros for your own use).
I know that most of this stuff is undefined according to the spec. But I (might) also know what my particular GCC version does in these cases.
What I know for sure, that those programs do output something, and it's not "I don't know" (the string) ; )
What I don't know, is what level of sophistication the author assumes.
C, for all it's simplicity, is a relatively complex language.
And for the fifth (return i++ + ++i;) it warns "Multiple unsequenced modifications to 'i'".
I would not skip those warnings.
I know why we still use C, but the use of C is inherently prone to security problems.
C does not provide bounds checking by default, so it can be forgotten (Heartbleed) and the lack of either static checking, RAII or garbage collection (Not as a library e.g. Boehm) makes memory corruption all but inevitable.
In conclusion, the power given to Programmers by C far outweighs any of its perceived downsides in real-world scenarios.
Any good alternative still allows you manipulate raw memory, but provide a safe alternative which makes it much harder to fuck up.
What power do I actually lose by using a safer language?
Now coming to your other point, in today's environment, it is true that you do not lose much for the most part when using a safer language because somebody else has done the dirty work in the implementation of the corresponding language's runtimes, compilers, libraries and ABIs. Without the latter you cannot have the former. After all at some point you have to move out of the cocoon provided by the language and meet real hardware (a good example is bare-metal programming on MCUs). And that is where C is needed and any challengers have to provide exactly similar "ugly, dangerous and unsafe" features if they want to dethrone the champ.
To be honest if was pretty obvious what the right answer was, but tried to answer honestly anyway ;)
I do not like this attitude (then again it was just one random dude).
You might want to listen. You’re getting the K&R comment and the downvotes because this does not work, ever. It’s a really, really bad idea. In recursive calls, it might not crash right away, but you will have bad data, the memory at the pointer address will have been overwritten by the next stack frame that’s placed there.
Don’t ever return pointers to local memory because the memory is “gone” and unsafe to use the moment your function returns. Even if you try it and think it works, it can and probably will crash or run incorrectly in any other scenario - different person, different computer, different compiler, different day...
Your comments about getting a warning and ‘However if you wrap the local’s address... it “works”’ should be clues. The warning is the compiler telling you not to do it. The workaround doesn’t work, it only compiles. By using aliasing, you’re only tricking the compiler into not warning you, but the warning is there for a reason.
Signal handlers allow C programs to respond to events outside of the normal control flow (see signal.h, etc.). This means that once a function, say fnc1, has returned, the memory on the stack that was used by fnc1 can end up being reused at any point in time. A signal, perhaps generated completely asynchronous to the program itself by a different process, causes a stack frame to be allocated (possibly on top of fnc1’s old stack frame) for use by the corresponding signal handler. This could happen at any time, even before fnc1’s caller gets a chance to use the pointer returned by fnc1.
I would have preferred to be told:
- yes and no. You'll get warnings if you try to return a pointer to a local, however, doing this and that, you can manage to do it.
- but once you have achieved that, the result will be dependent on the way the stack is handled (not really in your control). You'll feel some comfort doing this in recursive calls, however beware of signal.h.
But this isn't the answer I received. I guess C programmers do not know the difference between what you can do (however risky) and what you shouldn't do. Also when someone asks such "weird" questions, do not assume he's a beginner with no notion of what constructs he can handle safely, maybe he's someone trying to find the limits of C – and once these limits are identified it can be a good conversation starter about C's internal and the way various compilers differ.
Edit: also downvotes on HN are not like downvotes on Reddit: there's actually a limit (-2 ?). Below this the comment disappears. Conclusion: only downvote when the comment engages in antisocial behavior (not respecting the rules or common human decency, etc ...), not when you disagree with it. I always upvote an unfairly downvoted comment for these reasons.
You are mistaking some luck in having it not crash once for thinking that it’s okay in some situations. It’s not okay under any circumstances. That’s what makes this even more dangerous. Your program could crash at any time. It might run a thousand times and then suddenly start crashing. It might always run for you, and then crash on other people. But just because it runs once without crashing doesn’t mean it’s working.
A signal is not the only way your function’s stack can get stomped on the very next instruction after you return. Other processes and other threads can do it, the memory system can relocate your program or another one into your previous memory space. Recursive calls are guaranteed to stomp on previous stack frames when your recursion depth decreases and then increases, the previous stack will be overwritten.
Returning a pointer to a local stack frame is always incorrect. It’s not risky, it’s wrong.
BTW: you have the ability to see comments below the downvote limit, go to your profile settings and turn on showdead.
I didn’t downvote you, if that’s why you were trying to explain voting behavior to me, but you will find on HN that downvotes and upvotes both happen for a wide variety of reasons, and are not limited to either whether people agree, nor whether the comments are polite. Downvotes are often cast for comments that break site guidelines, for example just failing to assume good faith can get you downvoted. So can making blanket generalizations about a group of people, like the above “I guess C programmers do not know the difference...”. See the comments section here: https://news.ycombinator.com/newsguidelines.html
I sometimes upvote what appear to be unfairly downvoted comments to me. I usually upvote people who read and respond to me, regardless of whether I agree with them.
- Java: no
- Ruby: no
- PHP: no
- C: yes and noWhy do you still think there’s some yes in C? It’s not making sense yet that your memory is gone after you return? Returning a pointer to a local variable is exactly the same as calling delete or free on a pointer and then reading from it. You officially don’t own the memory after a return statement, so if you try to use it, then what happens is indeterminate. Again, since it doesn’t seem to be sinking in: it is always wrong to return a pointer to local memory. But, if you really really don’t want to listen, and you’re sure it works sometimes, then I say go for it!
[edit - that seems to be the point sigh]
Answered the rest, IDK.
Laughed! Great exercize.
i = 0; return i++ + ++i;
Could produce different results?
I could imagine the post-increment happening after the sum: giving 0 + 1…
This looks like a good explanation of why both are correct and why it’s confusing: https://stackoverflow.com/a/4445841
Sure, you've made your point, but you've made it in a ham-fisted way which doesn't really help people understand why a given undefined or implementation-defined behaviour is the way it is, and what things they should verify about the implementation in order to predict where their code will not work.
I disagree.
I can easily conceive of a (compiler, architecture, compiler options) tuple that simply crashes with an error at compile time or runtime with that code. Namely, some compiler for a 16-bit architecture with "sanitization" options enabled and optimizations disabled.
Integer overflow is one of the easiest "undefined behavior" cases to identify with mechanical checks. Much easier than bounds checking for example, where a general solution is quite tricky.
It's reasonable for some folks (especially working programmers who need to "get stuff done") to think that "what a reasonable compiler in their problem domain" would do is what "C" means. It's equally reasonable for other folks (especially compiler writers, verification experts, researchers, etc.) to think the ISO standard is what "C" means.
It would be great for the standard to be more "reasonable" and have less undefined behavior. But I, for one, cannot think of a more horrible, thankless chore than actually trying to make that happen. So much code is written in "C", and there are so many compilers and platforms, modern and legacy, that "C" runs on, each with their own notion of "reasonable," that it will take an incredible amount of work.
GCC and Clang are mostly compatible, and as far as the low-hanging fruit is concerned. If you consider them the authority, it generally resolves most interesting questions about what ISO decline to specify. I do not think that there is any great burning need for ISO to go and define things more rigorously.
Also, the vast majority of programming languages don’t have standards at all, and people still use them productively.
Java, JavaScript, Modula-2, Pascal, C++, Fortran, C#, F#, VB.NET, Eiffel, D, Ada, Common Lisp, Scheme, Python, Scala, Haskell.
k = 2
i = ++j+++++k++
i = ?
j = ?
k = ?
I think that's exactly the wrong takeaway here. Most of these have a well-defined result on a given platform (host + abi). You can measure that result. But it'll be different on a different platform. And the others don't have a well-defined result — a real-world compiler will produce a result, and you can measure that, but it might be different tomorrow. Unless you know the difference between platform-specific behavior and undefined behavior, you don't know which ones to avoid.
"Measuring" is exactly the wrong thing to do because often it is indicative of only your specific architecture/compiler.
C is essentially a portable assembler. It’s not enough to learn the language, you need to have deep understanding of the underlying hardware and compiler infrastructure.
The answer is: No, I do not know C or any other language for that matter. I know some implementations of various languages just enough to write sound and readable code. And as I forced to use a bit more languages than I like I rely on local/Internet search to keep my brains concentrated on accomplishing the actual task rather then effing them up trying to figure out some esoteric constructs.
Didn't quite know the purpose of the test, there's a difference between code for any machine and any compiler and gcc running on some vanilla x86, which is pretty common, and could have been the content (ex: you say it's undefined, that's obvious, everyone knows that already, but it's still deterministic... Here's a breakdown of what happens in practice, blah blah blah). There's a real difference between the kind of "knowing" here and say the kind of "knowing" with unallocated pointers.
If it said "esoteric implementation on exotic hardware" then it's easy, you know what they are trying to do. If it said C89 you also know what's up. But how it's presented, it's a guess.
This was endemic throughout schooling. Instructors would say "just do your best" and I'd be like "wtf? There's like 2,3, maybe 4 perspectives on this with different answers depending on how clever you're trying to be or what you're trying to get at... Might as well put "I'm thinking of a number 1 through 5" on the exam".
(You also seem to be assuming GCC is more predictable about undefined behaviour than it actually is.)
That's an assumption though. If it's undefined, the result may even depend on the order of some internal hash table which is affected by other code.
Undefined behavior means the standard says demons can fly out of your nose. Those demons might behave nondeterministically.
1. 8 (https://godbolt.org/z/Udcuj0)
2. 0 (https://godbolt.org/z/kZz1YQ)
3. -96 (https://godbolt.org/z/1Xq_Oo)
4. 1 (https://godbolt.org/z/MIo0s_)
5. 2 (https://godbolt.org/z/J0dsVW)
I agree, you can argue that it's unspecified, undefined, or whatever. It might not be well defined by the C specification, but none of these programs produce surprising output. Programming in the real world requires that you are able to read and write code like this, even if it requires that you investigate (and depend) the specific behaviour of your compiler/platform.
Perhaps “produce consistent results” would be more appropriate.
I'd like these things to be listed and marked "implementation defined", personally.
It’s an unfortunate truth that programming in the real world involves programmers who dare to explore these corners of C and claim to have answers to these questions. Stay away! Knowing C means knowing what is not defined as much as knowing what is.
The behavior of any program that evaluates `i++ + ++i` is undefined. The solution is not to find out how it happens to behave in some circumstances. It's to find clearer code that expresses whatever the original intent was.