So you think you know C?
hackernoon.com
hackernoon.com
If a person wanting to improve guesses correctly, did they actually learn anything? Can they use that next time? They definitely can if they get it wrong and have a wrong answer to start research with.
Edit - I see that I misread your comment, please disregard this. I thought the quiz had don't know and the other options you specified.
I understand, in abstract, I might be on an EBCDIC system, and things will be different. But that's not the reality of C. C isn't just a formal specification.
I spent a while on 1, thinking about how my compiler would handle that. 2 I was 90% confident about. 3 and 4 I was 100% confident about.
No, C IS just a formal specification.
What you are talking about is C on your particular compiler on your particular platform. If you are writing C code just to be used for that scenario then you can use all of the implementation understanding you want.
If, however, you want to write C that will have consistent behaviour across multiple compilers and multiple platforms, then you need to limit yourself to the behaviours that the standard guarantees. Otherwise the compiler behaviour may change and your program behaviour will change.
Even between different versions of the same compiler implementation defined behaviour can change (although assumptions about sizes probably won't).
It doesn't get you as far as the ideal goal which is to write code that will silently work on another platform. Static asserts mean that you then have to go and write more code when they happen.
Right up until an optimizing compiler throws your assumptions under the bus.
Compilers are allowed to (and both GCC and Clang frequently do) assume that your program will never invoke undefined behavior, and optimize accordingly.
See: http://blog.llvm.org/2011/05/what-every-c-programmer-should-... Also maybe: http://blog.regehr.org/archives/1307
This can lead to some pretty bad security flaws: https://lwn.net/Articles/575563/
You might be taking issue with the question itself, but I think it's a valid question because C compilers tend to be lenient with what they compile.
It's semantics. On many exams, "don't know" is a correct answer when you can't know because info is intentionally left out.
Also, why would anyone write ' ' * 13? For most of these questions, if your first instinct is "that doesn't look like good C code", your answer is better than the "right" answer.
However, "it's implementation specific and it's most probably this, or that if you use such and such a compiler flag" is different from "I don't know".
As a rule of thumb, don't assume things about C's abstractions, read the standard instead, or ask friendly humans who do if they check out.
To be honest, I hope compilers don't do such things. I would vastly prefer to see it run Tower of Hanoi simulation in Emacs at this point.
But evading bound-checking of 8bit math done in 32bit registers is totally reasonable (by the standard of usual UB optimizations), thanks.
It does, via bit-fields.
It is, but not required to be an octect. Quoting from the standard (n1570, 3.6):
byte ==== addressable unit of data storage large enough to hold any member of the basic character set of the execution environment
Later in a comment (page 44 in pdf, note 49): A byte contains CHAR_BIT bits
Casey Muratori, who's a better C programmer than me, is of the opinion that you should check your compiler's behavior rather than read standards.
(Not that you should take his word as gospel, I don't fully agree myself, but it's worth considering.)
It's much better to avoid undefined behavior entirely. The problem with this idea is that it's hard to learn and remember all the different kinds of undefined behavior.
>> Casey Muratori, who's a better C programmer than me, is of the opinion that you should check your compiler's behavior rather than read standards.
In Casey Muratori's case, he's the type of person who cares to understand what the compiler is actually generating under the hood. He doesn't exactly trust the compiler to always do the right thing.
There was an episode of Handmade Hero where he caught an optimization #fail of the compiler and tried to figure out what was going on (live). https://twitter.com/cmuratori/status/596775025023287297
Casey Muratori also has been developing in Visual Studio for a very long time. Visual Studio has a long, tortuous, history of being non-stanards compliant and buggy. And Microsoft until very recently, never cared to fix anything. (And it still doesn't conform to C99, let alone C11.)
Whatever C related updates you might see, are the ones required by the ANSI C++ for C compatibility.
For C99 and C11 support there is the clang frontend to C2, the Visual C++ backend, currently taking an LLVM role at Microsoft, being shared between clang, VC++ and .NET Native.
Given: 1) There are various optimizations that may kick in or not depending on UB. 2) The cleverness of the optimizer may and does(!) change between versions. 3) There's no "-warn-on-optimizer-changes" flag for any compiler.
Do you think it's realistic to compare compiler output whenever you upgrade your compiler? Do you think it's realistic for 1M SLoC projects? If not, do you take a statistical approach with random sampling? If so, I'd be very interested. If not... what exactly are you sure of and how?
There is one thing I have noticed in gcc and clang which is done against the standards: The sign of bit sized int in a struct.
eg: In a struct with member 'int a:8', The standard says that 'a' can be signed or unsigned (based on the machines default of the sign of char). But In gcc and clang, this is signed always regardless of the machine.
Wait, are you saying gcc/clang's behaviour is permitted by the standard, or not?
It's permitted. The same way as a char can be signed or unsigned: Implementation defined. But in this case, it's always signed regardless of the architecture (against how a char is handled).
From c99 draft (n1570), J.3.9 (Implementation-defined behavior): Whether a ‘‘plain’’ int bit-field is treated as a signed int bit-field or as an unsigned int bit-field (6.7.2, 6.7.2.1).
But say, another compiler (ARM compiler) treats this differently [0]: Untill version 5 of the compiler, such an int was unsigned by default. Later versions defaulted to signed (as do gcc/clang).
[0] http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc....
Do both.
Disassemble your compiler's output to see what horrors the optimizer inflicted on your code to "break" it. Read the standard to understand why it thought it was an okay optimization to do in your case, to avoid running afoul of the same mistake again. Ask what other optimizations are enabled by this undefined behavior, the better to spot them when the code your coworker's former boss's roomate, who interned for a month or two, invokes undefined behavior.
Repeat.
(To be fair, I haven't read the standards cover to cover - perhaps I should, I'm still occasionally learning about new and exciting forms of UB...)
(edit: i misread, it's always 1) To be pedantic, a size of char is defined by CHAR_BIT in limits.h. As well as all of the min and max values and lengths in bits for integers.
To add to the discussion, many high level languages don't have a hard min and max values (haskell, that is popular in these woods, i see has minBound/maxBound).
From the "So you think you know C?" test i answered... probably correct for the "standard" x86 implementation, while knowing that they are implementation specific (for 4 i'm not sure). The authors "Don't know" is his own, i would have answered the questions as "It depends", but that was not offered.
In the end i agree with the general notion here, that this is ugly C.
PS Integer wrapping is undefined as well, so "good" C code should work around it. (i miss this from assembly)
To be even more pedantic: it's the number of bits in a char that is defined by CHAR_BIT. However, sizeof returns the number of chars needed to store a type, where sizeof(char) is defined to be 1.
Nevermind me.
Signed integer overflow is undefined. Unsigned integer overflow is defined.
From ISO/IEC 9899:2011:
3.6
byte
addressable unit of data storage large enough to hold any
member of the basic character set of the execution
environment
NOTE 1 It is possible to express the address of each
individual byte of an object uniquely.
NOTE 2 A byte is composed of a contiguous sequence of bits,
the number of which is implementation-defined. The least
significant bit is called the low-order bit; the most
significant bit is called the high-order bit.
Edit: Reading further into the C11 spec, `5.2.4.2.1 Sizes of integer types <limits.h>`, it says that `CHAR_BIT` must be 8 bits (or larger). A google search suggests there exists some processors that have 16-bit bytes and others that have 32-bit bytes.It is worth remembering that POSIX (and Windows probably too) mandate 8bit chars so there is no point being defensive about it on these particular platforms. And I kid you not, I have seen people who are. Because "ISO C99 this and that".
Ever since, reading introductions to C that say things like "a char is usually one byte in size" make me cringe. I never bothered to check, though, what C89/C90 had to say on this.
LaTeX (and probably TeX alone, but I never use it without LaTeX) supports something similar, but AFAICT only in appropriate contexts (matching pairs of `` and '' are converted to the appropriate quote, as are ` and ', but they don't seem to affect "typewriter text", so the smart quotes stay out of my code snippets). Or maybe I just haven't noticed LaTeX munging all over my document.
One might argue, though, whether it's really a good thing to be resilient against crappy WordPress blogs, since it encourages running copy-pasted code from the web. But I guess that would happen anyway, regardless of language designer intentions and a little frustration with copy-pasting won't cause people to actually read the code they're running.
And compiler developers in the real world are also increasingly nervous at this state of affairs. The subset of C that works in practice is unspecified. There are plenty of optimizations that compiler developers would like to do, but they don't because they're worried about breaking some code somewhere that relies on undefined behavior (like the code you're talking about). At the same time, performance competition pushes compilers to exploit undefined behavior more and more heavily. There is (IMHO, well-founded) worry across the board that this situation is not sustainable. Experience with this problem motivated the design of languages like Swift, which try to clamp down on UB drastically—which also has the happy side effect of making articles like this one inapplicable.
Undefined behavior has already had practical consequences, such as security-sensitive VM escapes: https://bugs.chromium.org/p/nativeclient/issues/detail?id=24...
Back when I used to code C professionally on early 2000's, these type of issues were quite common.
Specially since regardless of what one reads on HN, or hears at most conferences, most of those companies, which tend to see IT as a cost center, usually don't employ any processes regarding code reviews, static analysis or unit tests.
Which in languages like C this means an heavy price when things blow up, or the way things are going with IoT.
So given the increase in CVE entries, this still matters a lot.
The list of qualifiers and disclaimers in that sentence made me dizzier than taking that quiz on dark corners of C.
Less is more.
However, I think that your original sentence is fine too. GP is being a bit pedantic.
CppCon 2015: Herb Sutter "Writing Good C++14... By Default"
https://www.youtube.com/watch?v=hEx5DNLWGg
When he asks the audience how many use such type of tools regularity, he gets around 1% as result.
I see this with the other languages as well, such practices are still seen as extra costs, shuffled to the end of the projects.
Unless the customer has software development as core business, there the attitude is quite different.
Very sad.
Or I could be way off base - I've never attended CppCon. And I do love me some ASAN when I need it ;)
So... could you explain what you're insinuating so I could respond to that instead of just guessing?
[1] Not sure about CppCon, but all "technical" conferences I've personally been to did have a lot of late-night drinking, but that was secondary to the technical stuff.
What I'm saying is that knowing these minute details doesn't make one a proficient developer with the language, nor does getting every one of this list wrong a bad developer. Sure, with 10 years of experience most people would run into one or two on the list and know about it; still doesn't make the premise that not knowing or caring about trivia competitions makes one not 'know' the language.
The moment one needs to work with more than one compiler, including different versions of the same compiler, one needs to become a language lawyer to keep some sanity.
[1] I know the source of error wasn't related to these exact questions.
https://gcc.gnu.org/onlinedocs/gnat_rm/Value_005fSize-and-Ob...
On second thought, who cares about reading its for squares and nerds. I would rather spend 2 hours Google searching what is an octet and finding out how enums are represented... :p
Under Keil C51, alignment is always 1. This is legal C.
Under Keil C51, int is 16-bit. This is legal C.
Under Keil C51, sizeof(pointer) is anything from 1 to 3. This is an extension, but a very popular one.
Don't assume that just because you can write crash-free code on x86, you "know" C.
I think this is the funniest thing in this all. You might cross all your t's and dot your i's, and you still may trip in a pitfall when your compiler/platform vendor happens to actually diverge from the standard. Such is life with decades old language with at least as many implementations.
We used to do that by hand.
For x86 in 16-bit realmode (i.e. when you're targeting PCs on MS-DOS), that is also true. Likewise, pointers can be 2 or 4 bytes.
On the other hand, if you are working with such an "exotic" (i.e. not ILP32/LLP64/LP64) platform, you probably already realise that a lot of other things are also very different and the fact that the syntax looks much the same and it's still called "C" is the least of your worries.
Microchip's PIC microcontrollers also come to mind as being an architecture with a C compiler, yet with very different characteristics from the "usual" x86/ARM. (AVRs, Z80s, and the like are more similar to 16-bit x86 restricted to a single segment, a "SIP16" data model.)
That said I ported a bunch of stuff from the AVR to a 32 bit ARM Cotrex and the mostly non hardware dependent stuff stuff just worked. On the other hand stdint.h is an old friend to me.
Amusing bit
int apple = 74000; // fails on 8 bit machinesTest on every platform and every compiler you support, or it's not going to work. Avoid undefined behavior, sure, that's probably good practice. But there is no need to bend over backwards for compilers from the 90s targeting the AS/400 if you are writing a linux application that only needs to support Ubuntu 16.04. If your only means of avoiding broken behavior is "we're going to try really hard not to do anything undefined or implementation-defined" you will fail. But if you pass your functional tests, your user doesn't care whether you are invoking undefined behavior with every function call.
On the other hand, I don't agree passing tests is enough. No matter how many tests you run, you can't guarantee there's no bug in your code. This is important especially when security is a critical requirement.
"Nobody in their right mind should use code like this in production".
You should avoid code like this at all cost. If you want to use this code because you feel smarter, you will be creating enormous amount of problems in the future and you are actually dumb.
The way to win the game is not playing it. Specially if you are programming nuclear plants.
If you see this code what you have to do is replace it, not understand it. It works today but when you change the compiler or the architecture 10 years from now, it will be such a huge pain in the ass to find all the bugs and undefined behavior it creates.
http://www.gowrikumar.com/c/index.php
I think this page contains high quality C puzzles.
If you know something similar, let me know, I would love to solve.
http://www.gimpel.com/html/bugs.htm
BTW (apologies), the correct past-tense is "In 2008, I started to learn the C...".
Most online puzzles uses the former (I have also seen this in hacker rank), but the standard says that the main function should be defined as void if no argument is accepted (or some implementation defined manner)
:}
I'm working right now on a system that mixes Aarch32 and Aarch64 cores. Communication between these cores has to be rock-solid.
Previously, I worked on a DSP architecture that could switch at runtime between a 16-bit more and a 24-bit mode. I still sometimes have nightmares. A particular bug that took months to resolve involved app code that expected one mode and interrupts that went into another. The compiler did the right thing most of the time. Most of the time.
int i = 0; int j = sizeof (i++);
// Thanks. :-)
Will it optimize: http://ridiculousfish.com/blog/posts/will-it-optimize.html
edit: it's actually about gcc rather than the standard itself, still pretty interesting though.
Yes, inc/dec operators take extra care. Structure packing requires platform awareness, similar to how endianness can matter for data manipulation.
(a) char may be signed or unsigned so
char a, b;
...
short c = (short (a) << 8 | b);
will not work to construct a 16-bit value from two 8-bit values.(b) NULL is not a useful macro. For example,
printf ("%p", NULL);
can have unexpected results.Those things mentioned in the article? Outside of DSPs and GPUs, you can ignore them.
(((short)((unsigned char)a)) << 8)
Which works but goddamn is it ever hard to read. Bracket highlighting makes it just bearable. Fortunately I've abstracted this kind of stuff out as much as possible so I only have it in a few places, but it still bites me in the ass sometimes. Curious if anyone has a cleaner way of expressing this?How do you figure?
#define NULL ((void *) 0)
or just plain #define NULL 0
The latter works correctly because a literal 0 in pointer context is interpreted by the compiler as a null pointer.If it is 0, the printf() provides no pointer context, so the 0 is interpreted as an int. That works... poorly if sizeof (int) < sizeof (void *)
Is there any reason for an implementation to elect to use the non-casted version of NULL? It seems like an oversight.
Even if by some chance half of your C team knows all of the ins and outs, subtle things like this can slip by even when under review.
"Deep C" illustrates the difference well between people who know C and people who know C. A bit of it is language lawyering, but some of it can be practical - although perhaps specific to certain projects or environments.
IMO this sort of knowledge is most practical if you're maintaining legacy software or in specific environments that you would otherwise be making a mistake in. People writing evergreen software or in specific targeted environments don't really need to worry about all these types of "Gotch'ya!".
And if that colleague won't learn, that sounds like an issue in need of a bit of active management...
Once an idea enters his head it's hard to dislodge it. He learned programming from a mediocre CS school (despite the expensive price tag, which they deserve for their medical and law schools, not their engineering disciplines). And when something works, it's hard to convince him that it's wrong (that is, correct by coincidence) or unstable (in this case, sizeof(int) and related issues). He isn't technically (one of his favorite words) wrong in many instances, but problems will arise when we change OS, compiler, or CPU.
I'm slowly converting him to seeing things my way. My main contribution has been to mentor the junior developers. Things will change as their contributions worm their way into our projects.
Maybe I was just too ignorant and was limiting myself to what Turbo C had to offer on the MS-DOS side.
Of course on Windows we had all those WORD, DWORD, BYTE, BOOL,... macros already, but I was only thinking about actual C standard bit-width types.
sizeof(int) >= 2
sizeof(long) >= sizeof(int)
That's all you can expect.
The best approach is still to not expect anything.
§5.2.4.2.1
CHAR_BIT is at least 8
USHRT_MAX is at least 2^16 - 1
UINT_MAX is at least 2^16 - 1
ULONG_MAX is at least 2^32 - 1
ULLONG_MAX is at least 2^64 - 1
§6.2.5p8 guarantees that integers of increasing conversion ranks obeys subset relations, and §6.3.1.1 lays out that short, int, and long have conversion ranks in that order.
Thus, sizeof(char) = 1 ≤ sizeof(short) ≤ sizeof(int) ≤ sizeof(long) ≤ sizeof(long long).
Not really "64 bit" tho-on IBMi pointers are 16 bytes with teraspace enabled.
The test could also have included questions on things like packing bools and enums into structures. Again critical to get righht when sharing or transmitting structs, or anything bigger than a char.
Things like this were especially a pain with industrial network protocols that we used in our applications during the transition from PowerPC to Intel on OS X. All the CFSwapInt16HostToBig() and CFSwapInt16BigToHost() stuff...
And earlier before I learned about '#pragma packed' it took me a while to find out where all these extra zeroes suddenly came from :-D
D, Rust, and Swift are the only languages I know of that have "systems programming" as a focus.
You might still try to use FreePascal.
If targeting embedded development, there are a few Basic, Pascal and Oberon compiler vendors to choose from.
And there is also Rust and Swift if you are happy to stay just on Apple OSes (Open source Swift is still a bit of WIP).
Simple. It keeps asking code questions about undefined behaviors that MUST NEVER be used in real world programming.
#StandUpAndLeaveTheInterview
unsigned int i = 16;
return (((((i >= i) << i) >> i) <= i));
? (I changed `int` to `unsigned int`.)It's basically a question about knowing when you'll get burned by the size of int varying.
(i >= i) // 1
((i >= i) << i) // 0x10000 or 0
(((i >= i) << i) >> i) // 1 or 0
(((((i >= i) << i) >> i) <= i)) // ((0 or 1) <= 16) == 1
Where did I go wrong?UB.
> If the value of the right operand is negative
> or is greater than or equal to the width of
> the promoted left operand, the behavior is
> undefined"
(http://www.open-std.org/jtc1/sc22/WG14/www/docs/n1570.pdf - ISO/IEC 9899:201x (C11 Working Draft N1570): 6.5.7 Bitwise shift operators)In section J.2 Undefined Behaviour, the standard mentions that it's undefined when "an expression is shifted by a negative number or by an amount greater than or equal to the width of the promoted expression"
[1] http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf
Why would you bit-shift the result of a logical comparison?
This code makes no sense, it's wrong, and should be rewritten.
I'm worried. What blew up?
I'm not C expert, by far; I only have an instinct that if it's about C and memory layout, I shouldn't rely on my understanding and check it instead. On every architecture I'm targeting.
I had my formative C experience in the DOS era with 16-bit compilers and then SGIs with 64-bit compilers (and I had the "fun" of porting grad student written code from SGI to Linux) so you get burned with expectations about int length. Packing makes it even worse as many RISC machines (like MIPS) are word addressed and non-aligned loads require a bunch of bit fiddling or you get a Bus Error (SIGBUS).
In this particular example the evaluations and effects are allowed to be interleaved:
pre = I + 1; // evaluating the pre-increment
post = I; // evaluating the post-increment
I++; // effect of pre-increment
I++; // effect of post-increment
return pre + post; // evaluating the sum
Simply put, pre and post-increment are allowed to occur "simultaneously".applause
Though I'd agree with the comments on how this is kind of silly. The answer is not always "I don't know" but rather, in some of those cases, "I didn't define my data types well enough to know for sure, depending on CPU architecture and the compiler".
Do you think you know the English language?