So you think you know C? (2016)
wordsandbuttons.online
wordsandbuttons.online
Turns out the author does mean according to the standard, but thinks that “I don’t know” is a synonym for both undefined and unspecified behavior. It seems weird to use imprecise terminology in a post that’s all about lecturing about standards compliance.
So it's easy to say: I don't know.
On a TI DSP I used this millennium, sizeof(char) == sizeof(short) == sizeof(int) == sizeof(float) == sizeof(double) == sizeof(void *) == 1. Each and every type was 32-bit. Each memory address pointed to a unique 32-bits, ie: (int *)0, and (int *)1 did not overlap on that system. As sizeof measures addressable units, not bytes, so they're all size 1.
Even on more standard systems, int is 4 bytes on systems using an ILP32 or LP64 convention, but ILP64 is a thing too, making int 8 bytes.
The crazy one is `long`, which is 32 bits on some platforms and 64 bits on other platforms.
I solved that problem by never using `long`, opting instead for `int` for 32 bits, and `long long` for 64. `long` should be deprecated for 32 and 64 bit platforms, it's not fixable.
(I’m not comparing the languages in any way here, just praising a clean notation.)
This was a semantic quiz, not a technical one.
Obligatory xkcd: https://xkcd.com/169/
In either case, is "I don't know" wrong? Given the information in this quiz, I don't think it is.
Things have stabilized quite a while ago. 99% of people who write C or C++ in 2024 will never use any computers where sizeof(int) != 4, or big endian processors, or systems where floats don’t conform to IEEE-754 standard.
Why shouldn’t they write code which requires these particular details from their compiler and the target processor?
That said, if you use a variable-width integer type (`char`, `short`, `int`, `long`, `long long`) instead of a fixed-width type (`int32_t` & such) for anything other than passing parameters to existing libraries (including the standard library) I'd say you're Doing It Wrong. If you actually intend to have a variable-width, use one of the `_least` or `_fast` types to make it clear that you didn't just screw up.
Thankfully C23 takes this approach, although it'll probably take forever until it's as adopted as widespread as C99 is now, and that's still not widespread enough.
But taking this test, would you really check "I don't know" if you do know it is UB, when there is no option "platform specific" or "UB"?
This quiz isn't an IQ test, people, and the questions are intentionally trick questions. It's essentially a form of cynicism intended to demonstrate major design oversights of the C programming language. So calm down and have a laugh.
Also, "I don't know" is the right answer. If you can't accept that, you may need to meditate more.
1) 8
2) 0
3) 160
4) 1 (both clang and gcc output a warning)
5) 2 (only clang outputs a warning)
While I like the idea of this quiz, I think it would be more powerful if it provided examples of compilers / architectures where these are not the correct answers. (I also think thorough unit tests would catch most of these errors)
In fact, by not doing so they are making a subtle implicit statement that it is uninteresting to consider actually attempting to execute these snippets.
The third paragraph of the "P.S" of the article (you have to press submit to see it) is the one that really gives the game away.
If the compiler has defined behavior (and you have unit tests for that behavior) on all of these platforms, I don't think it is a huge deal. (Ideally you wouldn't... but sometimes its an accident or unavoidable)
As an example, while struct padding (problem 1) might not technically be in the spec, it is a cornerstone of FFI and every new compiler (that supports C FFI) has a way to compile structs with the same padding.
To my original point, if the article had instead given examples of compilers + architectures that produced different answers, I might feel differently. However, just saying mentioning that these weird edge cases are undefined (in the spec) doesn't mean much to me.
2) I thought the type would be promoted to short. It turns out that the result of the arithmetic operation is promoted to int.
3) The signness of char is platform dependent. It is signed on x86 and amd64, but unsigned everywhere else. After seeing my mistake, I would expect this to cause the answer to be -96 on amd64 from sign extension when it converted into an integer, yet it is 160, which is what I would have expected from a platform where char is signed. If anyone knows why it is 160 here, please let me know.
5) This is a classic. I know to answer I do not know because despite having an Operator Precedence, C famously says that this is undefined. I have no idea why the standard does this when there is a clearly right answer. Java for example makes this have only 1 right answer.
This might be ASCII or EBCDIC or something else local to a specific hardware implementation.
https://en.wikipedia.org/wiki/EBCDIC
So, maybe 0x20, maybe 0x40, maybe something else.
At least you know that '0', '1', .., ''9' are contiguous.
The wrapping rule here is that signed integer overflow is UB.
Early C compiler projects (eg: The Hendrix Small-C of ~1982) would get patched by some to support the full C language and extended to cross compile to and from whatever machines were about at the time, System/360's, VAX, PDP's, early PC's, BBC micros, etc.
It wasn't always the case that char encoding passed on by default, there was always the option to insert a trans table whether compiling or dealing with data stored in not native form (similar to data in big end V little end).
There are a huge number of assumptions in that simple chain of events, and if any of them are wrong you get a different answer.
However we’re also allowed to do something like this: i=0, a=i, b=i, b=b+1, RHS=b (RHS=1), LHS=a (LHS=0), a=a+1, i=a, i=b.
Probably quite a lot of other things are allowed to happen. Usual disclaimer that a standard-compliant compiler is allowed to vaporise your cat etc as part of UB.
The thing to Google is “sequence points”.
So You Think You Know C? (2020) [pdf] - https://news.ycombinator.com/item?id=37541685 - Sept 2023 (86 comments)
Free e-book (~2 MB): So You Think You Know C? [pdf] - https://news.ycombinator.com/item?id=22958870 - April 2020 (8 comments)
So you think you know C? (2016) - https://news.ycombinator.com/item?id=20366940 - July 2019 (322 comments)
So you think you know C? - https://news.ycombinator.com/item?id=12902304 - Nov 2016 (198 comments)
So you think you know C? - https://news.ycombinator.com/item?id=12900980 - Nov 2016 (1 comment)
So you think you know C? - https://news.ycombinator.com/item?id=12900279 - Nov 2016 (9 comments)
Same title different articles:
So you think you know C? - https://news.ycombinator.com/item?id=4657317 - Oct 2012 (13 comments)
So you think you know C: the Ksplice Pointer Challenge - https://news.ycombinator.com/item?id=3125891 - Oct 2011 (98 comments)
Primarily, I have determined that I really don't like dealing with manual memory management. For all the tasks I'm interested in, C's performance gains over a GC'd language are marginal at best. Number crunching? Julia is competitive with C. Web servers? The JVM will handle it just fine. Microcontrollers? For what I do Lua+NodeMCU or MicroPython does everything I need it to. Reasonably fast command line application? Go's got you covered.
When I do use C in 2024, I pretty much always cheat and use the Boehm GC, which is fast enough for whatever I need it for. I'm not smart enough to know if I handled pointers correctly, and I don't know that I want to spend the time to get smart enough.
Obviously systems C and manual memory management has a place in the driver and kernel world, and if you're genuinely good with it then my hat goes off to you, but I don't feel like it buys me enough today to use it much.
However you might well enjoy https://dtolnay.github.io/rust-quiz/
Like the C++ quiz, "Undefined Behaviour" is a valid answer, however, the quiz questions are about safe Rust, so that answer is always wrong.
I still get more than half of them wrong unless given far too long to think about it.
>
> What is the output of this Rust program?
>
>
> // JavaScript is required (sorry)
Tears is the output of that Rust program.
EDIT: I will never understand HN comment newlines.
(Point is, some "quizzes" are made cynically to demonstrate a flaw in some design. Don't get your ego all bruised up because you got some questions "wrong" on the internet.)
The password game is a game first and a commentary on how stupid "password rules" are second. The game makes sense in our cultural context (where stupid password rules are a thing) but it would be fun (but less popular) anyway.
The comments remind me of the VAX programmers in 1990 who wrote code assuming pointers and ints were the same size, and if yours weren't, they told you to get a "real computer".
That solidified my opinion that C and C++ are very different languages despite naive view being that C is just a subset of C++. They're fundamentally different.
C++ also broke implicit void pointer conversions, but at least that one had the reason that it was incompatible with function overloading. Not that anyone involved with C++ would tell you this. Instead, they provide bad C code that breaks the strict aliasing rule and claim that breaking implicit void pointer conversions somehow follows from it when in reality that is a non-sequitur:
Using "unused" volatile also works, but it incurs SRAM traffic, which might slow down DMA and cause other effects. So ideally it should be simple empty loop compiled to `b .` instruction (jump to itself).
Another infinite loop usage is for abort() or exit() handler, as a crash place. For this use-case, using "unused" volatile does not cause issues, because device is already crashed, however it still feels wrong.
Why? If you are an experienced C or C++ programmer you know about these quirks. Replacing an infinite loop without side effects with a crash sounds about as "bad" as optimizing it away.
Importantly (see for example https://blog.regehr.org/archives/140) it's not that the standard previously explicitly forbade this sort of optimisation, it's just that it wasn't very clear and compiler writers interpreted it differently.
So in summary, yes, they're different languages, but "modern" C is is less different for this particular thing.
namespace {
void Panic() { while (1); }
} // namespace
extern "C" void kernel() { Panic(); }
The clang++ really does optimize this away:
https://godbolt.org/z/1Wf135918
The same thing happens when I let clang++ compile a C version as C++:
https://godbolt.org/z/MxcGfWhej
However, if I compile it as C with clang, I get an infinite loop:
inline uint64_t makeRemainderMask( ptrdiff_t missingLanes )
{
// This is not a branch, compiles to conditional move
missingLanes = std::max( missingLanes, (ptrdiff_t)0 );
// Make a mask of 8 bytes
// No need to clip for missingLanes <= 8 because the shift is already good, results in zero
uint64_t mask = ~(uint64_t)0;
mask >>= missingLanes * 8;
return mask;
}
TIL the language standard defines right shift operator for unsigned types in a weird way, making so for uint64_t argument, a >> b is equal to a >> ( b % 64 ). I expected zero on the output for b >= 64, not that.The reactions a number of people have to C++ code that is also valid C code or close to being it are ridiculous. Some times, they even deny C++ code is C++ code. :/
In C and C++ it's undefined behavior (Really it should be implementation defined)
uint64_t mask = -( missingLanes < 8 );
mask >>= missingLanes * 8;
I only care about the AMD64 because the surrounding code uses AVX2 intrinsics, i.e. my new version is hopefully good enough. I’ve tried conditional operator to not rely on these details, however VC++ failed to generate fast code from them. It emit branches, and that code is kinda performance critical. li t0, 128
li t1, 1
sll t2, t1, t0
You'll get t2 := 1 since t0[5:0] == 6'b0 (i.e. no shift). It's a very sensible solution IMO if you don't have trapping arithmetic since you don't have to do anything special to handle illegal shifts, it just works.I love the C language, but there are now so many UBs in the language, it is painful to use.
That said, I am curious what the official answer is myself.
Wisdom is understanding how likely your code is to ever run on any platform where sizeof(int) is not 4, and if the answer is "not really," then stop worrying about it.
To succeed as a C programmer, you need both knowledge and wisdom.
This is just pointless complexity. It's why D's cast syntax looks like this:
cast(T)x;
i.e. cast is a keyword.https://docs.python.org/3/library/typing.html#typing.cast
typing.cast(typ, val)
https://haxe.org/manual/expression-cast.html cast expr; // unsafe cast
cast (expr, Type); // safe cast1) it'll be a compiler error anyways
2) of course there's room for confusion. It could be
cast(T) var
Or it could be
cast(var) T
And you must remember which is correct.
result = (a)(b); (foo)(a + b)
Or this: // Comma expression or function call arguments?
(bar)(first = side_effect++, second) (a)(b,c)
is that a comma-expression or an argument list? typedef long T, U;
// T is an argument T is a typedef
// vvvvvvvvvvvvvvvvvvvvvvvvvvv ~~~
enum {V} (*f(T T, enum {U} y, int x[T+U]))(T t) {
// T is an argument (until the end of function)
// vvvvvvvvvvvvvvvvvvvvvv
long l = T+U+V+x[0]+y;
return 0;
}
[1] https://hal.science/hal-01633123/file/jourdan2017simple.pdf int16_t a( int16_t a, int16_t b) { return a*b; }
uint16_t b(uint16_t a, uint16_t b) { return a*b; }
int32_t c( int32_t a, int32_t b) { return a*b; }
uint32_t d(uint32_t a, uint32_t b) { return a*b; }To make things weirder, I tried 8-bit and 64-bit versions. The 64-bit version behaved like the 32-bit versions, but the 8-bit versions had no undefined behavior checks. Why is 8-bit so different?
Write a variadic macro `CLEANSE_MACRO_ARGS` that can be used within one of C's unhygienic macros, to turn it into a hygienic macro without reducing the prettiness of the macro body.
In standard C, this requires C23 and only works for macros that are not used as expressions. Or you can use GNU extensions and make it work for expressions and work even on old compilers.
int x, y;
x ^= y ^= x ^= y;
I was using this for years until I realized that it is UD.That said, even if it were not, you did not define x and y. Reading them is also undefined behavior, which is what I initially though you meant until I read what followed them.
Good to have had that mistaken bit of errata corrected.
PS: 'sizeof with parentheses' looks so weird to me.
int main(void) { return -1 == (~1 + 1) }
I am fairly confident that the answer will be the same on every system on which that is run, but technically, the C standard does not guarantee that it is the same unless you use C23.
I still hope that you will have look at what are the standard type sizes before trying to develop or compile something for a weird platform or compiler...
a. Your target ecosystem(s) probably follows a few norms in cases not defined by the standard. If it works on the target ecosystem(s), nobody cares if it's non-standard. Testing on the target ecosystem, preferably automated testing, is the standard that matters.
b. Some of these norms are extraordinarily powerful. For example, check out NaN boxing[1]. Standard? Hell naw. Useful? So useful that it's practically mandatory for dynamic PL interpreters to compete on performance.
c. So let's say you're targeting an ecosystem that does some atypical stuff, i.e. doesn't follow the norms mentioned. Well, joke's on you, off-the-beaten-path ecosystems typically also have bleeding-edge tooling that doesn't implement the standards. So again, testing on the target ecosystem, preferably automated testing, is the standard that matters.
d. So let's say you're targeting a lot of ecosystems, so many ecosystems that you can't reasonably test on all of them, like if you're writing NetBSD or the JVM. Okay, well, the edge cases of the C standard still don't matter because you just shouldn't be using any of the off-the-beaten-path parts of the C standard anyway, since those parts are the parts which aren't implemented in lots of obscure architectures.
e. And sure, if you really just enjoy being annoyingly pendantic you can go and learn the whole C standard and go write your own blog test with integrated quiz and then condescendingly tell people they don't know C, with the implication that you do know C. But the reality is that knowing those weird edge cases of C isn't useful because if you actually use those edge cases of C then nobody else on your team can read your code.
f. And lastly, there's plenty of defined behavior which you shouldn't use, too. Just because you know what it does, doesn't mean the next person reading your code knows.
My answers to the questions in the test:
1. 8 on most ecosystems. If it matters, test it, but if it matters you probably need to use one of the non-standard extensions mentioned to do anything about it, so the standard didn't matter so much did it?
2. 0 on most ecosystems. But frankly, don't use `short int`, ever. Use int with the limits defined in `limits.h` if you just want the most efficient integer width, and if you need a specific integer width use the the types in `stdint.h`. Yes, I'm aware that `int` isn't defined as the most efficient in the standard, but `int` is one of the first things people implement in a C compiler so the likelihood of coming across an int that isn't the fastest integer size is absurdly small.
3. Play stupid games win stupid prizes. Who cares if this is defined? Even if it were defined it would be bad code. Don't multiply chars, FFS.
4. Play stupid games win stupid prizes. Obviously bit twiddling when you don't know teh structure of the bits is bad. The fact that Endianness is relevant in bit shifts and yet is not brought up by the author does in fact show the dangers of smugly pretending you know C better than other people.
5. Play stupid games win stupid prizes. Who cares if this is defined? Even if it were defined it would be bad code because depending on the pre/postfixedness of ++ is hard for most people to reason about.
[1] https://craftinginterpreters.com/optimization.html#nan-boxin...
thats both: funny and sad
The things I like about C aren't really about the language itself. There just aren't many programming languages that do these things well
1. Cross platform across Linux, MacOS (Intel and Arm), Windows (Without the use of MinGW And Cygwin), *BSD
2. Ability to create an executable
3. Small programs that start quick
4. Simple syntax that does not obfuscate when allocations are occurring
For GCed languages GO is probably the one that best fits this criteria for most people.
Sometimes I wish someone just took C and improved it. Instead of looking at C++ and thinking "what features should C have that C++ has" It should look back at C and go "What features does C not have, that it should"? Some that come to mind are
1. A better syntax for creating pointers and dereferencing them
2. Fixed const so that it works with array initialization
3. Modules instead of headers
4. Cross platform string types that encoded length 5. Hash tables
6. Remove preprocessor macros and replace common use cases with things built into the language
7. Build system that didn't rely on Make, Cmake, Ninja, Autotools, configure etc
8. Easy interop with C
Maybe this is a lot to ask, but most languages I see are solving waay more problems then this.
2,4,5, can be done without leaving C, which is why most people homebrew their own solutions. I think Pascal actually comes close, but it's held back by the fact that the editor support is pretty poor. Pacal-Mode in Emacs is not good, and Lazarus with its multiple floating windows is infuriating to have to alt+tab multiple times to switch to and from it. Plus I could never actually get it to build my programs properly without throwing some dwarf errors.
Recently I've been playing around with Gambit-C and it seems to suit most of my needs. It compiles to C, and then the C code is compiled with Clang, MSVC, or GCC depending on what your platform is. I get the benefit of a language that has more features than C with it being R7RS (but still being small because it's Scheme) and implementing the most common SRFIs (including hashes), but with all the portability of C. It also has dead simple interop with C using the C-Lambda and C-Define special forms (you can write C directly inside the Gambit code)^1 which means you can leverage C code with hardly any effort. There are tradeoffs with this approach, but IMO almost all of them are ecosystem related which can be fixed, as opposed to language related, which might be impossible to fix without breaking compatibility with existing code
1. https://www.deusinmachina.net/p/gambit-c-scheme-and-c-a-matc...
But UNIX being free beer gave another push to C, that those languages did not have.
I think that should be implementation-defined behavior, not unspecified behavior? IIRC unspecified behavior in C is not required to be known or consistent.
"So you think you know C, huh? What does this horrible piece of code do that no one would ever write and if you see it should be nuked from orbit? <ridiculous contrived example>"
I just clicked I don't know on all of em, cause I realised what the metagame was from a mile away.
Some of these things are at least a little insightful, but overall, kind of a waste of time past a certain point, unless you work on a compiler or something like that. Much more constructive would be making resources on best practices for things like memory management, string handling, knowing where the footguns are in the stdlib and other common libs, managing complexity etc. Most bad C code isn't because of some standards gotcha, it's because of those things.
But if I'm in an existing codebase, and I see a certain volume of this kind of "potential UB everywhere" code, my first instinct isn't let's spend however long trying to understand what exactly this code may or may not be doing/relying on according to dusty corners of the standard and the datasheet(sometimes you may have to, but I find it's rare). I prefer to approach it as "what is this code supposed to do, and how can it be done more sanely? Usually I find I can replace the crazy code with sane code in a fraction of the time it would take to fully understand the crazy code.
/s
My personal gotcha was finding out I was relying on shifting an unsigned integer by the bit size of the integer to be 0, as if "you shifted every bit out of the integer, leaving 0". Shift any size under the integer's size in bits? Yep, those bits were shifted out leaving 0. Shift that last bit? Nope, that's Undefined Behaviour, and suddenly it's not 0 just because you changed a flag.
Most of us would say that it has a value of 32 as defined by ASCII. This would be wrong on a platform that was using EBCDIC.
We all have to make assumptions in our programs, which are sometimes wrong.
RedHat Linux has also been ported to this platform; it's probably ASCII, but not sure.
Oracle has a native client on zOS that does transparent conversion between ASCII and EBCDIC. It was probably built with such a compiler.
There was a paper that the original port of Research UNIX ran as a client on an IBM TSS/370 kernel, which likely was not ASCII.
For what it's worth, this is a common sentiment people have when they first come to blows with formal. "If you can't verify everything, why verify anything?" is a reasonable question to ask but it's not a practical position to hold since formal does increase the quality of software/hardware. Don't argue for less formal, fight for more!
The point is “undefined behavior” is more about gaps in the standard than bugs in programs.
And not least of all, we've learned more about programming language development than when C was created.
I really wish we could let this die. All it shows is that if you provide insufficient context, you get insufficient answers. There are better ways to point out C's shortcomings.
We can know what is happening in memory in a human readable way. And that is the beauty of it. If you can’t tell what’s happening in memory to some degree, then it’s not that you don’t know C, it’s that you don’t comprehend computer architecture.
They are technically `undefined` according to the C standard, but are the behavior of every mainstream compiler. So much of the world's open-source code depends upon these that it's unlikely to change.
Using clang version 15.0, the first 2 produce no warning messages, even with -Wall -Wextra -pedantic. Conversely, the last 3 produce warning messages without any extra compiler flags.
The behavior of the first two examples are practically defined even if undefined according to the standard.
Now, when programming for embedded environments, like for 8-bit microcontrollers, all bets are off. But then you are using a quirky environment-specific compiler that needs a lot more hand-holding than just this. It's not going to compiler open-source libraries anyway.
I do know C. I knowingly write my code knowing that even though some things are technically undefined in the standard, that they are practically defined (and overwhelmingly so) for the platforms I target.
That said, warnings do not necessarily mean that the code is invoking undefined behavior. For example, with if (a = b) GCC will generate a warning, unless you do if ((a = b)). The reason for the warning is that often people mean to do equality and instead write assignment by mistake, so the compilers warn unless a second set of braces is used to signal that you really meant to do that.
Third and fourth are only defined in some implementations.
A better test is -Weverything.
Have we crossed the point yet that the majority of new microprocessors and microcontrollers sold each year are 32+ bit yet? Most devices I’m familiar with still have more 8 and 16 bit processors than 32 and 64 bit processors (although the 8 bit processors are rarely programmed in C).
Then you admit that the microcontroller world presents exceptions.
You've now arrived at "I don't know" the answer.
The article never said "using mainstream C compilers".
Something that a lot of C programmers do...
Most people who write for typical desktop and mobile computers don't do C. They tend to do C++ or other, higher level languages. Those who write C tend to do either quirky embedded code, or code that is highly portable, in both cases, knowing about such undefined or implementation defined behavior is important.
If you intend on relying on such assumptions, make it explicit, for example using padding, stdint, etc... On typical targets like clang and gcc on Linux, it won't change the generated code, but it will make it less likely to break on quirky compilers. Plus, it is more readable.