Who Says C is Simple?
cs.berkeley.edu
cs.berkeley.edu
C is a small language. C is gussied-up assembly. It is the ballet of programming language.
People who say C is simple really mean that it doesn't do anything behind my back.
A minute to learn, a lifetime to master.
The language is fairly simple in terms of features and aspects you have to learn to use it, but concepts like pointers and direct memory management are difficult to master. Programming in C is like building a building out of bricks; bricks are relatively simple objects, building a building out of them is not a simple task.I assume you never compile with optimisations turned on then?
Same was with pointer arithmetic and function pointers.
Right now I am struggling mightily with monads mostly because the voice in my head tells me the thing all of the internet is singing - they are hard don't bother.
It's when there are bugs in code which uses pointers that the weird can kick in and really give you a headache.
=D. Just need to be the discord to the Internet's harmony.
> The answer depends on whether the optimizations are turned on. If they are then the answer is 3 (the first definition is inlined at all occurrences until the second definition). If the optimizations are off, then the first definition is ignore (treated like a prototype) and the answer is 4.
struct tun_struct *tun = __tun_get(tfile);
struct sock *sk = tun->sk;
unsigned int mask = 0;
if (!tun)
return POLLERR;John Regehr asked people to submit str2long() implementations that didn't execute undefined behaviour. Despite being well warned about avoiding overflows and other undefined behaviour only 35 of the 78 submissions passed the test suite[2].
Even a trivially small 2 line C function can result in different results on the common compilers depending on which compiler and which optimisation level you use[3]. Compiler engineers struggle to agree on and correctly implement the "simple" rules of C. To me that's an indication that they aren't so simple.
I would say that C doesn't do anything behind your back as long as you don't ever use an optimising compiler or you pay very very close attention to the standard. If you forget then suddenly your check for pointer arithmetic overflow has been "helpfully optimised away" behind your back for reasons that are not immediately apparent at all[4].
I do agree that apparently simple rules can lead to complexity in use, and C is full of this too. For example you would think in a lower level language like C viewing a chunk of memory as a different type would be trivial, but the strict aliasing rule means you have to be a language lawyer to understand what is allowed and what isn't[5] as it makes the intuitive solution into undefined behaviour (which is extra pernicious since most of the time it will work as intended, until somebody compiles the code with a smarter compiler at a high optimisation setting).
[1] http://blog.regehr.org/archives/963
[2] http://blog.regehr.org/archives/914
[3] http://blog.regehr.org/archives/482
[4] http://pdos.csail.mit.edu/~xi/papers/stack-sosp13.pdf example 1
While that is certainly nice, I suspect it makes it hard to really compare based on just pagecount. I doubt prose and such a formal definition like that are equally dense.
[1] http://mythryl.org/my-Mythryl_is_not_just_a_bag_of_features_... (go to the middle of the page or search for "pages")
Thus YOUR SPECIFIC CPU's instruction set may be well-defined. But if you say "your CPU" to a group of people with different CPUs, there may be no simple statement that is well-defined and generalizable across all of them.
That was how it worked in the 1990s. Nowadays, a C programmer needs to figure out which undefined behaviors are justifiable and which should be avoided at all costs because they will be used by the compiler to justify optimizations. And Signed arithmetic overflow used to be in the first category, now it is in the second one. So is the use of uninitialized variables.
Not to appeal to authority, but I worry about these things for a living:
http://blog.frama-c.com/index.php?post/2013/07/11/Arithmetic...
http://blog.frama-c.com/index.php?post/2013/05/20/Attack-by-...
If you cannot be bothered to read that much, then please simply compile int f(int x) { return x + 1 > x; } at different optimization levels with the compiler you already have, and observe the values for f(INT_MAX) in each case.
If you have a relatively small number of users who understand the gist of the language, it can be expressed in a few pages.
If you want the language to be useful beyond a trivial description, you'll have to add some complexity, which leads to weaknesses.
If your language becomes world-scale popular, you're going to have to spend specification space dealing with things like explaining that under certain conditions yes in fact a Boolean variable can have a value other than "true" or "false".
By not-doing-stuff-behind-your-back, I mean that you can map the code you write into machine instructions in many cases. An optimizer will move stuff around on you but it is typically local(-ish) manipulations. That's much less black magic that garbage collection. You have to manage memory yourself.
By being-ballet, I mean that it is hard. I am a swing dancer. The best dancers I know all took ballet classes as kids. None of them dance ballet today. It is probably coincidence/age but the best programmers I know all spent a decade or more working in C. I think working in C gives you a level of understanding about what a computer does that a higher level language doesn't. That said -- if you want to get stuff done, use a higher level language. I'm really good at C coding but I'm x2-10 faster when in C#.
Well I think that is where we disagree. I don't find it easy to hold all the rules of C in my head. It's plenty large enough that there are things I hardly ever use. Even the common parts of the language like arithmetic on integer types can get complicated very quickly if you want to be sure your code contains no undefined behaviour or works correctly for INT_MAX etc. Often you have to understand not just what the standard says but also what your compiler/target architecture does for the many implementation defined things.
>Yeah, there are edge cases (important ones even) and you can build god-awful complicated expressions if you want.
You don't have to write long or complicated expressions for things to get tricky. That was the point of this example: http://blog.regehr.org/archives/482 It's a simple function yet mainstream compilers got it wrong for years.
I don't consider these things as edge cases because they come up all the time and have caused countless serious bugs in real world C code.
>By not-doing-stuff-behind-your-back, I mean that you can map the code you write into machine instructions in many cases.
That is becoming less and less true with modern compilers. Vectorizers will kick in at different optimisation levels and depending on various heuristics that I'm not sure even the compiler authors would be confident in predicting for more complex code. They can perform a lot of complicated transforms. Undefined behaviour means lots of code can be modified in fairly unintuitive ways.
>An optimizer will move stuff around on you but it is typically local(-ish) manipulations
clang includes a link time optimizer: http://www.llvm.org/docs/LinkTimeOptimization.html#example-o...
By your argument, it feels like you'd say the language is large because you have to understand that rule if you really cared about consistent results everywhere. I agree with the author -- small but not simple.
Platforms differ and it leaks into the language precisely because it is such a simple language. It is simple like HTML 1.0 is simple. It is simple like the Bill of Rights is simple. There are a limited set of rules but there a lot of undefined behavior as a result.
I guess you believe HTML is complex because you have to understand CSS these days to do anything.
2. I find it interesting that you use what is clearly an edge case and then argue that because they are common it is not an edge case.
The example of "int foo(char x) { char y = x; return ++x > y; }" is almost the textbook example of an edge case. Seriously, don't trust me. Ask around and see if you can find 10 people who know C well that would consider this mainstream (excluding embedded developers).
There are countless serious bugs in real world C code because (a) there is so much damn real world C code and (b) it doesn't exactly protect from shooting yourself in the foot.
My experience working with a big program that ran on Windows, 2-5 flavors of UNIX, and VMS (both DEC and Alpha) is that the bulk of the real world errors in C code do not have anything to do with undefined behavior across platforms. They have to do with memory management (null pointers, buffer overruns, etc) and poorly written macros.
3. I agree with you that adding optimizers introduce a whole set of things you have to hold in your head that push you into 'large' territory. Just like programming on a GPU makes you rethink everything about how you organize code and writing for embedded code has its own set of rules.
But how does that make the language large?
My point is that a language that says "we manage memory on your behalf inside of a VM" is doing a lot more for you.
Boy, are we beating this thing to death or what...
...
I like C quite a bit but I wouldn't want to make a living programming in it today. I drop back into C when I have a compute kernel that needs it but 99% of the code remains in C#. C is small but too simple for the problems that I'm solving today.
Merriam-Webster definition of "Terse" :
1: smoothly elegant : polished
2: using few words : devoid of superfluity
Even then, there might be dragons waiting for you.
I do.
C is like chess. There are few simple rules (compared say to C++ spec). But knowing the rules doesn't mean you'll end up beating Kasparov. It still takes skill and practice to be a good C programmer.
So, the language itself, is pretty simple compared to other popular programming languages as far as it has a simple syntax. I can teach someone the rules of chess pretty quickly, It doesn't mean I'll create a chess grand-master in a day or two.
> I do.
I disagree. C comes with a lot of edge cases and subtleties that can surprise even people like myself who 'know' C.
Some examples: Exact semantics of restrict, C99 inline semantics (eg I wasn't aware that it's possible to make a non-extern/non-static inline definition into an extern one with a single redeclaration), effective typing rules and whether it's possible to circumvent them, whether or not it's undefined behaviour to cross boundaries of the sub-arrays of a multi-dimensional array if it doesn't happen in a single expression, ...
With C++ and its typical libraries it is a bit different. It has so many features (templates, classes, streams, friends, shared pointers, unique pointers, distructors, constructors, polymorphism rules, and combination of those) that code gets complicated without using obscure features, just sticking to the standard ones gets hairy and needs someone who knows the whole spec.
This needs to be qualified. Simple compared to what? C is simple when compared to C++, Java, and many other languages. Obviously, C is simple because it lacks syntactic sugar, classes, polymorphism, templating (generics), memory management, etc. So, uh, yeah. It's simple.
The fact that bad code can be written in a language doesn't really make the language non-simple. I could write bad code in any language (often times I do!) - so I'm not sure how any judgment about a language can be made with these kinds of examples. As far as the GCC/VC examples are concerned, they are a non-issue. Shitty compiler keywords are shitty[1]. We know. This is one of the many pains of writing cross-platform code in compiled languages. These examples are contrived and I highly doubt most come from production code.
[1] http://stackoverflow.com/questions/3437404/min-and-max-in-c/...
Programming in assembly generally requires understanding registers, how the processor works, lots of branching variation, a few other ways to loop, memory organization on the processor, and at least tenths or hundreds of different opcodes, some with surprisingly subtle differences between them.
So, no, ASM is not simple the way C is.
Languages have a complexity budget, and C blew its on a squillion different kinds of integer and a bunch of arbitrary-seeming rules to minimize the amount of typing needed to change between them. That was a good tradeoff in its day, where hand-optimization was practical and program source needed to be small. It's not an appropriate language to use now outside of very specific circumstances.
This questions appears twice in the article.
Note that shifting a 32-bit integer (no matter the sign) by 32 is undefined behavior (§6.5.7, 3). I'm using Clang, and its `int` has 32 bits even on a 64-bit system.
C is deceptively simple.
return ({goto L; 0;}) && ({L: 5;});
was valid C. GCC 4.4.7 chokes on it, using -std=gnu99. Is this a C11 thing?Conclusion: I kind of lost interest at this point. :-) I don't quickly know how to get it to compile using GCC, though.
These problems are all optional.
If a programmer wants to write obfuscated code (as in the examples), C allows that.
If a programmer wants to write clean, maintainable and documented code, C allows that too.
What these examples demonstrate is not that C is "hard", but that it's powerful.
Just because C won the battle with those languages, it does not mean we need to live with its design issues ad eternum.
- Proper arrays with bound checking, which can locally be turned off, if required for performance reasons
- Explicit operation for converting arrays into pointers
Because you see, if it is compiler specific, it is not part of the language.
- remove crufty alternate syntaxes such as trigraphs and K&R-style definitions
- a multiple-pass compiler which removes the need for explicit prototypes/header files
• Check for overflow would be awesome indeed.
• Undefined behavior avoids massive performance penalties on hardware that wouldn't match the defined behavior, so it's a feature and unlikely to go away (compilers may warn you though).
Yeah, with luck they will be part of C++17, you just need to wait 4 years for them to be defined and then around 5 more for all major compilers, across all OS to support them.
They are not even being discussed for the next C standard, and thus similarly to C blocks, it will remain a clang language extension.
2. Pedantic but it's nul terminated, null is something either the same or completely different depending on the implementation. Also what would you recommend for non null strings? passing a struct around of *s and size_t len, _that_ is a horrid idea.