How can C Programs be so Reliable?
tratt.net
tratt.net
There's no reason why you can't write correct C code, or correct assembly code for that matter. The challenge is to do so without wasting a lot of time: Any amount of time that you spend consciously thinking about correct memory management or hand-optimizing your opcodes could probably be spent doing something more important, unless you are working on one of the few problems where that kind of optimization is actually the bottleneck.
Of course, the flip side of having to think about every layer is that you get to see and potentially tweak every layer. It's nice to work on something transparent. It's nice to know what is going on down there among the sockets and the buffers. I've been thinking about practicing some C for just that reason, and it seems to be why the OP likes C. But I don't anticipate being very efficient when writing my own web server in C. My website will be better if I just install a big pile of other people's C and get on with designing or writing.
Crucially, this linkage does not depend on whether there are libraries for what you're trying to do or experience level in the language. If coding a particular feature takes 1000 lines of code in C vs. 100 in Python, you'd have to assert that a programmer can write (and maintain) code 10x faster in C than Python. Since roughly the same amount of conceptual energy goes into each line of code, this is tantamount to the assertion that a programmer writing C can think through concepts 10x faster than when the same programmer is writing Python.
There's a reason the field invented higher-level languages, and it doesn't entirely have to do with novice programmers.
I think the overarching point is experience with the language. I for one have seen and written incredibly long and convoluted programs in a high level language which was largely due to a weak understanding of the massive API's that come with it. And it would have taken roughly the same amount of time to read and understand what the API functions do and how to use them as it would in C.
In the end, it's all relative and it's totally dependent on experience with the language and what you need for your application.
So with this definition of efficiency, less lines of code does equal more efficient code.
The problem is, your code will appear to work fine even if you don't complete the chore. It isn't until your program blows up that it'll even occur to you that there was more work to be done.
Would love to see a reference to this.
But come on. While let-me-Google-that-for-you requests are annoying, they are in toto less toxic to threads than comments like yours; at least the lame question generates a factual answer.
may be hard, but not more difficult than creating the run time for any other language such as Java or Python -- and they are all written in C.
PS: PG is in love with LISP in large part because he deals with problems well suited to the LISP domain. But, if he had been writing drivers he would have gone in another direction. The real lesson is if you wanted to write a great XML editor use something in the LISP family or something that's closely related to it not that LISP wins on all fronts.
Python doesn't produce smaller programs compared to C http://plg.uwaterloo.ca/~migod/846/2011-Winter/projects/Simo...
Sure, we throw in the almost 100%-Python Bazaar, at 200k, which CVS beats. But CVS vs. any of these other source control systems is not really a fair comparison, IMO, and it's still blown out of the water by Mercurial.
I'm not saying that this proves that Python code is smaller (comparing source control systems to each other is completely unfair, since they differ so much in feature set, platform target, and code quality), but it certainly does not disprove it.
That means that when you lay out your program the most important parts (memory map and failure modes) are clearly visible.
IF you are a good programmer.
And that's the reason there is an obfuscated C contest, if a C programmer sets his or her mind on being deliberately hard to understand that same power can be used against any future reader of the code. Incompetence goes a long way towards explaining some of C's bad reputation. You can write bad code in any language, but none give you as much rope to hang yourself with as C (and of course, C++).
the building blocks are so simple and transparent
that you can follow the thread of execution with
minimal mental overhead.
I do not agree.I've seen plenty of code that does weird things with pointers, like passing around a reference to a struct's member, then to retrieve the struct decrementing a value from the pointer + casting. Or XOR-ing pointers in doubly-linked lists for compression. And these are just simple examples.
I've seen code where I was like "WTF was this guy thinking?".
My biggest problem with C is that error handling is non-standard. In case of errors ome functions are returning 0. Some are returning -1. Some are returning > 0. Some are returning a value in an out parameter. Some functions are putting an error in errno. Some functions are resetting errno on each call. Some functions do not reset errno.
Also, the Glibc documentation is so incomplete on so many important issues that it isn't even funny.
Yes, kernel hackers can surely write good code after years of experience with buggy code that they had to debug.
But for the rest of the code, written by mere mortals, I basically get a headache every time I have to take a peek at code somebody else wrote.
Yes, that happens. But I've seen that in COBOL, Perl, Pascal, Java, PHP and in Ruby as well.
> In case of errors (s)ome functions are returning 0. Some are returning -1.
That's not a feature of the language.
That's not a feature of the language.
Well, yes, but it's kind of nice when you've got exceptions with stack traces attached.Some people don't like exceptions, but I do.
Serious C projects tend to come up with this stuff on their own, often with better adapted implementations than the "plain stack trace" you see in higher level environments. Check out the kernel's use of BUG/WARN for a great example of how runtime stack introspection can work in C.
That's not the worst that exists in C :-) Let me quote a dietlibc developer from http://www.koders.com/c/fid1639C203A2255EB1FA11DC6A68D74FEB2...
/* Oh boy, this interface sucks so badly, there are no words for it.
* Not one, not two, but _three_ error signalling methods! (*h_errnop
* nonzero? return value nonzero? *RESULT zero?) The glibc goons
* really outdid themselves with this one. */That's not a feature of the language.
The inconsistency is a natural, expected, unavoidable result of the language forcing, er strongly encouraging, use of an unsuitable error reporting mechanism ("find some value in the range of the function's return type that isn't in the range of the function, and use it to indicate an error"). This wouldn't be an issue with exceptions or tuples / multivalue return like some languages allow.
I thought that references are a feature of C++, not C. Personally, I never really got references... They are just a kind of magical pointers that programmers can forget about, but they make the code much less readable and can interact in funny ways...
Of course C has references because C has pointers. References in C++ are just constant pointers.
This is incorrect. You cannot have a reference to nullptr, for example. Pointers and references are different beasts, nowhere in the C standard does it refer to pointers as references. The underlying representation in compilers does not imply equivalence.
C is a very handy portable assembler, though.
http://james-iry.blogspot.com/2010/09/moron-why-c-is-not-ass...
(Why yes, that does sound like something badly in need of refactoring. And illegal reliance on implementation details.)
A pointer type describes an object whose value
provides a reference to an entity... int i;
/* Iterates over everything except the last n elements of array... right? */
for (i = 0; i < length - innocent_little_function(); i++)
do_something_with(array[i]);A function call in a for loop's conditional doesn't fit your definition of wacky?
I probably should have used sizeof, even though that doesn't make sense there.
that's not weird, that's a pretty standard way to enqueue structures on singly/doubly linked lists... it's made somewhat prettier by offsetof/CONTAINING_RECORD though
I mean, can't you pass a reference to the whole structure instead? I prefer pointers to void* to the whole thing, with a normal cast later, instead of seeing pointer arithmetic.
I'm not a C developer, I just play around -- I've seen for example this practice used in libev, passing around extra-context along with the file-handler in events callbacks.
That seems really ugly to me, as they could have added an extra parameter for whatever context you wanted to be passed around.
For example, if you want to trigger some operation after some number of second elapses, you can preallocate some structure and then just populate it and fill it in when the timeout fires. Timeouts are usually just signals or other contexts where you have no way to return failure, so it's important that it be possible to always handle that case correctly without the possibility of failing.
one advantage is it produces a generic linked list API. you can write routines to traverse, add, and remove elements from the list without caring about the structure of data stored in the list. if you use offsetof, you can also have the list data for a structure at any position inside of the structure instead of the beginning. some systems do that so they can store header information at the beginning.
you can also have elements enqueued on multiple lists. you might say that if you're doing that, you have bigger problems, but sometimes shakespere got to get paid.
Another way to think about all the offsetof() stuff is that it's emulating multiple inheritance in C. You can think of structures as inheriting the "trait" of being a participant in a container; the "pointer-arithmetic-and-cast" idiom to move from a container entry to the corresponding object is isomorphic to downcasting from the trait to the object that contains it.
Interestingly it is not possible to express this pattern in a generic way within Java's type system.
You can write obfuscated code in any language. The point about C is that the mental model is very simple. There's no magic happening anywhere, so if you can parse the language, you can figure out what's happening line-by-line pretty easily.
In C, what you see is what you get. :)
#define int doubleHad me laughing though :)
Without exceptions, the advantages of C++ are not that great compared to the hassle it needs to get running in kernel mode like dealing with name mangling, static/global constructors, etc and hassling with compilers.
Incompetence is REASON for C bad reputation(if any).
I think C is one of these "other languages" because "here be dragons".
I watch a lot of people bang out C++ code as if it's totally safe, and fail. I see a lot of people hammer out C# code and say, "to hell with you, you don't even have .NET!" And so on. But today, when a programmer sits down and writes a C program they must sit and think out what they're doing and why—with no abstractions like OO to make an easy solution.
There are so many ways to blow your head off in C without knowing you left the opportunity in the program, that it forces a competent programmer to think differently about how they code. And a newbie? Well, if they aren't scared stiff about blowing a hole in their system, they should be! ;D
And C doesn't change often, unlike other languages.
I'm no C programmer, but I've seen C code for years and translated it into whatever language I'm using at the time. I have tremendous respect for UNIX/Linux, and a great many C-powered programs. Thanks for your work on them, guys and gals.
That's true for any programming language. Sadly, far too often, programmers are unable to afford taking the time needed to think about what they are doing or understanding what happens under the hood of the libraries they link against.
Not just incompetence. Also bad language choice (usually due to legacy).
If programmers don't get enough time to properly test and review the code, which needs to be done very thoroughly in C, it's easy for even experienced developers to shoot themselves in the foot.
C is very good (let's say irreplacable) for low-level hardware and OS code. This is code that needs to be verified and tested very well.
On the other hand, using C for run-of-the-mill business projects or higher-level stuff on a tight deadline can be a very bad idea. It results in a lot of overhead for programmers to think about the details of error handling, buffer sizes, pointers, memory allocation/deallocation and so on, especially getting it right for every function. It is a recipe for screwups.
In this case it is very useful to have garbage collection, bounds checking, built-in varlength string handling, and other "luxuries" that modern languages afford you.
The first fundamental purpose of any programming language is to provide abstraction via functions. This implies that following the thread of execution is never easy and the blocks are never simple. It's pretty much a wash, with special mention for languages in the Hindley-Milner family.
The second fundamental purpose of any programming language is to provide specification abstraction via replaceable modules. This is where C fails. It is common practice in C culture to not specify interfaces in depth (we are all good programmers, aren't we?) and the implementation via manual virtual tables makes it painfully difficult to find the specific implementations in the code base.
The original idea of the language (or at least a major part of it) was to be a portable alternative for the many processor specific assembly languages in use - rather than having to write the same functionality for each one, you could write it once in C and then compile it for each platform. If that's your aim, then you will end up directly manipulating memory, and you open yourself up to that whole class of errors - memory leaks, array overruns, pointer arithmetic mistakes. All C gives you is portable access to how processor hardware works, with a few conveniences (y'know - function calls).
If you want to protect against these problems you have to add some extra layers of abstraction between the language and the underlying hardware, and that comes at a cost. That cost is mostly performance, but thanks to Moore's law these days that is a much lower priority hence the abundant use of higher level languages - Java, Python etc.
My point is that C is how it is _on purpose_. This direct access to the hardware comes with some downsides, but they aren't 'flaws', they come hand in hand with the power.
The real reason that most C programs in daily use are so robust, is because they are ages old. Many, many man-years have been invested in the production of e.g. BSD, unix tools, POSIX libraries, and even web browsers and word processors.
Why do we use Javascript and even PHP to program web-applications? Because we need fewer lines to get the same result. Moreover, given the correlation between number of lines and number of bugs, shorter programs are better. If we had been limited to C "web 2.0" would have been decades away.
C allows you to write interpreted languages that execute with a speed high enough to afford you more abstraction.
Let's give credit where it's due, but really C could have been Ada in any of those examples and the results would have been about the same (especially in PHP's case).
$ ls -1 git-*.sh | wc -l
25
$ ls -1 git-*.perl | wc -l
9
Compare this to the built-in commands: $ ls -1 builtin/*.c | wc -l
92
or all the C files: $ git ls-files '*.c' | wc -l
306At least with Java, pointers/references/objects are either null, or valid. Uninitialized references won't compile, and a null dereference blows stack at point of first use.
Having said that, I like pointer and bounds checking, but wish Java had significant memory management options, for times when you were willing to trade some safety for speed.
This is straightforward enough in C: http://gaiustech.wordpress.com/2011/09/09/segmentation-fault...
A segfaulting program will dump core, which the programmer can use to get the stack trace. I consider this to be better UI than printing the trace at a likely bewildered user.
$ ulimit -c
unlimitedSounds like a feature.
It is amazing that the people who rely on high-level languages think they can stomp on lower-level languages like this, without even realizing that virtually _all_ the features they talk about are made possible by low level languages and are implemented in them.
New C programs get written all the time. They work. They move your world, every day, like clockwork.
Actually, no. When the documentation says
This function shall fail if:
[EFOO] Could not allocate a bar.
it doesn't mean that this is the only possible failure; POSIX states that functions "may generate additional errors unless explicitly disallowed for a particular function".Except in very rare circumstances, when you make system or library calls you should be prepared to receive an E_NEW_ERROR_NEVER_SEEN_NOR_DOCUMENTED_BEFORE and handle it sanely (which in most cases will involve printing an error message and exiting).
if ((fd = open(myfile, O_RDWR | O_NONBLOCK, 0644)) == -1) {
switch(errno) {
case ENOENT: case ENOTDIR: case EACCESS:
case ELOOP: case ENAMETOOLONG: case EPERM:
warn("Cannot open file");
goto choose_file_to_open;
case EISDIR:
if (chdir(myfile) != 0)
warn("Failed to enter %s", myfile);
goto choose_file_to_open;
case ENXIO: case EWOULDBLOCK:
enqueue_open_callback(myfile);
return;
default:
err("Cannot open file");
}
The above is well-documented in open(2); compare, for instance, http://docs.python.org/library/os.html#os.open.(Also, open(2) has all these options for a reason; think symlink races. Python's open() is not sufficient.)
default:
err("Cannot open file");
case.I'm not sure what era the author is referring to, here. In the late 80's, Turbo C broke the price barrier for a decent MS-DOS C compiler at the $79-$99 price range. Shortly after that, Mix began offering their MS-DOS Power C compiler for $20. Tom Swan's book "Type and Learn C++" provided a tiny-model version of Turbo C++ on a disk provided with the book.
The GNU ports djgpp and GCC were available for MS-DOS and Windows in later years.
> the culture was intimidatory;
I'm again wondering what time-period he's talking about. When I started learning C in the late 80's, most of the trade magazines were full of articles that used C as the primary language for whatever programs or techniques were being presented. Dr. Dobbs Journal was full of C code. Before Byte quit publishing source code, one could find a fair amount of C there. Of course, the specialty magazines like The C/C++ User's Journal and the C Gazette contained nothing but C and later C++ code.
> This is a huge difference in mind-set > from exception based languages,
Yes. C is a language that was designed two decades before Java.
At first, I was really taken aback by the author's take on C, but as I tried to digest why he has these perceptions of the language, I ventured to guess that a number of developers who came of age when languages with more modern niceties were available probably also have this view of C. From the perspective of someone who has been able to use more modern languages, C must seem like a rickety bridge that could be dangerous to cross.
A number of points that Mr. Tratt makes, though, pertain to the programmer; not the language. Certainly there are library routines that allow for buffer overflows, like gets(). It's been known for quite a while ( since the Internet worm was unleashed in 1988? ) that fgets() should be favored so that buffer boundaries can be observed. Certainly people writing their own functions may not write them correctly, but this is a matter of becoming conversant with C. It's a matter of attaining the right experience.
Also, as I recall, to be a certified ADA-compliant compiler requires having a ton of libraries.
The notion that it's some kind of commitee-created monstrosity seems largely to have come from ESR's wildly inaccurate writeup of it in the Hacker's Dictionary. It's actually a fairly small and nice language.
The object oriented extensions didn't follow the "standard" Java "." syntax, though, which probably hurt Ada 95's uptake more than it should have.
Would I write a complete web service in C? Probably not. Would I write a fast image manipulation/modification library for that specific website if needed in C (or C++)? Probably -- because I like the performance gain when I'm converting 10.000 images.
I love the fact that you can just build components in different languages and then glue them together so you can build awesome products.
It starts to pay off when you write that package either as a service with a large number of users or if you make a general purpose library for inclusion in lots of other programs, especially if they are written in other languages.
The reality is that doom and gloom about C is overrated. Sure, you can shoot yourself in the foot easily, but most competent programmers will do just fine. That's been my general experience and I don't think I've spent my career surrounded by rockstars :-)
It happens.
Added in edit: Just for reference, I didn't down-vote you. Not least, one can't down-vote replies to one's own submissions or comments.
I see all my posts sink. The only reason I bother to post is it gets maybe 50 people to read and that is better than 0 I figure.
In C, errors don't trickle down and you need to deal with them in each level of abstraction, which can be totally useless and time consuming.
Time consuming, yes, absolutely. "Useless" I don't understand at all. It's structured exceptions that more often seem useless to me.
The more layers of code an exception bubbles up (or trickles down) through, the less the exception handler can know about where it happened, why it happened, or what the resulting state of the program is. Very often, the only "handling" that can be done is making a report of the exception.
It seems to me that the most useful exception handlers, the ones most likely to actually salvage the situation and allow the program to continue to work, are the ones that immediately follow an exception-throwing call, the ones that don't allow any trickle-down.
But those are degenerates, of course. They're functionally equivalent to C-style error return codes.
One problem with C-style error handling is that having error handling at all levels makes it impossible to reuse code without tweaking it. Say you handle an error by printing a warning message to stderr, now you can't reuse that code in anything that doesn't want error messages printed there (maybe it doesn't want errors printed, or it needs to localize them, or it's using stderr for something else like in strace). So error handling in C kill reusability.
Another problem with C-style error handling is that the number of error types increases as you pass the buck up levels, but C doesn't have any good way to express more than one type at a time. Say one function returns true or false and another returns an error constant like errno. When one function calls the other, what do you return at the top level? You could map one to the other, but you've lost maybe important information about the error. So error handling in C kills composability.
And a third problem is documentation. With no standard error types that are known to the compiler it is rarely possible to tell the caller they forgot some error handling.
Exceptions address the first two problems, and checked exceptions the third one. A common misunderstanding of Java checked exceptions is that they force the caller to handle the error, when what they really do is force to caller to document the errors it can generate.
The declared exceptions (and exceptions vs errors) in Java are a nuisance, IMHO. For example, use any kind of "dependency injection" ("strategy" pattern, driver plugin(s) for other, older languages) via reflection, and suddenly all of your exceptions have transmogrified into errors :-(
Also, setjmp: http://en.wikipedia.org/wiki/Setjmp.h if you really really really like exceptions and want them in C.
Heavily used ones are as reliable as other heavily used programs, but barely any C programmers even use clang (static analysis) or even the elderly lint and its more modern cousins.
This on top of half of people calling themselves C programmers are really C++ programmers (they really are quite different how you use them in the correct manner), I don't really think he's correctly summarizing the field at all.
edit: I have been a C programmer for most of my career, including embedded linux, cli linux (including research robotics), and C-Servers to communicate to the above
I'm not some guy who just knows python and bitches about "the hard compiled languages" (although I do like python and ruby and objective-C).
Oh, if only that were true. I've seen some not-so competent programmers churn out lots of C code (and then move on to C++ in order to do some real damage)
The only low-level language that has any innate claim to reliability is C++ with proper use of the RAII idiom.
Or rather a very carefully chosen subset of C++. See http://yosefk.com/c++fqa/exceptions.html#fqa-17.3 for some of the problems.
By the way, what about ADA?
Be aware that it is a b&d language.
I think some of Yosef's advice on C++ is out of date WRT to the C++11 standard too. shared_ptr is now the recommended smart pointer, for instance.
My goodness look at that, how can those tightropes be so reliable?
The point is that, as mentioned elsewhere, C is "build your own system" level -- VERY LOW (level). The overhead of wrapping primitives with strategies for your app is minimal, once you know it's an issue.
Thank you for allowing me this nostalgic indulgence, hackernewers. I know for at least a few of you, it will resonate.
But C is also ugly because Duff's device is possible. Think about it from an optimizing compiler standpoint. It takes a serious amount of effort to turn Duff's device -type control flow into an intermediate representation that can be somehow optimized. Now compare that to a language that is based on some form of extended lambda calculus.
Arguably, indeed. The analogy is quite simple - a gigantic roulette wheel with 2^$membusbits slots, except the numbers are sequential. The ball is the pointer and pointer arithmetic involves moving the ball around the wheel.
When you're referencing that row in Excel, you don't copy around its value, but rather, you copy the address of the value. That way, if you change the value, any other cells that reference it will also fetch the new value.