Taking C Seriously
subfurther.com
subfurther.com
I can't say the same of... well, any other language I've used in the past 5 years.
Not saying it disappeared, but there is now a bifurcation between C and C++ programmers.
In all fairness, development continued past 1973. Most C programmers today would have some trouble even reading the code from the 1978 first edition of K&R; and the void* pointer, that the author refers to later in the article, was a part of ANSI C (~1990).
It would be nice to know if there are any changes from the April draft, though...
Still, whoa -
> #define cbrt(X) _Generic((X), long double: cbrtl, \
> default: cbrt, \
> float: cbrtf)(X)
Fun stuff.That C is still in wide use after 40 years is a testament to the elegance of its original design but let's not get carried away.
How about the Linux kernel?
And this came without thinking about it... surely in two more minutes or by going to ohloh I could fire a couple more large bodies of non-trivial low-level C code at you.
And don't forget Git! ;)
Please pick gcc to mention or, indeed, any other C program ever written, before libbfd ;)
I tried writing Linux kernel drivers; it was a horrifying mess. Tragically, it would have been easily managed in C++.
Now when can we get telemetry support? ;)
If you are thinking of Return to Castle Wolfstein, that was C as well, using the Quake 3 engine. But id didn't even make that game.
The author is pretty active here on HN. Very well documented design and goals. Also well documented for what problems it does not solve.
I think around 20k LOC. Seems pretty non-trivial.
Assuming that you meant only 'non-trivial': aside from the embedded space, C and GCC still represent the first-tapped resource in many companies. C is terse, well-known, fast and predictable.
I'd go marginally further and claim that, done correctly, gmake and a proper directory hierarchy remain the most effective way of organizing and maintaining a large software project.
People have codebases that are millions of lines of code in C++ that I don't think would even be possible in pure C.
I have lots of reasons why this is true, many of which are specific criticisms about C++. There is one issue that is industry specific: Everyone working in embedded has either come from the bottom up (deep firmware in assembly or machine code) or the top down (Java in school), and when you get a dozen people where half of them treat C++ like "C with classes" and the other half treats C++ like it's Java, it's a guaranteed disaster.
For low level systems work, C++ can be considerably more difficult than C to organize and port large programs.
It isn't "C++ is always worse than C" it is "C is often better than C++ for a given task, depending on task"
Each task would come with a modification, and an error to make in implementing the original task. You'd recruit undergraduates, who hadn't used the language before. Some would be given the original program to modify, others the broken version to fix. You'd measure how long they took to do it.
This way, language communities that gamed the machine benchmarks would pay a price on the human ones.
As well as C refusing to go away, there's a noticeable surge for C#, Objective-C and Lua, and a substantial erosion in the popularity of trendy languages like Python and Ruby.
Then you have luajit - also a small, easy to embed solution, that approaches standard "C" written code, and it's main weakness, from what I understood (Mike Pall had a post about it) is it's garbage collector.
On top of that, ffi bindings make "C" calls extremely fast, when the jit is active, and not very bad when the interpretter runs (but slower).
And the lua/luajit license allows you to embed it in your application, without the need to share source code.
Mike Pall is working on ppc, arm versions (and I think there might be already ppc jit).
Not last to forget - awesome community (comp.lang.lua).
What would be really interesting is to see someone highlight specific cases where this approach ultimately fails to measure up in performance with using pure C.
I would think that the LuaJIT approach would be tens of times more maintainable for a sufficiently large application, so it's really imperative here that we ask 'Why not?'
One area which is not easy translatable is OpenMP (www.openmp.org), inlined assembly, and SSE packed floats. But that's okay, and even then there is probably a better alternative - a language more suited to such tasks, instead of "C" - OpenCL (www.khronos.org) or DirectCompute.
But for general coding, it's very very good.
For example, read this: http://www.jwz.org/doc/gc.html
> In a large application, a good garbage collector is more efficient than malloc/free.
My point isn't necessarily to disagree with the article, but to point out that the article has practically nothing to disagree with. It has no substance.
To elaborate, the actual problem with C is that you have to deal with ownership semantics manually. In languages with things like uniqueness typing or automatic reference counting that problem goes away. Common to all these languages and all GCed languages is the need for memory management--clearing references or maps or just generally indicating (with the language's particular idioms) its lifetime. Sometimes in a GCed language all that ownership gets untangled for you for free, but this may actually be a maintainability hazard. A trivial change might suddenly start retaining objects forever. See Haskell, where a seeming perfect program may suddenly gain space leaks upon mere removal of, say, a print statement.
Some GC implementations might be faster in practice if they can move around memory and improve cache locality and reduce fragmentation. On the other hand, some introduce long collection pauses (often an issue with C# on XNA for example), and the default malloc() on many systems is very slow and tends to fragment. But "sufficiently smart" GCs avoid these issues.
Ideally the solution is a hybrid approach: for data with obvious lifetime (bound to a scope or a certain area of the program execution) you want to actually show those intentions in the code or types. For short-lived objects that you're working with or object graphs you want garbage collection. This is what generational GC simulates, but I always find it asinine to fiddle around with references (possibly having to null out) and deal with non-deterministic deallocation when the lifetime is clear. Sigh. One day.
> for a large, complex application a good GC will be more > efficient than a zillion pieces of hand-tuned, randomly > micro-optimized storage management
vs:
> Note that I said a good garbage collector. Don't blame > the concept of GC just because you've never seen a good > GC that interfaces well with your favorite language.
Ok, so we're comparing 'state of the art GC' to 'zillion pieces of hand-tuned, randomly micro-optimized storage management'. Astoundingly, we come to the conclusion that GC Is Awesome and MOAR Efficient.
It's clear that you can make great code in C, if you're a good programmer and get enough time to plan things like memory management done right. Enough examples of that.
But when you don't, C is a terrible language. It is very verbose, it makes you repeat yourself and the macro system is so error-prone to the point that it's usually forbidden to use. So you have to resort to custom code generators (we have at least three!).
Most companies don't want to be in the business of worrying about buffer overflows and memory leaks and segmentation faults. They want the requested functionality implemented robustly ASAP. If there is less code to be written, there is less to test and worry about, so a high-level language that takes those concerns (largely) away is a great help.
So my advise would be to limit using C to the performance-critical parts (found using benchmarking), and only make developers work on that which understand every detail about C. Also, give them enough time to test every nook and cranny threefold.
And later, in 2000, he states "Today, I program in C." http://www.jwz.org/doc/java.html
JWZ is an entertaining read, but unless he's updated in a post somewhere in the last 11 years, it isn't clear to me where he comes down on C at the moment.
Btw, I program in C.
At the moment, his main focus seems to be his bar, so perhaps he isn't the best person to ask about the relative merits of C versus modern garbage-collected runtimes.
Not really the point, though: The point is that blanket statements about the inefficiency of gc amount to superstition, and should not just be let stand.