The cost of forsaking C
blog.bradfieldcs.com
blog.bradfieldcs.com
Knowing C also helps me write more performant code. Fast Java code ends up looking like C code written in Java. For instance, netty 4 implemented their own custom memory allocator to avoid heap pressure. Cassandra also has their own manual memory pools to try to improve performance, and VoltDB ended up implementing much of their database in C++. I've been able to speed up significant chunks of Java code by putting my "C" hat on.
I would recommend every college student learn C, and learn it from the perspective of an "abstract machine" language, how your C code will interact with the OS and underlying hardware, and what assembly a compiler may generate. I would consider learning C for pedagogical purposes to be much more important than C++.
Do note that these are two very different things. The C abstract machine as defined by the standard is sometimes so different from the actual machines you're going to run your code on, that you get these fine undefined behaviour things that everybody's on about.
Please learn both, indeed, and stress the differences; then highlight the advantages of C. C is a dangerous, but powerful tool, and the necessity of warning students about things should not inhibit learning C.
C is sharp, yes. It's also a universal interface. Almost every language has bindings for C, and C has bindings for almost everything.
In my mind this is the #1 reason to learn C. And it's unlikely that C++ will replace C in this role anytime soon as C++ has many more ways for an ABI to go wrong.
Once whole operating systems are written in Rust maybe Rust will be the most important language but until then it's C. C++ doesn't really count. I don't think you can do any serious C++ without knowing C really well.
Really it wasn't too long ago that the STL just wasn't good enough for high performance code and we rolled our own containers/pooling/everything. Modern C++ is a very much a different beast these days.
I will say if you don't understand pointers you're in for a world of pain in C++.
This reads to me as if you ascribe those attributes to C, which simply isn't true; C++ has an equal (to C) concept of pointer, allocation/freeing (and, with RAII, I would argue superior).
Okay, but is that knowledge coming from coding in C, or it coming from debugging low-level code? I don't know what tool(s) you used to do that debugging, but if I assume for a second that you were using GDB, wasn't the main benefit of your past C experience (in this case) the exposure to tools like GDB, rather than your C knowledge being key in enabling you to dissect that JVM issue?
* How the code is executed on modern hardware and operating systems.
* How various language features might be compiled to assembly
* Memory layouts for things like stack frames, how a simple heap allocator might work, array layouts, layouts for other structures like linked lists, etc.
* The relative cost of various operations and how to write code that is cache friendly.
* How debuggers generally work and how to debug code.
Could these things be taught equally well or better in an undergraduate seeing using rust? Honestly I don't know; I know very little about Rust. I can say that I think C++ is a worse language for this purpose because of the language's complexity and because of features that aren't well suited for the above purposes.
I have heard, for instance, that's it's actually very difficult in rust to write a linked list without dropping to unsafe code. I would consider this a bad thing for the above purposes.
Does anyone know this at this point? With pipelining, out of order execution, vectorisation, register renaming, speculative execution, buffered micro-op fusing and what have you, it's very hard to say general things about how the code executes on the actual hardware.
I'm arguing that it's a matter of degree, not of kind. People make it seem like it's the latter.
Viewing all of code and data living together in a big array is important, I think. You don't understand e.g. dynamically generated thunks and shims without seeing the duality of code and data at the execution level.
Get your cache misses in order(which is easy to understand in the C model) and you're ahead of 95% of the curve(plus you get a 10-50x performance bump as a reward).
It is also hard in C to write a linked list without dropping to unsafe code. In fact, it is hard in C to add frigging signed integers without dropping to unsafe code.
In my experience young devs using C get very confused about wtf pointers are really about and what's the deal with the heap, stack, etc. Whereas they understand the concepts better in assembly despite the verbosity.
I'm not arguing for a deep understanding of assembly but if the goal is to "think like a computer" than at least a basic understanding is incredibly helpful. Or at least less confusing than C.
I actually cut my teeth on TI's TMS34010 asm. Simple RISC (lot's of registers), flat address space, and bit addressable. I loved it.
Then again, I admittedly learned 8080 and then Z80 and x86 first, so any RISC feels too simple; but those who learned any RISC first probably feel the exact opposite.
modern MIPS has "ext" and "ins", which are probably an improvement in that regard.
> plentiful-yet-unhelpfully-named registers
what's wrong with MIPS register names? my only comparison is x86 and that has been awful. I'm constantly looking up which registers are what.
They're not really names, just "addresses" of a different form --- and of course, 0 is not really a register. x86 register names are supposed to be mnemonic, because certain instructions are only usable with, or have shorter forms when used with, the "right" register: Accumulator, Count, Base, Data/Divide, Stack Pointer, Base Pointer, Source Index, Destination Index.
To me the MIPS registers seem pretty logical too, though: $zr for zero, $sp for stack pointer, $a0-3 for arguments, $t0-7 for temporaries and $s0-7 for saved registers. Much better than the plain numbered registers on some assembly flavors.
When I did my undergrad, Ada was the primary language of instruction for first year, with x86 assembler taught as a subject in the second semester.
The idea being to get us thinking about programming correctness from the beginning, then to teach how the machine worked once we had a good foundation.
By second year we were just expected to know C, and coming from assembler, things like pointers were something I never had a problem understanding.
Start high to get a good grounding, go low to understand how everything works, then return high to write code that has a good grounding in how it will be executed on the machine.
If the only goal. The OP listed four.
#2 is a great reason to learn Latin but the "influence of C" is identical to the influence of Fortran or Pascal or Algol and of equivalent magnitude to (what evolved into) Common Lisp or Scheme.
And #4 is true eventually but it only matters when the software is broken.
Here wa my roadmap:
(1) I first learned assembly by learning computer architecture from Tanenbaum's book -- Structured Computer Organization -- and related course at the Vrije Universiteit Amsterdam. This taught me a toy architecture (Mic1) but it give me a rough idea of how assembly worked.
(2) Then, later I took a course in binary and malware analysis and all that we were required to do was to read x86 assembly in IDA and interact with it via GDB -- and using the command: layout asm, which gives a layout for all the register values and upcoming instructions.
And once I have a debugger available and understand it well enough, I can learn quite well since I can make little predictions and check if they are right or wrong, and so on.
So while it's realistic to embrace C for many tasks, it's wrong to convince yourself that "the rest follows" from C.
> the mill uses a novel temporal register addressing scheme, the belt, which has been proposed by Ivan Godard to greatly reduce the complexity of processor hardware, specifically the number of internal registers. It aids understanding to view the belt as a moving conveyor belt where the oldest values drop off the belt and vanish.
But what I really had in mind are systems that communicate by passing messages. Distributed systems certainly have an "architecture", but it spans many machines; and communication occurs via messages (RDMA being an exception).
Even modern CPUs contain non-Von Neumann features like multiple cores, pipelines, and out-of-order execution, so the line gets blurry. To a large extent modern CPUs enable C-style programming with a lot of contrivances to hide the fact that they're not quite Von Neumann anymore. Dealing with the different architecture becomes the compiler's job.
"Thinking in C" hinges on the idea that the size of memory is several orders of magnitude larger than the number of cores, and that you can only modify these words one at a time.
Of course those aren't really CPUs to be programmed with software, they're another level of abstraction down. But they're pretty common and hardware description languages are vastly different from C. The inherent massive parallelism of FPGAs & the resulting combinatorial-by-default (sequential only when explicitly declared) languages requires a very different way of thinking.
Being without GC, I learned the trade offs. Obviously, being able to explicitly allocate and free memory was a big deal. But at the same time, I learned to appreciate the complexity and pitfalls of GC.
Threads are great in Java (as are their parallels in other languages): this is the work to do, put the results here, and everything is cleaned up. Meanwhile, just trying to get the results back from a forked process was CRAZY. Implementing and managing my own shared memory to pass messages back and forth...wtfbbq.
"Think like a computer," not really. More like, "understand what all goes into a thread, an async call, or a garbage collector. Gain more understanding into the corner cases and exactly how much this thing is doing for you. So maybe setting something to null or pre-allocating objects to reduce heap work isn't that big a deal."
C taught me to be patient with higher level languages, to better understand what they were doing for me, be able to reason about how they may be doing these things wrong, and to think critically about how to coax these features into behaving the way I want them to.
Really any lower level language can teach these things. The point here is that C is a nearly universal language that can teach you to reason about these things.
Also, the blog throws in "and C++" in parentheses. C++ is another beast altogether -- it diverged from C a long time ago -- read Scott Meyer's "Effective Modern C++" if you think C and C++ bear any resemblance.
1. You need to know C because C is popular, and
2. You need to know C because C is a lowish-level language.
I can't very well argue with the first point -- indeed, nearly everything fundamental is made in C or the C-like subset of C++ still -- but I oppose myself to the second.
I do think you should learn a lowish-level language, but there are good reasons to make that language something like Ada instead; something high-level enough to have concurrency primitives yet low-level enough to force you to consider the expected ranges of your integers and whether or not you want logical conjuctions to be short-circuiting. Something high-level enough to have RAII managed generic types, yet low-level enough to let you manipulate the bit layout of your records. High-level enough to have exceptions and a good module system for programming in the large yet low level enough to have a built in fixpoint type and access to machine code operations.
Unless you target C specifically for its popularity, there are other options out there, filling the same gap.
In practice Ada programmers, because they actually have a proper method of expressing integer ranges, are going to be much more considerate of such things.
My very controversial opinion is that C is obsolete for new programming projects. For example, everything that C can do, C++ can do and better/safer/faster. Rust is also getting there (if not there already). C is not Pareto optimal anymore.
Rust (and collectivelly all languages that claim to replace C, or just have a significant cult following insisting that everything be rewritten in them) are fine languages on their own merit but I can't see them replacing C any time soon. The main problem with these is that they were designed by people who don't like C. People who do like C would probably love a "safer" alternative to it and similar to it, but languages like Rust just aren't appealing for the same use-cases. The only language that kind of works for this is Go, which was written by people who, on the whole, do like C.
You can also find examples for the other way around, too: Web browsers are all written in (mostly) C++ for example.
Games are another example. John Carmack also admitted that switching from C to C++ was the right thing to do for Doom 3:
"Today, I do firmly believe that C++ is the right language for large, multi-developer projects with critical performance requirements, and Tech 5 is a lot better off for the Doom 3 experience." http://fabiensanglard.net/doom3/interviews.php
> Top 4 kernels are written in C: NT, Mach, Linux, BSD.
Parts of NT are written in C++ btw.
1) It is doing a disservice to the students to tell them that you won't use much of what you are made to learn.
2) A nontrivial amount of people will attempt to use it seriously anyway, sometimes with disastrous results.
3) It usually indicates teacher laziness and/or disinteredness in finding alternative ways to teach.
4) Surprisingly often, when you are not supposed to use X, it is actually useless to learn X. (That is NOT the case in this situation -- there are reasons to learn C -- but the students can't know that unless told so.)
5) There are almost always better alternatives. I have never been unable to find a Y which fulfills the same criteria as X, yet in addition is also usable for serious things. It's just that sometimes you have to look a little harder for it.
In other words, I think it is a teacher's categorical responsibility to find a means to teach what they want in a way that is practically applicable by the students right away. And failing to do so should be taken as a strong hint that what they want to teach may not be the thing they should teach.
I wouldn't say this if I didn't firmly believe that you can teach most things in a way that grants the student immediate practical applications. And I don't say this in a political way -- of course everyone should have the liberty to teach whatever useless thing they want in whatever shitty manner they cn come up with. I'm viewing it more as requirements on any teacher who wants to call themselves good, or tell themselves they are doing their students a service and improving mankind.
Rust (and maybe Go) do seem like very good long-term replacements for C, but imho it's too early to say that. Until there's a "full stack" of software (OS to user-space/desktop), you still kinda need to know C if you want to be able to understand how your whole computer works.
AFAICT, Rust was created to be a safer C, and Go was created to be a faster Python; completely separate domains.
That would surprise me; Go has little in common with Python.
But also because it makes it easier to deal with asynchronous actions and has facilities for higher level programming should there be room for it.
Words escape me.
It's not like books on C stopped being written by 1990.
That's trivially true. Pick anything as your bottom abstraction layer, you can almost certainly find something below. Unless you're working on the foundations of mathematics or trying to come up with theories of quantum gravity you haven't hit the bottom of the abstractions.
It's also not as though C is the bottom layer even for programming. There's assembly. Then HDLs like Verilog or VHDL. Then electrical engineering... You have to stop somewhere. My knowledge spans from some electrical engineering through C up to Python/Java/Lisp levels, but I don't (currently) know much of anything about Javascript or actual semiconductor engineering or more than basic physics...
C is pretty good as a foundation for programmers. It's terrible as a foundation for computer engineers, and it's pretty much at or above the ceiling for what's expected of electrical engineers. It's not guaranteed to remain a good foundation for programmers. Massively parallel architectures like GPUs don't tend to fit well into C's abstract machine model.
This is mostly true today, but I hope that over time we'll see more languages fill the gap between hardware and high-level languages. There's no fundamental reason why that middle layer has to be in C, we just have had a shortage of viable alternatives.
C is much more compact.
I bet Stroustrap occasionally sees some C++ code and says "huh ... I didn't know you could do that."
To be fair, that is not an entirely uncommon thing to say when watching a user work with something you made.
C++, Go, Rust... these are very viable alternatives to C.
It is, but that's not how Go's growable stack works. Go's is grown by the language runtime and relocated when necessary, using its GC infrastructure, so you really can't put non-Go stack frames on it.
For me the cost of forsaking C would be not to experience 100% programming freedom/power (which comes with 100% responsibility). And yes, C also gives you the freedom to do any nonsense :)
I was always fascinated how little resources are needed to run C programs, here an example (not by me): https://github.com/barteksalsa/avr-tiny12-disco/blob/master/... ATTiny12 mikrocontroller with: 1 KiB program flash rom, 32 bytes of "ram", 3 level hardware stack
"The C language combines all the power of assembly language with all the ease-of-use of assembly language."
It's also been said to have been developed with the idea that compilers for it should be really easy to write, so let's just not spec half the language. ;) Freedom comes as a consequence, not independently.
One of the first things I do on a new Unix login is alias the rm, cp and mv commands to -i so they prompt before clobbering files. I spend a lot of my day at a command prompt, but I like my guardrails.
P.S.: alias the rm, cp and mv - that I do too, I think everyone had his 'rm -rf $stuff' moment (-_-)
> no language level support for memory hierarchies
By this, I assume you mean cache hierarchies (for example). What languages do?
A couple of well-known articles/discussions about that universe:
A Forth Story: https://news.ycombinator.com/item?id=3813966
Lisping at JPL: https://news.ycombinator.com/item?id=7989328
I've found it hard to go from there to link up with the more abstract stuff. Not so with python, because the mental energy I spend reading python code is very little, which enables me to move very fast.
Reading sklearn source has been a revelation to me. I can't imagine having to read a similar high level equivalent in C.
We give students four reasons for learning C:
1. It is still one of the most commonly used languages outside of the Bay Area web/mobile startup echo chamber;
2. C’s influence can be seen in many modern languages;
3. C helps you think like a computer; and,
4. Most tools for writing software are written in C (or C++)
[Not really TL, but good context for reading comments referencing the numbers]At least there is a cute image presenting “The C Programming Language” in the context of Plan 9: http://9front.org/img/000000wq9nh.png
https://github.com/ransomw/exercises/tree/master/lleaves/plu...
... only gripe with this article is the couple of "... and/or C++" clauses: i personally hope to be forgiven for forsaking C++ for new projects, because it is not the 1970's anymore, and i suppose there are alternatives for almost every use case as well as externs for plain-C interop with existing C++ libs.
it's more like C+{100} so many are the additional feature-ughs.
and I say this as somebody proficient in C.
You don't have to reinvent every level of the stack and connect it all together(your imagination can fill in the gaps), but this removes the mystery of computing.
So... C99 and C11 aren't things?
From the K&R2's [1988] Preface:
"We have tried to retain the brevity of the first edition. C is not a big language, and it is not well served by a big book".
Well, that line would go out window with a C11 update. Along with, from the introductory chapter:
C is not a "very high level" language, nor a "big" one, and is not specialized to any particular area of application.
Oops; the page count in ISO C more than doubled.
But its absence of restrictions and its generality make it more convenient and effective for many tasks than supposedly more powerful languages.
The absence of restrictions, whereby C programmers were wrestling particular behaviors out of their code with the cooperation from a particular set of compilers, has lately been turned into a license to wreck non-portable code that used to work for decades.
Who wants to update a classic book in the environment in which you have to warn the reader of undefined behavior at every turn.
Sometimes there is just unavoidable complexity for extremely useful features like thread support. Keeping stuff like that out of the standard doesn't necessary make the language simpler if you are forced to rely on other standards anyways ...
Certainly not by means of the abhorrent garbage that has been added to C for threading.
POSIX is a standard. POSIX started as a branch of the Unix C library that didn't belong in the language. So for instance that's where "open" and "ioctl" went, while "printf" and "malloc" went to C.
There is no need for C to start adding its own versions of POSIX requirements (let alone toy versions); it just adds complexity.
Now what it means is that we can end up with some mongrel project where we have POSIX threads being spawned in one module, ISO C threads in another and so on.
'The most recent edition of the canonical C text (the excitingly named The C Programming Language) was published in 1988; C is so unfashionable that the authors have neglected to update it in light of 30 years of progress in software engineering.'
Of course you can't see that 'it' refers to the book when you quote it with half the sentence missing.
Knowing your allocations and performance at a low level can influence usage of a high level language and have a huge impact.
Can't even count how many times I've read the sources of some program to figure out what some random button does (or is supposed to do) -- back when I messed around with blender I used the code more often than the wiki to figure things out and found tons of bugs that way.