Teaching C (2016)
blog.regehr.org
blog.regehr.org
https://nostarch.com/Effective_C
It is written by Jens Gustedt (https://icps.icube.unistra.fr/index.php/Jens_Gustedt)
In my own Masters program we had a class on software systems which heavily featured both practical coverage of C and lots about debugging and observability. One of the most aggressively _useful_ classes I’ve taken.
Already in the mid-1990s it was possible to use C debuggers as poor man's REPL, have tracepoints, scripting debugging sessions, visualize data structures, track down memory leaks and corruption.
Also ignore the stdlib as much as possible, especially the string functions.
The C stdlib APIs are essentially ancient leftovers from the K&R and early C89 era and should have been modernized when C99 came around. Especially for newbies, paying too much attention to the stdlib may ingrain bad habits.
All the wide string stuff (wchar.h and wctype.h) as well as locale.h is all pointless today.
Everything related to filesystems and file IO is just the bare minimum that's acceptable for simple UNIX-style command line tools, but not really useful in modern applications (most notably, any support for non-blocking IO and memory-mapped-files are missing).
Arguably, complex.h doesn't even belong in the stdlib (I wonder why such an esoteric feature was even considered).
Better support for custom allocation strategies would be nice (e.g. each function which allocates under the hood should accept an allocator argument instead of being hardwired to malloc).
I'm not even asking for a more feature-rich stdlib, a stdlib2 with the obsolete parts removed and modernized IO and memory management functions would go pretty far. But then of course there's the problem that the C stdlib is also accidentially the de-facto system API on some UNIXes.
Sure but POSIX supports that and it may as well be the stdlib for those of us in *nix land.
Including a more complete set of common operations (particularly string and vector operations) would be helpful as well.
etc.
It sounds complicated but the VM, in Python, is basically:
r = {}
mem = []
boot(mem)
pc = 0
while True:
op, args = decode(mem[pc++])
if op == “+”:
d, a, b = args
r[d] = r[a] + b
elif ...:
...
The “source” which gets “compiled” can also just be Python function calls that emit op codes — no need to get bogged down with lexing and parsing!The thing you focus on is the semantics of representing a function call (and return) in machine code and how to manage the stack. It made C make a lot more sense e.g. defining locals up front so the compiler knows how much stack space to use. I wish I could remember more — this was all from teaching A-Level Computer Science a few years ago.
I find C easier to understand than the algebra of other languages such as Haskell, Ocaml, Rust because of how C statements correspond to assembly.
Memory is just a big grid of numbered locations and computing is logistics between grid locations. Execution is transition between memory locations.
LEAQ or "&" (C addressof) is an address calculation or LEAQ for one of those grid locations.
True in the 80s maybe. Not remotely true now.
objdump -b binary -Matt,x86-64 -D main.bin -m i386 ;
(note: this command places the assembly and C code side next to eachother, I think you have to compile with -g debug flag with gcc)You can work that the assembly implements the semantics of C's virtual machine model and it is this that I find easier to understand than some algebraic systems. I kind of think of struct->ptrarray[indx]->struct.something corresponds to the calculation of ultimately a single memory address even if there is various shift lefts or adds or ors.
Java has its template interpreter which is interesting to read about and there is copy and patch JIT compilers.
Sun's Forte Pro was one of the first IDEs to have such capability for Java.
https://godbolt.org/z/hjW4T934q
And that capability is not remotely unique to C - look at the list of languages Godbolt supports! Not all of them support linking source code to assembly but a large number do: C++, Rust, Zig, Ada, even Dart.
The book contains some production-quality data type implementations that aspiring C programmers ought to read.
It also primed us for more involved classes, like operating systems, compilers, databases, and so on.
But I have to say, it is a language that students really need to fundamentally know if it is going to be used down the road. Sounds obvious, but when I TA'd in more advanced classes (i.e. anything after "introduction to programming"), many of the students that were struggling, were really just struggling with the language - not necessarily the CS concepts. A lot of them were even students that had programming experience prior to enrolling, but had mostly used higher level languages.
Going through college I saw exactly what you are saying when it comes to struggling with the language. However most of the individuals I ended up assisting on top of my studying, one-on-one tutoring (if you will); the most glaring thing I noticed was not just a lack of understanding with how the language worked, but more importantly to our discussions, they didn't seem to understand how the fundamentals of computing. How memory is addressed, what exactly are pointers, how de-referencing works in actuality, heap vs. stack, passing to functions to returning values, etc. Once I would walk them through the basics it seemed that they started to understand that while they weren't able to really work effectively within the lower-levels, they would eventually begin to intuit things for themselves much more easily.
I don't believe that all that is required to learn how to write code, however if you know how a machine works, you tend to be able to figure out how to treat it or use it. Like with a manual transmission, if you watch visually how a clutch engages, you'll tend to not burn it as much while letting it do it's job, since you know the fundamentals of how it works. It's no longer the "abstract" or "too complicated to understand". Giving people that knowledge gives them the ability and the confidence to delve deeper and genuinely engage with new ways of doing the same thing they've been doing before.
The actual problem is doing a half baked job, teaching it as if PDP-11 were still current, without any consideration for safe programming practices and modern tooling.
Teaching C (2016) - https://news.ycombinator.com/item?id=32798826 - Sept 2022 (94 comments)
Teaching C (2016) - https://news.ycombinator.com/item?id=18334476 - Oct 2018 (211 comments)
Teaching C - https://news.ycombinator.com/item?id=11668396 - May 2016 (152 comments)
The explanations were well-done. They gradually build up the examples. They actually did warn about common gotchas even then. Their examples mostly worked today with only minor tweaks. I used it with Clang’s and MS Visual Studio’s static analyzers.
https://www.gnu.org/software/hurd/
You can debug drivers on the fly with gdb, something much more difficult to do under OpenBSD, NetBSD and Linux-(libre).
Also, you can run Hurd inside Hurd, a la chroot. But no root permissions are needed.
https://www.gnu.org/software/hurd/community/gsoc/project_ide...
Yep, C is the Unix and Unix-like ABI, and an odd form of C it's under 9front/plan9's ABI/API too, but this makes the userland much more powerful than the typical permission-isolated Unix system where the user has near no capabilities even for mounting non-external media (by default).
While also teaching students that static code analyzers are very far from infallible and can waste a great deal of developer time with false positives.
It happens about once every month or two that someone posts on the sqlite forum, "I've found this serious bug..." and, after some back and forth, it's revealed that they used a static analyzer or, even worse, just pasted a single internal function from the sqlite core into ChatGPT and asked it to analyze the code for them (completely free of any calling context). Roughly 9 times out of 10, static code analyzers are flat out wrong in their analysis of that particular source tree. It's very rare that a report in sqlite which stems from a static code analyzer is actually correct.
The advice given to such posters is invariably: if you can demonstrate code which tickles the bug your tool is claiming to have found, it will be treated with priority. Until then... not so much.
One answer might be that although students might spend most of their later professional lives in higher-level languages, they should have some experience in a low-level language.
The ideal such language would not have memory management (i.e., not Java / Python), no memory management help from the type system (so no Rust), and no fancy facilities to build abstractions (so no Rust or C++). This language should, ideally, still be a little more pleasant to work with than assembly (so, no assembly). For historical reasons, we would still like unchecked array accesses and null-terminated strings.
Now C is an answer to this question, but it need not be the only answer. We could instead teach students a fictionalized version of real C. So nice 2-s complement behavior on overflows, buffer overflows segfaults and other bad behavior, but mostly of the predictable kind. We could define undefined behaviors as the instructor sees fit.
As much as we can, we can try to make this fictionalized language agree with the semantics of gcc -O0. We can then study gcc -O0 -S itself as an empirical artifact, and use it to understand x86/x86-64/ARM/MIPS assembly.
Finally, while we're at it, we can use the Bourbaki dangerous bend symbol (https://en.wikipedia.org/wiki/Bourbaki_dangerous_bend_symbol) to repeatedly remind students that what they're studying isn't a real language, that real C has sharp edges (including undefined behaviors), and that things like ASAN / Valgrind / ... exist.
I worry that this article is confusing that we'd like students to eventually understand (some understanding of C and an appreciation of the fact that it is an important language which still needs to be handled with care) with the didactic process of reaching that understanding. Wittgenstein's ladder is a thing.
The "sharp edges" of C in either user mode or kernel mode when running on an x86 or arm platform are true implementation semantics that you must deal with while trying to realize genuine "computation."
Programming languages are just collections of human shorthand for managing these semantics. Which makes me feel that inventing this language would be putting the cart before the horse.
Agree. Xv6 is the reason why I used C. I learnt more about various OS concepts (process scheduling, file system etc) by hacking Xv6 code. Linux is too complicated for my taste.
As a mobile app dev, I don't use C. Well.... I do. Very ocasionally though. Only for interacting with specific Android hardwares.
Knowing C also helps in understanding a lot of other languages.
https://www.reddit.com/r/learnprogramming/comments/18zz72t/d...?
If you know of a better place to ask such questions, let me know.
As such I never understood what is so hard about pointers, we were already close friends before learning C.
The first thing learned with assembler is pointers and addresses. So I don't know how hard it would be to get them coming from other languages.
But I did first learn to program in Basic and Fortran, which gave zero insight whatsoever when subsequently learning assembler. I was pretty baffled by it for a while, until I suddenly "got it".
Diving into Z80 was a matter of performance, not additional features.
> I’ve always taught C as a side effect of teaching operating systems, embedded systems, or something along those lines....One might argue that we shouldn’t be teaching C any longer, and I would certainly agree that C is probably a poor first or second language.
So this would be an upper-level class: an introduction to C programming for people with Python/Haskell/etc experience and a decent general understanding of computer science.
Has anyone read K&R and regretted it? I am working through all the exercises and have really enjoyed the process.
I am their first experience at university and for some of them C us their first programming experience.
In this case my only advice is just keep repeating the same subject in more than one lesson.
What is suggested into the article seems to me more about "how do I run C in production".
But good suggestions anyway, I'll review them to see what I could use if it makes sense maybe when we do lab.
The thing to keep in mind is that *you* are the one throwing the razors and knives from the moment you go beyond printf("hello, world!\n");
I'd never recommend a company to start building anything in C. In term of the various team sizes and staff churn over the lifetime of a company, it's too risky.
But afaik most companies, particularly the growth hacking ones, are either Go or Kotlin or at that level of abstraction. Rust is slowly eating at C++ and I'm pretty sure no company in their right mind today start a new product in C++. It will slowly become COBOL.
C on the other hand still has new C written every day. I'm a polyglot and I love C, but I never had a job in C. I hope some day that happens. The modern tools aforementioned, e.g. static analysis, clangd, are really good.
The thing I love most about C is that there is a direct relationship between me doing something stupid and feeling pain for it. In python, I can write crap all day without feeling it. Some day I'll feel it, but not today!
I like small binaries and fast code. I like that when a developer tells me I did something wrong, they can prove it, and it's not some stupid high level language "the more you knoooow". I like control over things.
I love modern C and I wish it well!
The library support for rust in key areas is terribly lacking and not everyone is a Microsoft that can afford to build the world from scratch.
As just one example a few valiant folks develop plugins and realtime audio with rust, but serious work that needs to get done today happens in C++.
I say this as someone who loves rust and zig, but these types of statements make me feel like HN is way out of touch with the industry. New products are constantly made in C++. That doesn't make it good (and i don't think companies building in C++ are making the best choices), but to say that no big tech company is writing any new code in C++ is just not correct. I see this happen every day at my job.
I also would definitely say more C++ code is written every day over C, although I'll say, I am not as familiar with the embedded world, but I know one HFT/Gaming/Robotics firm will have millions of new C++ lines every day.
I get exhausted by people denouncing C in favor of the language du jour, and suggesting it's been obsoleted by newer and especially more abstract languages. A hammer and nails are still useful for specific jobs, even if you have screws.
All tools have their trade offs. Even when you have a battery powered impact driver, sometimes you still need to pick up a basic hammer.
If you had a job in this you’d know that they are not ;)
Experience says otherwise. C code bases have serious longevity due to how many people know it and can jump in and contribute. There just aren't that many patterns and abstractions.
You and me both. I'm also a polygot but I originally started on interpreted languages like PHP and Python. I then learned some C, which was quite frustrating before I learned how to reason about memory ownership and hold myself to some idioms. Oh and Valgrind.
I then rewrote a bunch of projects in Rust and while it lead to correct and working software, it didn't spark the joy that C did for me. I don't exactly know why and at times I almost feel ashamed to mention this. I do hope there's a future where there's a version of C with some more substantial changes/improvements though, perhaps taking a lesson or two from Rust or Zig (eg string type w/ length).
The first thing I'd borrow from Rust are Option and Result for better error handling.
What's interesting, modern C, is promoting this move as well. Don't just return an integer, return a struct result_t with a fail bool or error bool in it, as opposed to some const char* pointing to null. Do this more and more and your C code starts becoming a lot more digestible and modern (although common sense still applies to not go overboard with these constructs in C, but you can set up a nice contract type API design within your code base).
Excellent! Such wisdom is rarely found in random forum threads on the internet.
I also wouldn't recommend starting a new thing with it if the whole team wasn't well-versed with it. On the other side of spectrum, today's world is well-served with python, c++ and rust (that battle is ongoing, I like both), java, heck even JS. There's something special about C though and I've yet to see something replacing it's place. Closest was D1, but that came and went.
I'm pretty sure this does not match reality.
> C on the other hand still has new C written every day
So does C++.
They also point out that C is a bad first language. This shouldn't be relevant to many Computer Science courses, but I have seen too many Electronics students who are taught C first, or indeed as their only language.
(Although C++ wouldn’t be my first choice to begin with for a first language)
Someone who knows nothing would benefit from Pascal a lot even in 2023, even if the practical relevance of the language nowadays is nil (not due to lack of merit, but due to the social dynamics, as a commenter on the OP's original blog righly said) - by the way, R.I.P. Niklaus Wirth. I guess it would be a bit like studying Latin or classical Greek to learn about grammar.
Python is "conceptually worse" than Scheme and Pascal - for learners at an academic level, IMHO - but of course practically more useful/valuable, from an industry point of view.
For example you want fat pointers, particularly for slice types. On a PDP-11 spending two registers for these types is extravagant, today this seems ridiculous, but C still doesn't provide any fat pointer types. So either you have to roll your own (and with them libraries of code to use them) or put up with whatever was a good idea in the 1970s. The most famous slice type is Rust's &str or C++ std::string_view, a fat pointer for referring to text - strings in other words, but not the mutable, owned, auto-growing strings you might associate with higher level languages, this is just the simple concept of some text. C can't do that, what C gives you is a pointer to a byte, pointer arithmetic and a stern admonition to stop when you reach a byte with a zero value... or you can roll your own slice types.
I admit that I still have a soft spot for C; once I finally understood pointers I did most of my projects in C; it helped that I had (and still have) a love for systems programming. To this day I can write C in my sleep even though it's been two years since I've last written a significant amount of C. But in grad school I got bit hard by the Lisp and Smalltalk bugs....I went from a big Bell Labs fan to a Xerox PARC fan, and in my professional career I've been largely coding in Python for the past five years since it's now the lingua franca of machine learning.
But I wouldn't recommend C as a first language; I feel it's too much for absolute beginners. I'm torn between Python and Scheme; my feelings right now is that Python is a good introductory language for helping people gain programming experience and allowing students to build interesting things using Python's extensive libraries, while Scheme is an excellent vehicle for teaching how programming languages work at a high level; I have a soft spot for The Structure and Interpretation of Computer Programs (this was the introductory CS textbook at MIT from the 1980s to the late 2000s when MIT switched to Python) and I used it as part of an upper-division course on programming language principles and paradigms at a university where Java is the introductory language.
You can read about how experts write their C code (like https://nullprogram.com/blog/2023/10/08/) but you aren't going to appreciate why they decide to do this. Indeed, beginners need to be able to blindly follow rules before they can critique them or invent their own coding styles.
typedef struct { i32 value; b32 ok; } i32parsed;
Defined a Result Type.
The biggest problem with C is that it lets you do anything to include things you don't want to do and should not do.
"The programmer knows best" means you can easily write to arbitrary memory addresses, smash the call stack, allocate memory and never free it, etc. and as long as the syntax is correct, your program will compile.
Sometimes, in interesting low-level code, embedded code, or similar, this unrestricted behavior is needed: you write seemingly arbitrary values to a memory address because it is memory-mapped hardware and that is where the control register receives instructions. But in most application-level code, it is bad behavior and just causes segment violations.
"The programmer knows best" is the primary cause of security flaws in C (and C++ because of its heavy compatibility with C) since even experienced C programmers make mistakes or find their code being used in unanticipated ways.
It's great if you're coming from assembly language and the de-factor standard for embedded systems.
Without understanding what code gets generated, there are so many footguns it can be dangerous at best. Knowing things like calling conventions and how the compiler interacts with memory are really important.
The assembler perspective is important, but not for correctness, rather for optimization. When you optimize it starts mattering how much register pressure and cache pressure etc you have.
Like I guess the point I have to make is that if you are writing C in 2024, there is likely a good reason, and if you don't know what's going on in the assembler, I feel like people are playing with fire.
I've found that I'm much less accurate when writing code for non-GCC/Clang compilers because my mental model of what's going to be generated isn't accurate enough without the years of experience I've had looking at the outputs of those specific compiler families.