Teaching C (2016)
blog.regehr.org
blog.regehr.org
Edit: I used to teach C to a group of undergrads and one question that came up often was around the special significance of the main function and order of definition of functions.
The RMS book addresses this very succinctly -
Every C program is started by running the function named main. Therefore, the example program defines a function named main to provide a way to start it. Whatever that function does is what the program does. The main function is the first one called when the program runs, but it doesn’t come first in the example code. The order of the function definitions in the source code makes no difference to the program’s meaning.
https://lists.gnu.org/archive/html/info-gnu/2022-09/msg00005... If you understand basic concepts of programming but know nothing about C, you can read this manual sequentially from the beginning to learn the C language. git clone https://git.savannah.gnu.org/git/c-intro-and-ref.git
cd c-intro-and-ref/
make c.pdf # requirements: curl, links, perl, texinfo (Linux), texi2html (BSD)
#!/bin/sh
set -v;
x0=c-intro-and-ref.git;
x1=https://git.savannah.nongnu.org/cgit/;
test -d $x0||mkdir $x0;cd $x0||exit;
for x2 in Makefile c.texi cpp.texi fdl.texi fp.texi texinfo.tex;do
echo url=$x1$x0/plain/$x2;
echo output=$x2;
echo user-agent=\"\";
done|curl -K/dev/stdin -s;
case $(uname) in :)
;;*BSD) texi2html --no-headers --no-split --html c.texi
;;Linux) makeinfo --no-headers --no-split --html c.texi
esac;
# personal preference for ~15-inch screen: 4 space indent, 60-70 max chars per line
links -width 70 -dump .html \
|sed '/^ *Link:/d;/^ *\*/s/\*//;s/^/ /;/Jump to: /{N;N;d;}' > c.txt;
exec less c.txt; printf 'the form\n' \
| tr a-z A-Z
golfs to printf 'the form\n' |
tr a-z A-Z
but that only saves like one or two keystrokes but is probably good to know for shell trivia questions or when implementing something that supposedly can handle shell continuation lines1. An extreme example
https://raw.githubusercontent.com/drudru/qhasm/master/qhasm-...
Not exactly true. Consider:
#include <stdio.h>
int foo(double d) { return (int)d; }
int main() { printf("%d\n", foo(3)); return 0; }
/*int foo(double d) { return (int)d; }*/
This prints different results depending on whether `foo` is defined before `main` or after `main`. This is caused by implicit function declaration.gcc will warn about this if the -std=c11 switch is thrown, but by default it will compile without error or warning.
(Of course you know all this much better than I do, but someone younger reading this thread might be mystified.)
This order dependency is why C source code tends to be written bottom-up, rather than the more natural top-down.
gcc /tmp/testf.c
/tmp/testf.c: In function ‘main’:
/tmp/testf.c:4:18: warning: implicit declaration of function ‘foo’ [-Wimplicit-function-declaration]
4 | printf("%d\n", foo(3));
| ^~~I think the biggest issue with C is that it's taught as the programs you are building is foundation of everything while brushing under the rug that in also all course work you are not build C on bare metal you almost always are building on on OS which runs an implemented hardware which are extremely important. But by teaching your building on this all powerful language makes overconfident programmers. I've known so many developers who jump the gun and think they are write what will happen on the computer without understanding how high in the clouds they normally are.
A couple of resources I found in the comments:
* A short list of safety rules for C: https://en.wikipedia.org/wiki/The_Power_of_10:_Rules_for_Dev...
* A much longer document, 42 rules for coding in C and C++: https://pvs-studio.com/en/blog/posts/cpp/0391/
Update: also, a 210-comment thread from 2018: https://news.ycombinator.com/item?id=18334476
C feels like a very simple language, but there are lots of subtle corners which probably don't need teaching at first.
On the other hand, I think it's important to hammer home "undefined behaviour" early. There are too many guides that say writing outside the bounds of an array "writes to memory outside the array", which simply isn't true. It might write, it might not, who knows, undefined behaviour means all bets are off.
In general I feel a lot of practical teaching of C is teaching about undefined behaviour, which is something many other languages (Java, Python, Haskell, Rust), either don't have, or where they do beginners won't stumble across it.
https://git.musl-libc.org/cgit/musl/tree/src/string/memchr.c
and the user has completely disappeared along with their comment/question.
Anyhow, you can think about the __GNUC__ part as an attempt to fast-forward the loop by testing size_t-length chunks for the character by XOR-ing the chunk with a repeated string of that character. The first loop aligns the pointer for the fast-forward. The final loop is then used to test the unaligned part, which when not using __GNUC__, would be the complete loop.
I felt the need to answer because I'd done the work.
C still is, and will continue to be for the foreseeable future, the lingua franca, the least common denominator. Acknowledging how C comes with a very... interesting set of tradeoffs that make it uniquely well-suited for certain purposes and at the same time incredibly dangerous is a worthwhile proposition if one is aiming to truly understand C development.
I would highlight the following:
* Structured programming support, which means nested loops and conditionals without a primary need for "goto" jumps, enabling a sense of "depth" that is missing in the "flat" world of assembly
* Expression-oriented syntax, meaning that operators (even those having a side-effect) return a value of a certain type, and can be nested, again enabling recursive program structure versus a flattened one
* Global symbol allocation and resolution, which means that a programmer uses names rather than addresses to refer to global variables and functions
* Abstraction over function calling conventions, which enables the programmer not to worry about function prologues and epilogues and the order of pushing arguments on the stack or in registers
* Automatic storage management, meaning that a function-scope local variable is used by the programmer with its name and the compiler decides whether to put it in a register or at a certain offset in the stack frame
* Rudimentary integer-based type system that has the distinction between a scalar and a fixed-size collection of scalars laid-out sequentially (arrays), and special integers called "pointers", supporting a different set of operations (dereferencing to a certain type and adding or subtracting other integers from them, without any safety guarantees whatsoever)
Nothing more, nothing less. Not understanding these foundations is the source of major pain.
Additionally many of the C features had already been sorted out in JOVIAL, NEWP, PL/I, BLISS among others about a decade before C was born.
C was solving the issues of UNIX v3 design, that is all.
Plenty of languages can be used to teach low level programming concepts.
In the context of platform ABIs, sure. The widespread stabilization and ossification of C ABIs is a boon for the rest of the ecosystem but it's entirely at the expense of the C language/stdlib. Hence the performance advantage of projects like fmtlib.
Notwithstanding its ubiquity C is in many ways "The Sick Man of Asia". Every major C compiler is written in C++ with tooling heading the same way. The dominance of C++ in the heterogeneous space has accelerated this trend and spread it to many HPC libraries. Even foundational bits such as Microsoft's UCRT or llvm-libc are written in C++.
On the current trajectory C will become the next Fortran, i.e. a widely used language which is nonetheless unable to support itself.
I'm not a native speaker - this is a genuine question - does it sound weird to say "least common denominator"?
The web isn't a perfect corpus, and Google isn't a perfect corpus analysis tool, but with those caveats...
Google results for "least common denominator": 738,000 / Google results for "most common denominator": 656,000
Additionally, as you point out, Wikipedia lists both as synonyms.
Your original comment was correct and idiomatic English.
It teaches us various Linux/Unix concepts like signal, threads, file I/O, socket, etc. On some chapters, we'll be guided to re-implement various built-in tools like who, ls, sh, pwd. Very interesting from developer point of view.
I wonder how the claim holds up? Certainly not well in the embedded space, and Rust mainline kernel development seems still a few years away at least.
I'd argue there's some transition in space fsw to c++, but that's not really significantly different than C+classes, since most c++ features and std:: are not allowed.
When does automotive, iot, and other spaces anticipate tansitioning?
I'm not sure if this is representative, but most of my recent FLOSS projects have all been C89 libraries.
The rationale is that it's trivial to add C bindings in any language, it's extremely fast to build, and it runs everywhere.
IMO, targeting C89 is only warranted when you know for a fact it's needed. After all, not only is C99 a good improvement but toolchains stuck on C89 are not the type of toolchain I'd want to support.
Until somewhat recently C89 was the latest version of C that was supported by the MSVC compiler. Support for C99 was always terribly broken and Microsoft chose to claim it did supported C99 but without supporting mandatory features, which meant they did not in fact supported C99.
This sad state of affairs only changed significantly in 2020, with a lowkey announcement.
https://devblogs.microsoft.com/cppblog/c11-and-c17-standard-...
Likewise, even with the updated C11 and C17 support, it isn't 100% there, only good enough, because C++ is what matters for Windows developers when not using .NET.
https://www.autosar.org/news-events/details/autosar-investig...
Note that MISRA is also adopting AUTOSAR on their standards.
True, but AFT gives us at least 8 years.
I feel that last phrase exemplifies why this should be given quite a lot of attention. It is also the sort of thing that will lead a student to a deeper insight into the language and how it differs from the other procedural languages that they are likely to already be familiar with (the same can be said for memory management, BTW.)
Such a course might be something of a grind, but if so, it would inculcate the right sort of attitude for programming competently in C!
For dessert, take a quick look at some Obfuscated C winners.
ISO/GNU C is unfit for programming classes of devices as ubiquitous as smartphones, or virtually any type of SoC. There is a reason CUDA/OpenCL/ROCm/SYCL exist and why they can't be programmed like usual C if you want performance.
If you look at the assembler an optimising compiler produces from C, it can take hours to learn how to map the C to the assembler. Also, all the business with "undefined behaviour" (which they will hit soon, and often) isn't how computers work either -- assembler has (very little) undefined behaviour, if you write to a random memory location it is written to (unlike in C where maybe it is, maybe it isn't).
My recommendation would be to teach them whatever is the quickest way to something they enjoy, wether that be Python, Minecraft, Roblox, Javascript for web dev, or C if they (for some reason) really want to start by doing kernel programming or something like that.
You can learn all the "fine details" later, they have a lifetime to do it :)
It's weird to teach a language while also ripping a hole in one of its main abstractions and assuming stuff about how that works, in my opinion.
It seems so prevalent that locally-scoped variables are often referred to as "stack variables" in casual conversation, but I'm curious of cases where it's not true...
> 6.5.2.2.11: Recursive function calls shall be permitted, both directly and indirectly through any chain of other functions.
This either means using stack frames, or heap-allocating activation records and tying them together with references. Second is a much rarer approach.
I do use UBSan/ASan as well, but I can't enable them for production run, and I can with no-strict-aliasing, which is at least safer than without-it.
Let them learn C when they need to use C. Consider it a specialist language that is only used for certain tasks, like Fortran or COBOL -- not something everybody has to know.
The question then becomes, why do we want more evil people in the world?. Does that make any sense? because that's exactly what you get when you don't use Rust.
Another problem recently was that compiling Rust required downloading a binary blob from Mozilla. That's a no-go for many projects.
Meanwhile Microsoft has finally acknowledged picking C for Azure Sphere was a bad idea for its overall security story, and is now adding Rust support.
Just telling on yourself, "I can't write software without training wheels", "I couldn't be bothered to learn how pointers work", so flagrantly.
Rust is unspecified, lacks real battle testing, and lacks any substantial track record.
Comparisons to C are comical, and discussing Rust as being a C replacement as if it were a forgone conclusion is just... I mean I can't think of any way to phrase this that won't result in a ban - so use your imagination.
Rust is probably a fine language. A lot of y'all down the rabbit hole need a reality check though.
> Just telling on yourself, "I can't write software without training wheels"
I think both of these reactions go too far, in opposite directions. Of course the fact that C is difficult doesn't mean that we should stop teaching it. It could just as well mean the opposite, that we need to teach it better, certainly as long as it stays in widespread use. But at the same time, we should acknowledge that C is difficult even for experienced professionals, and that "not knowing how pointers work" isn't the main reason every large C and C++ codebase on earth has memory corruption vulnerabilities.
(I'm sure that's not literally true. Someone somewhere must've written a lot of perfect C code. But I think the usual posterchild for well-tested C is sqlite, and even sqlite has had memory corruption issues in the wild.)
I know enough about C to know that I can't write it safely 100% of the time, especially when you introduce things like parsing untrusted input and threading. Thinking you can do this safely, and thinking you don't make mistakes suggests you actually don't know as much about C as you think you do.
There number of subtle and unexpected things that cause UB are pretty concerning. Most of our software that we rely on day-to-day is filled with subtle bugs, many of which will eventually be exploited and used for RCE and other nasty things. I don't understand how that couldn't concern you!
To be clear I don't think we should stop teaching C or anything that extreme. I don't think it should stop being used completely either. Mostly just that we should prefer safe languages when possible and practical, or use hardening features when we do use unsafe languages, like bounds checking for example. A lot of times I think we shouldn't even prefer rust, a lot of userspace software can be written in a GC'd language without issues.
But just because it's unsafe doesn't mean it should never be used.
It's unsafe in the sense that it's nearly impossible for even clever, experienced programmers to write nontrivial amounts of code in it that don't have foot-shooting behavior in some form.
No rooting, no jailbreaking, no way to break out of the dystopia they will inevitably try to create.
"Those who give up freedom for security deserve neither."