Teaching C
blog.regehr.org
blog.regehr.org
The issue that many students face in learning low level C, is that they don't learn assembly language programming first anymore, and they come from higher level languages and move down. Instead of visualizing a Von Neumann machine, they know only of syntax, and for them the problem of programming comes down to finding the right magic piece of code online to copy and paste. The idea of stack frames, heaps, registers, pointers are completely foreign to them, even though they are fundamentally simple concepts.
I agree with this. I found another article that made the transition for someone coming from something like Python or Ruby easier. It shows how to use gdb as a sort of C REPL. https://www.recurse.com/blog/5-learning-c-with-gdb
The quick feedback loop certainly makes things like pointers, arrays, etc, more clear.
What about teaching C in the context of something like AVR programming, where you have to worry about those sort of things because there simply isn't any abstraction on top? That was where I first encountered C and I think learning it in such a constrained environment helped me appreciate/understand the utility of C a lot more.
I don't agree. That leads people to incorrect conclusions like "int addition wraps around on overflow" (mentioned in the article), "null pointer dereferences are guaranteed to crash the program", "casting from a float pointer to an int pointer and dereferencing it is OK to do low-level tricks", and so forth. C is a language that implements its own semantics, not the semantics of some particular machine. Confusing these two ideas has led to lots and lots of bugs, many detailed in John's other blog posts.
It might be useful to teach these intuitions to beginning programmers who already know assembly language before learning C (though are there more than a vanishingly few number of those anymore?) But teaching assembly language as part of, or as some sort of prerequisite for, teaching C strikes me as a waste of time and likely to lead to wrong assumptions that students will eventually have to unlearn.
Your point about teaching assembly as being a waste of time also has some merit. Of course people still do program in assembly, but it is only becoming more and more rare to actually need someone to program assembly, which is why it just isn't emphasized as much any more and it makes C seem like an even stranger language for someone who started on Python or Javascript.
No, it is special. Dereferencing it is undefined behavior. It is not guaranteed to result in a load at address zero.
The main reason I keep coming back to assembly language with C, is that C cannot do anything that assembly language cannot do. It only places additional restrictions on assembly, which is a bit easier to grasp (bitwise operations, add, subtract etc.). Once you understand the fundamental operations of the processor, you can start to learn the copious corner cases that is the C programming language.
I know where you're going with this, but I don't agree with this statement. C provides the abstraction of structs and unions, which do not exist in assembly. Arrays are also an abstraction which does not exist in assembly - yes, assembly provides the mechanisms to easily compute addresses for arrays, but that is different than having a named entity for many objects. The direction you're going is that C is a thin abstraction over assembly, which I agree with. But that is different than saying C has a one-to-one mapping, which is what I think your statement implies.
That's basically true. Except that an optimizer is allowed to transform things as long as the observable behavior remains the same. And undefined behavior means more than "just that one line is undefined" ( http://blog.llvm.org/2011/05/what-every-c-programmer-should-... , http://blog.llvm.org/2011/05/what-every-c-programmer-should-... , http://blog.llvm.org/2011/05/what-every-c-programmer-should-... ). These two facts conspire to make undefined behavior surprising to many programmers.
For instance -- using one of Lattner's examples -- on some architectures it's a little expensive to check if a loop variable had wrapped around. So the optimizer will omit the check for wraparound if it can prove wraparound is impossible. That is, if it can prove that n + 1 > n. But the Standard only requires wraparound for unsigned integer types. So the optimizer can also omit the check if it can only prove either n + 1 > n or n is a signed integer. In this case, a bounded loop turns into an infinite loop, but "undefined behavior" includes that kind of transformation.
More to the point, Linux had a severe security bug where they dereferenced a pointer, then checked if the pointer was NULL before returning the value ( https://lwn.net/Articles/342330/ ). The optimizer removed the check for NULL because if the pointer was valid when it was dereferenced, the check was unnecessary; and if the pointer was NULL when it was dereferenced, then dereferencing it was undefined behavior and "remove an NULL check" is a valid transformation in undefined behavior. Then moving that load to a different point in the function is also a valid transformation (as long as it doesn't affect observable behavior), and a few more transformations could make the NULL dereference do something completely different from what you expect.
unsigned short *a = 0;
((void (*)(void))a)();
which may cause a jump to address 0 (usually the reset vector).On R8C with the NC30 compiler there are no warnings and the output is:
MOV.W:Q #0H, -2H[FB]
MOV.W:G -2H[FB], R0
MOV.W:Q #0H, R2
JSRI.A R2R0
It jumps to 0. But on PIC 16 with the XC compiler a warning is emitted warning: (1471) indirect function call via a NULL pointer ignored and the statement is not compiled at all. printf("a=%p *a=%d", a, *a);
if (a == NULL) return;
An optimizing compiler will likely remove the NULL check completely because the printf above it has undefined behavior if a is NULL.(I think I read about this somewhere, but it was a while ago, and I never really programmed for DOS (well, one little program to set the serial port to some specific settings, but that was about ten lines including whitespace and a little comment explaining what the code did).)
The compiler might emit machine code that attempts to read the memory at location 0. It would also be perfectly within its rights to optimize away that branch of code (if it can prove that it always dereferences null). The code might even appear to work now, and completely break in the next version of the compiler.
And if one isn't writing for high performance, probably one shouldn't be using C. I like the suggestion in a sibling about teaching C in the context of programming a microcontroller. I think this might bridge the gap a bit: encourage the right intuitions, without creating dangerous misconceptions.
For example, here are some of the things that confused me when I first learned C:
- why isn't there an actual string data type? (just char*)
- why do some people use "char" to store numbers?
- whats the deal with pointers?
- why are pointers and arrays kinda sorta interchangeable?
Until I learned how things work at the assembly language level, I could not gain an intuitive understanding for why C works the way it does.For example, I could imagine a machine with identical memory layout to C, but that supported a hugely parallel, variable-size data bus where an operation like "x = y" (string assignment by value) could copy an entire string in a single operation.
The reason C doesn't support this is because generally at the assembly language layer you can only operate on memory in register-sized chunks, and every load or store of a register's worth of memory takes time. So assignment of a string by value requires a loop, just like it does in C.
Understanding the why is interesting - like it is for any language. And like in many languages, the why tends to be complicated and somewhat arbitrary at times. If you're really interested in the why rather than the what, a book on the history of C will probably be more useful than learning assembly.
True, but I think knowing how arrays (and pointers to arrays) work in C is one of the main hurdles for a lot of people who are just starting out.
This became especially apparent to me after recently helping my friend get familiar with C.
I can see how some people can get confused by this when you consider the fact that structures can be copied using a straight assignment, while arrays can't.
People naturally try to find similarities when learning something new, so it took a little while for him to _really_ get it. I think his mind kept trying to think of a structure essentially as an array of variables, when that's not really the case.
Then it becomes apparent why one cannot copy strings "by assignment" and can structs: it is in general impossible to know runtime size of string at compile-time. C strings have no structure known at compile time. Structures, on the other hand, are there to enforce structure on data.
There is a pretty neat real-world analogy here: copy machine. String copy must be done character-by-character in a same way that book or document folder would have to be copied page-by-page. On the other hand engineering drawing must be copied whole. It may contain references to other drawings, you can still with some struggle extract individual parts, but it is copied as a whole. This analogy relies on a fact that drawings are single-page and but nicely encapsulates the "strings are arrays are pointers" idea: folder may be empty, may be single page, but it is impossible to know without attempting.
struct { char name[30]; } tgt, src = { "Einstein" };
tgt = src;
Since there are no strings, it cannot be true that strings are pointers. They can be arrays, and array sizes are known at compile time. It's just a 40-year old cop-out that we can't copy by assignment all types based on their `sizeof` size.You are stepping on the same rake as beginners: generalizing from a specific case instead of applying the general case to specific circumstances. I have stressed the word "in general". You may treat it like a cop-out or you can say that there are no special-case semantics here. Not sure if it was intentional, but your example is rather tricky. Until you step out of the box and see that we are no longer dealing with strings/arrays/pointers here, but structs, that have a bit different semantics.
> Since there are no strings
Yes, there is no explicit string type in a language, but somehow we do use strings in C. Semantics. We can semantically treat a particular block of memory as a string, time-series, binary tree, etc.. There is simply no special case (explicit language support) for strings.
> it cannot be true that strings are pointers. They can be arrays, and array sizes are known at compile time.
What about `malloc`? What about passing arrays between compilation units? I have covered this in SO[1]. Note that I never explicitly pass pointers, yet `sizeof()` thinks I do. Array sizes can, in some circumstances, be known at compile time in a specific program block, but not in general.
I'd say there are +/- 3 types of languages (core, no stdlib, etc) in this context: 1) provide common-special-case exceptions 2) wrap all cases in an easy-to-use interface 3) provide general case syntax. 1) languages with `=` and `eq` (Perl?) 2) languages with object-identity (Python?) 3) C
[1]: http://stackoverflow.com/questions/19589193/2d-array-and-fun...
You can copy/initialize strings in C without ever writing a single loop by using strcpy. As for hardware, you need not have something so exotic. x86 for instance has a string copy instruction. The C strcpy function is often compiled to it.
You can copy entire structures by value by doing "x = y".
This ends up being implemented as a memory copy loop, but I'm not sure what happens with the padding bytes. There's a chance they get copied as well, but I really don't know.
In the code base I work with, it's common to see arrays inside of a structure which they're the only member of. This makes it a little easier to copy them, though I'm not sure if that was the intended purpose.
Something like this would be defined in a header...
typedef struct { char myArray[50]; } Test_Struct_T;
Hence: you can't use memcmp() to reliably compare structures for equality.
template<typename T, size_t N> class array { T data[N]; };
That wasn't in the original language, though it's a pretty old extension.
That's why I don't recognize that feature. I guess we didn't have it back when I was doing a lot of C.
You mean like x86 (at least as far as the ISA is concerned) [1]? :)
I mean, this isn't just me being pedantic and annoying: I think it goes to show that C's machine is quite different from a real machine.
[1]: http://x86.renejeschke.de/html/file_module_x86_id_279.html
I'm constantly surprised by how poorly documented the actual operation of current processors is, and how few people seem to care. In one way, this means that the abstraction is working, and no longer does anyone need to look behind the curtain. In another way, like the move to teaching only higher level languages, it feels like something essential is being lost.
Reminds me of this story of a game developer who magically got their game under the limit at the 11th hour
http://www.dodgycoder.net/2012/02/coding-tricks-of-game-deve...
With the caveat that unless you precede them with signed/unsigned modifiers, it is not portable.
Because apparently writing ptr = &arr[0] like in older systems programing languages was too hard implement.
So, the C model neither makes sense nor matches assembler of today's CPU's. We've certainly developed ways to implement it efficiently. It wasn't designed for that, though.
https://en.wikipedia.org/wiki/PDP-11_architecture
It wasn't that PDP-11 made C implementations more efficient. It's that C was a BCPL specifically designed to compile easily and run fast on their PDP-11. That's why I can't overemphasize C's actual history vs the lore that people repeat. It's literally an ALGOL language with every feature that couldn't compile on 60's and 70's era hardware chopped off with some extensions added latter.
Worked fine for a PDP-11. Yet, forcing its memory model or tradeoffs into a language used on different hardware can cause unnecessary problems. In contrast, Hansen's Edison language deployed on PDP-11 had only five statements (extreme simplicity haha) but would map efficiently to most architectures. As would Pascal and Modula-2 that inspired it & were safer.
Contemporary x86 and RISC CPUs are what I was comparing the PDP-11 instruction set to. I don't see any fundamental differences. Painting with very broad strokes, the PDP11 ISA looks reasonably close to x86. And those minor differences in more modern RISC actually map better to C than the PDP 11 does -- for example, status flags being replaced with jumping based on register contents. Implicit widening to words is a weak mismatch for x86, but it's a pretty good match for modern risc (no need to mask out top bits in registers), etc.
I looked at your links, and I'm still not seeing how other C maps better to a PDP-11 than it does to modern CPUs. The only thing I'm seeing in the pastebin rant is that CPUs are fast enough and memories are big enough today to support more expensive features, which I can agree with.
Again:
> Worked fine for a PDP-11. Yet, forcing its memory model or tradeoffs into a language used on different hardware can cause unnecessary problems.
What parts of its memory model or tradeoffs made it into C? I can't find any specifics that you're basing these claims on, only assertions that it's true.
In fact, the usual complaint associated with C is that it left the memory model so loosely specified -- initially to allow it to match any hardware -- that optimizing compilers can use the looseness to do really strange things to your code.
Fair enough haha. Ok, my memory loss is hurting me on examples. I might have just been the little things adding up. I do recall two from security work: reverse-stack and prefix strings. MULTICS, UNIX's predecessor, had both with significant reliability and security benefits. Reason C had null-terminated strings was PDP-11's hardware and one personal preference/opinion:
"C designer Dennis Ritchie chose to follow the convention of NUL-termination, already established in BCPL, to avoid the limitation on the length of a string caused by holding the count in an 8- or 9-bit slot, and partly because maintaining the count seemed, in his experience, less convenient than using a terminator."
Now, on reverse stack, my memory is cloudy. Common stacks have incoming data flow toward the stack pointer in a way that can clobber it, even leading to hacks. MULTICS had data flow away from the stack pointer with an overflow dropping into newly allocated memory or raising an error. C language (and most) implementations use regular stack. I think it was because PDP hardware expected that with a reverse stack requiring high-penalty indirection. I could be wrong, though. I know a reverse stack on x86 gets a performance penalty and key traits of x86 come from PDP-11. A CISC with reverse stack would have problems with C.
The pointer stuff. Lots of the pointer stuff, esp arrays, comes from efficiency needs for running on a PDP-11. This by itself is why we can't map C easily to safer or high-level hardware. The CPU at crash-safe.org, jop-design.com, and Ten15 VM come to mind. PDP-11 model doesn't support safety/security so neither does C.
These are a few that come to mind that carry over into modern work trying to go against C's momentum. Hardware, software, and compiler work.
That's not a restriction of C, but a way to get more out of your memory on a restricted system; If your heap grows up and your stack grows down (or vice versa), then you can keep using growing both until the two meet, at which point you've used all the available memory. However, if they both grow in the same direction, you need to statically decide how much to give each one, which will lead to waste if you're not using much stack or heap:
[heap-->| |<--stack]
vs: [heap-->| |stack-->| ]
But, again, not something that C cares about; you have a number of architectures like Alpha (IIRC) where the program break and the top of stack move in the same direction.https://www.acsac.org/2002/papers/classic-multics.pdf
Really old stuff. Relevant quote: "Third, stacks on the Multics processors grew in the positive direction... if you actually accomplished a buffer overflow, you would be overwriting unused stack frames rather than your own return pointer, making exploitation much more difficult."
I can't find the original paper showing the penalty on x86/Linux. However, this one does the same thing for different reasons with many details:
http://babaks.com/files/TechReport07-07.pdf
Key point: "The direction of stack growth is not flexible in hardware and almost all processors only support one direction. For example, in Intel x86 processors, the stack always grows downward and all the stack manipulation instructions such as push and pop are designed for this natural growth."
So, on such an architecture, you can't directly use the stack operations to do the job: must implement extra instructions without hardware acceleration. The stack on x86 is effed-up and insecure by design. If C's stack is fixed, there's still a mismatch between it and x86 ASM. Itanium at least provided stack protection among other security benefits.
Note: I'm not saying this is sufficient to stop stack smashing. Just that reverse stacks are a better idea than the ludicrous concept of making unknown amount and quality of data flow toward the stack pointer. Definitely reduced risk a bit but how much takes more assessment.
I think his point that "The idea of stack frames, heaps, registers, pointers are completely foreign to them..." is the critical point.
Require students to write a toy operating system and compiler in C. They will understand things obscenely well by the time that they are finished.
Another option is to introduce students to DTrace and require them to use it to answer questions about kernel and userspace behavior. Ask the right questions and they will learn all of the things you want them to learn from the process of answering the questions.
This seems like a particularly uncharitable thing to say about people who write code in a 'high level' programming language. You're alleging not just that they don't understand the fundamental low level workings of a computer, but that they don't even understand how to write new programs in their language?
Though there is some truth to that. It's far more likely that copy-and-paste X is a workable solution in a higher-level language than a lower-level one.
A high-level language emphasizes portability and abstraction; a low-level languages emphasizes performance and implementation details.
So...the comment sounds harsh but in reality is a reflection of the success of high-languages.
I've seen so many people that just can't wrap their heads around pointers, but it makes so much more sense when you've gotten down to the nitty-gritty level and built up from there.
[1] http://www.amazon.com/Introduction-Computing-Systems-gates-b...
[2] http://www.amazon.com/Structured-Computer-Organization-Andre...
http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
(HN truncates the text of the URL, but they're all different)
I've been coding C professionally for over a decade, as required for firmware/embedded development, and those posts have instilled the fear of god in me.
But why? Yeah, hitting UB can be a terrifying idea but rarely happens in practice.
In two decades of C programming, I have hit an UB bug exactly once when a piece of code was ran on an ARM platform for the first time. It took a little bit of staring at disassembly and reading some docs to sort out but it wasn't the end of the world.
Understanding the basic cases is a good idea but the darkest corners of undefined behavior are only important if you're a compiler writer like Chris Lattner is.
Undefined behavior has been the source of numerous security exploits in the past, and will only get worse as modern optimizing compilers become more advanced.
It was rare for students 25 years ago to learn assembly first. Can't say it made it harder to learn or that students had issue (spent a couple of years in the 90s teaching C part time to contractors). They had issue with language beginner things. Pointer arithmetic, or confusing pointers/arrays, but can't ever remember anyone having a particular issue with switch. Fall through was a C thing, they accepted it quite happily, and forgot break sometimes as learners do. People seemed to have far, far more difficulty getting comfortable with C++ and OO than C. The new C++ programmer was much more dangerous than the new C programmer!
C was often taught as first "real" language. You'd introduce pointers and here's how that aspect of computers work as part of the same scribble on the whiteboard. Same for memory allocation, stacks, heaps and byte sizes/packing. The fact that C was so directly close to those concepts made grasping them that much easier.
We lost a lot when we moved beyond expecting people to be aware of those basics. PHP isn't even sure itself what data is. Being able to pack your data or have app data that's optimal would be appreciated by those "few" smartphone users outside SV where dropping data or fallback to GPRS happens often. Data is rarely thought of in terms of size, it's just a blob of some types/objects. Little surprise when the app spits JSON of epic size and spends half its time "thinking".
At least it's not XML...
C maps nearly 1:1 onto simple processor and memory models, and most importantly, gets out of your way and lets you get on with solving your system programming problems. Before C, just about any meaningful system programming task required a dive into assembly language. In that context, C was a huge win. It is also what makes C the langauge of choice for embedded development today.
Of course, system programming problems are not the bread-and-butter of most develpers today -- and a good thing, too. We can now build on top of solid systems and concentrate on delivering value to the customer at much higher levels of abstraction: the levels of abstraction that are meaningful to customers.
I dearly love Python because it allows me to work at levels of abstraction that are meaningful to the user's problem. I dearly love C when I want to wiggle a pin on an ARM Cortex-M3.
In my mind, CS education should start by teaching problem decomposition and performance analysis using a language like Python that provides high levels of abstraction and automated memory management. Then, just like assembly language was a required CS core course back in my day, students today should spend a semester implementing and measuring the performance of some of the data structures that they have been getting "for free" so that they understand computing at a fundamental level. Some will go on to be systems programmers, and will spend more time at the C level. Some won't ever look at C again, and that is OK.
In the end, CS education is about how to solve problems through the application of mechanical computation. The languages will evolve as our understanding of the problems evolve and our ability to create computing infrastructure evolves. CS educcation should be about creating people who can contribute to (and keep up with) that evolution.
I agree with that. I've proposed it myself. I'll add that I prefer them starting with a more type- and memory-safe, but low-level, language like Component Pascal so they can learn low-level thinking but appreciate safety features & good language definitions when they learn C afterward. One can emulate a decent bit of that in C and HLL that compile to C. Maybe they'll remember enough to make something valuable.
Far as industry, I'm not feeling as strong doubts as you looking at the quantity and diversity of this list:
https://wiki.haskell.org/Haskell_in_industry
There's some real ass-kicking going on in industry with Haskell even if it's niche use and has obstacles/issues from that perspective.
I never get to use it at work, but it helped improve my skills when using FP concepts in C++, JVM and .NET languages.
That seL4 matched Haskell and C... along with open-sourcing a key tool (AutoCorres)... makes me think this is achievable. My method isn't formal proof but would be more accessible. What do you think?
http://repetae.net/computer/jhc/
It doesn't support all the language features, though.
Besides their standard one, GHC also has an additional LLVM and an older, now deprecated, C backends:
https://downloads.haskell.org/~ghc/7.6.3/docs/html/users_gui...
The UHC also has a C backend, but not fully implemented,
http://foswiki.cs.uu.nl/foswiki/Ehc/UhcUserDocumentation#A_6...
Sounds a good idea, on the other hand have you ever used Frama-C? Never used it, but it seems being used in this kind of scenarios.
re generating C. The key thing is whether it generates human-readable and -editable C. Most of those tools are used to just feed and piggyback on a C compiler. In my scheme, the Haskell is like the high-level, executable spec with C being equivalent. Such tools might be a start on my goal. I like that JHC has no garbage collector. That's promising for same reason Rust having no GC is. :)
re Frama-C. Oh, yeah, that's good thinking as it's already been used for plenty C verification, even standard library. I was thinking of encoding Haskell specs in Frama-C somehow but not sure as I'm not formal methods specialist. Anyway, you might like where such techniques got their start for mainstream languages:
http://apotheca.hpl.hp.com/ftp/pub/dec/SRC/research-reports/...
Another possibility using something like the "Tiger book" and write a ML -> C compiler for the C subset you care about, but by then maybe contributing to Ada (GNAT), Rust or Swift would be better use of time.
I also forgot to mention on my previous comment that Idris and F* also have C backends, but they might suffer from the same problems.
Links describing it & to various projects http://goto.ucsd.edu/~rjhala/liquid/haskell/blog/about/
C verifier download http://goto.ucsd.edu/csolve/
What keeps you from wriggling that pin in Python?
Turing completeness dictates that any Turing complete language is equivalent to any other Turing complete language. In addition, you can do high levels of abstraction in C. The preprocessor and void pointers allow you to do some rather nice things. Structured programming also let you build things up. The advantage that Python enjoys is that it comes with a large number of library functions already provided and there are easily discoverable third party libraries. There are other differences, although every difference has a trade off. Garbage collection bloats memory requirements. Being interpreted means errors that can be caught in advance at compile time occur at runtime.
Did you catch that "Cortex-M3" part after "ARM"? No MMU, no OS, very limited SRAM -- CPython doesn't go there. I am, however, a huge fan of MicroPython, but that discussion too OT for this thread.
Sorry to pick on you, but this is a good example of misuse of this fact in an argument where it's not really relevant. Turing-completeness only relates to functions that take some string as input as produce a string as output (since that's all the turing machine model can do). In contrast, real-world programming languages interact with a machine or operating system, and not all languages provide the same interfaces or even run on the same machines.
I can easily implement a Turing-complete language whose only allowed system calls are reading from stdin and writing to stdout. Despite being Turing-complete, it will never be able to spawn threads, connect to a network socket, or even allocate memory on the heap.
We didn't use any fancy IDE's and were told to stick to VIM, we also had to compile with the flags -ansi -Wall -pedantic which alerted you to not only errors but warnings when we compiled our code if it didn't meet the C90 (I think) standards. It was a lot of work crammed into 13 weeks but it had one assignment which I thoroughly enjoyed.
Tic Tac Toe (Ramming home using pointers, 2D arrays, bubble sort for the Scoreboard).
Debugging a bug-riddled program (My favourite).
Word Sorter (Using dynamic memory structures, memory management by having no leaks, etc).
The debugging one was very different from most other assignments I had done at uni to date and the teacher said he recently introduced this assignment because the university had received feedback that students debugging skills weren't the greatest. They could write what they were asked to just fine, but when it came to debugging preexisting issues quite a few struggled. We got given a program with around 15 bugs and you got marks depending on what was causing the bug and a valid solution to fix it. This forced us to use tools such as GBD and Valgrind to step through the program and see where the issue was and to be much more methodical.
I really enjoyed C and when I find a bit of time outside of work and study I'd like to explore it more.
EDIT: Meant to add: fantastic article, wish my Intro to C instructor had read it...
#define ONES ((size_t)-1/UCHAR_MAX)
Reading through the loop, it seems as though it will create, e.g., 0x01010101 for a 32-bit machine with 8-bit bytes. And sure enough, if you calculate ((2^32)-1)/255 that's exactly what you get. But I never would've known that without going through the code and proving to myself that the definitely of ONES actually makes sense.If you write code like this and there are never any bugs then fine, I guess. But there will bugs.
Edit: And most likely in the excellent book "Hacker's Delight" as well.
Absent macro obfuscation, it is easy to reason about what a snippet of C does and how it translates down to machine code, even taken out of context. In C++, something as innocent as "i++;" could allocate heap memory and do file I/O.
The downside is that C code can become quite verbose, and to do anything useful, it takes a lot of ground work to basically set up your own DSL of utility functions and data structures. For certain applications, this is an acceptable tradeoff and gives a great deal of flexibility. I think teaching this bottom-up approach to programming can be quite useful - in a way, it mirrors the SICP approach, albeit from a rather different angle.
The question is, why are there not more languages that have the same paradigm, but also add basic memory safety, avoid spurious undefined behavior, provide namespaces, with a non-stupid standard library, etc.?
You mean Algol, NEWP, Mesa, PL/I, PL/M, Modula-2, Ada and similar?
A few of them are older than C.
Not recommending this route today as I just happened to start with BASIC. More like combining a simple, easy-to-compile, safe-by-default language w/out GC and with macros that extracts to portable C. Should make programming in C easier while providing all the benefits of C as my system did.
I've been eyeballing Nim language for this since HN commenters pointed out it's close to the goal already:
https://github.com/nim-lang/Nim/wiki/Nim-for-C-programmers
Someone might even be using a subset of it as a high-level C language. I'd be interested to know if anyone reading is doing something like that. Plus any other language besides Nim that can closely map and extract to C without its issues.
"The question is, why are there not more languages that have the same paradigm, but also add basic memory safety, avoid spurious undefined behavior, provide namespaces, with a non-stupid standard library, etc.?"
...basically just described Modula-3. It meets your requirements, was easy to read, had concurrency, had decent stdlib which had some formal verification, and could act as low-level as you needed with "UNSAFE" keyword. Brilliant design given all tradeoffs it balanced. It had some commercial uptake and was used in CVSup for FreeBSD.
https://en.wikipedia.org/wiki/Modula-3
Note: Important to not ignore it once you see "garbage collection." The GC was optional with a single keyword determining whether you or it handles a specific variable. Let's one pick and choose their battles with fate. :)
Note 2: The Obliq distributed programming language was an interesting project based on Modula-3. The SPIN OS, written in Modula-3, let you link code into a running kernel in a type-safe and memory-safe way for reducing context switches for performance.
(Regular HN readers will recognize the author of the top review on Amazon.)
[0] http://www.amazon.com/Interfaces-Implementations-Techniques-...
I learned C by myself many years ago but it's only until recent I have been using it for big projects.
Reading Redis' source code was a great aide, xv6 is also amazing to learn systems programming.
Learn C The Hard Way is also a good read, but not as your main book, since it goes too fast. Other invaluable resources are: Beej's Guide to Network Programming and Beej's Guide to Unix Interprocess Communication
A good advanced book is Advanced Programming in the Unix Environment
Programming in C
C Primer Plus
K&R (obviously)
21st Century C
Modern C (also mentioned in this post)
Understanding and Using Pointers in CUnfortunately, it appears to have been out of print for a while now.
Another book I can highly recommend is "The New C Standard: A Cultural and Economic Commentary" (http://www.knosof.co.uk/cbook/cbook.html). It takes apart the C language standard (C99) pretty much sentence by sentence, explains what it means and also contrasts how C99 is similar to or different from other languages (C++ mostly, but also, say, Fortran or Pascal).
Only thing I didn't like was goto chain part. I looked at both examples thinking one could just use function calls and conditionals without nesting. My memory loss means I can't be sure as I don't remember C's semantics. Yet, sure enough, I read the comments on that article to find "Nate" illustrating a third approach without goto or extreme nesting. Anyone about to implement a goto chain should look at his examples. Any C coders wanting to chime in on that or alternatives they think are better... which also avoid goto... feel free. Also, Joshua Cranmer has a list there of areas he thought justified a goto. A list of great alternatives to goto for each might be warranted if such alternatives exist.
Only improvement I could think of right off the bat on the article outside including lightweight, formal methods like C or stuff like Ivory language immune to many C problems by design that extract to C. Not saying it's a substitute for learning proper C so much as useful tools for practitioner that are often left out. Astre Analyzer and safe subsets of C probably deserve mention, too, given what defect-reduction they're achieving in safety-critical embedded sector.
1. What book do we assign?
2. What should we lecture?
3. What sort of code review work should we have students do?
4. What kind of assignments should we use? But only to say that he won't cover it in the article!
This is the almost the exact opposite order of what is most useful in terms of learning. Yes, some people (especially auto-didactic and well-focused students) are able to learn tremendous amounts on their own through books. But they are a relatively poor tool for teaching, compared to active learning methods. Lecture can be great, but usually is passive and worse than useless.
I want to acknowledge the importance of defining what you will teach and what successful (end-of-course) students look like and how to assess them. After you've decided that, it is proper to devise assignments and assessments, and then to decide on lectures and supplemental materials that support students in completing the assignments and assessments successfully. The time students spend should be active and practical - not that readings can't be provided, but they should be on-point and meaningful. Proper application of Instructional Design principles and theories of learning can make a world of difference for students.
But kudos for thinking about it, kudos for thinking about feedback mechanisms, and kudos for
PS: Obviously, I believe C has a great place in the curriculum - shouldn't leave undergrad without it!
When I first learned C in highschool I got a few books on C which all seemed to have the word 'Beginner' in the name. 'Absolute Beginners Guide to C' is one I remember in particular. I think having multiple books is pivotal because as a beginner if you encounter an explanation that doesn't make sense to you it is very hard to reason around it. You probably have very little prior knowledge, almost everything you know and learn up to the point where you get stuck will be contained in that single book, and if you don't know any other languages you can't make any connections to help yourself out. The reason the second, third, or fourth book is so important is that it will have a slightly different explanation that might make something click in your brain.
This has truly become something I try to keep in mind, considering a) I've later, sometimes long after starting on someone else's code base, learned a useful rationale for why they did some of the previously more inscrutable things in their code, and b) ended up writing a few things like that myself.
Documentation is key to understanding these systems, but it isn't sufficient. Often you are presented with a nicely documented mega-function, which while anyone can read through, but is very hard to reuse a portion of when needed. In breaking it apart into smaller chunks, you necessarily scatter some of the reasoning about why a particular approach was taken from where it was originally used, or at least where the weird behavior is required. You can either reproduce large chunks of the documentation at many different points in the code base, and hope it doesn't get out of date as the systems it describes in other files is slowly changed, or keep the documentation as fairly strictly pertaining to the code immediately around it, in which case the knowledge of how the systems interact can get lost.
Whenever you encounter code that seems to make no sense, it's better to assume there's some interesting invisible state that you need to grok, than that the programmer was an imbecile or amateur. The latter may be true, but assuming that from the beginning rarely leads to a better outcome.
Edit:
I'll share my favorite example of this. At a prior job, we had a heavily used internal webapp written in Perl circa 1996. It was heavily modified over the years by multiple people, but by the time I was looking in on it in 2012, it was a horror story we used to scare new devs. The main WTF was that it was implemented as one large CGI which eschewed all use of subroutines for labels and goto statements, of which there were copious amounts. The really confusing part was that they were used exactly as you would expect a sub to be used, just with a setting a few variables and a jump instead, so we always scratched our heads as to the reasoning for this. There was even a comment along the lines of "I hate to use goto statements, but I don't know a better way to do this, so we're stuck with this."
Fast forward a couple years, and I'm migrating the webapp to a newer system and Perl, and I discover the reason for this. At some point it was converted to be a mod_perl application, and the way mod_perl for Apache works is to take your entire CGI and wrap it in a subroutine, persist the Perl instance, and call the subroutine each request. The common problem with this is that because of this any subroutines within your CGI can easily create closures if they use global variables. The goto statements really were intended to be used just like subroutines, because they were likely switched to in an attempt to easily circumvent this problem. Now, there are better methods to combat this, such as sticking your subroutines in a module, and having your CGI (and then mod_perl) just call that module, which is what I ended up converting the code to do, but the real take-away is that the original decision, as impossible to defend as it seemed, was actually based in a real-world trade-off, and at the time it was done may have actually been the correct call.
There is some initial magic, where they have you import cs50.h which is full of black box functions in the beginning but other than that it's a good example of teaching beginner C.
It seems to me that while we know how to teach C properly today not many places do because they don't do as they say.
If you do stuff like that a lot you can easily end up wondering why your program is so slow even though the profiler says that there are no standout slow parts. Everything is just uniformly slow because their coding style doesn't consider the amount of work each statement maps to at the machine level.
Since someone always brings it up — I do believe as some point students should be exposed to a low-level language that doesn't to GC to expose them to memory management. Just not as a first language. I'd rather give them Python to start, let them get their feet on the ground and somewhat comfortable with it — give them success to start and hook their interest. C just tended to defeat the students, and needlessly. I spent so much time explaining things that to them must have felt like a rocket scientist telling them why their water bottle rocket exploded into flames on the pad.
And C needs knowledge in all fields recombined to be really used freely. Know one of those fields not - and you will be like a wanderer on a frozzen lake, doomed to trust those who know to guide you by ramming posts of no return where the ice gets thin.
Its also about taking a sledgehammer to all those certaintys people have about computers from marketing and personal experience as consumers.
I was going to write a Kindle book in the beginner's guide to C using Code::Blocks and its IDE because it is FOSS cross platform software. I found out it is a lot harder than I thought it was.
I learned C in 1987 at a community college still have the book on it that is written for Microsoft C, and we used Turbo C and Quick C for some of the assignments. Most of the programs I wrote can still compile and those that get errors or side effects can be debugged easily.
Most of university was Java and "use what you want", only a bit C++ for "Computer Graphics 2" (which I never did)
I found it a bit sad, but on the other hand I never needed it.
Transferred that to a UC and we definitely touched more ASM and C in the OS courses.
* CS 447: http://cs.pitt.edu/schedule/courses/view/447
* CS 449: http://cs.pitt.edu/schedule/courses/view/449
* CS 1550: http://cs.pitt.edu/schedule/courses/1550
447 and 449 are required, 1550 is optional.447 is almost entirely MIPS assembly, and goes into hardware architecture as well https://people.cs.pitt.edu/~childers/CS0447/
449 is a C class, using K&R 2nd edition: https://people.cs.pitt.edu/~jmisurda/teaching/cs449/2164/cs0...
1550 is more specifically about operating systems, using Tanenbaum's book: https://people.cs.pitt.edu/~jmisurda/teaching/cs1550/2164/cs...
So at least today, there's one class that's all C stuff. Almost all of the rest was Java, while I was a student.
Perhaps someone could explain what I'm missing. It's exactly the behavior that I see using gcc-4.8 and Apple llvm-7.3.
Not sure it's possible to teach green frosh 'why does industry use an old language' and the static analysis ecosystem (easier to teach skills than wisdom). But I applaud these people for trying. This feels like real programming.