Learn C The Hard Way
learncodethehardway.org
learncodethehardway.org
What kinda bugs me is that whenever people go to teach C, they make out like it _has_ to be a low-level exercise, as if writing in C suddenly means you can't use abstract data types or object-oriented style or name your functions properly or have Unicode support.
For example, people teach libc string APIs like scanf() and strtok(), which should almost never be used. (See http://vsftpd.beasts.org/IMPLEMENTATION for one take.) Instead, use http://git.gnome.org/browse/glib/tree/glib/gstring.h or write your own like http://cgit.freedesktop.org/dbus/dbus/tree/dbus/dbus-string....
If you're going to display user-visible text, you are pretty much required to link to GLib or another Unicode library, since libc doesn't have what you need (unless you want to use the old pre-unicode multi-encoding insanity).
Don't use fgets() and other pain like that, use g_file_get_contents() perhaps, or another library. (g_file_get_contents is in http://developer.gnome.org/glib/stable/glib-File-Utilities.h...)
You need help from a library other than libc to deal with portability, internationalization, security, and general sanity.
Maybe more importantly, a library will show you examples that in C you can still use all the good design principles you'd use in a higher-level language.
I told someone to "use a string class" in C recently for example, and they said "C doesn't have classes" - this is confusing syntax with concepts.
C requires more typing and more worrying about memory management, that's all. It doesn't mean that all the best practices you know can be tossed.
There's a whole lot to be said about how to write large, maintainable codebases in C, and it can even be done. It's not something I would or do choose to do these days, but it can be done.
One other thought, two of the highest-profile C codebases, the Linux kernel and the C library, have extremely weird requirements that simply do not apply to most regular programs. However, a lot of people working in C or writing about C have experience with those codebases, and it shows.
I certainly agree that C is most appropriate when you are "low in the stack," just above the operating system and maybe implementing something like a virtual machine. I don't make a habit of writing stuff in C for no reason and the vast majority of programming these days isn't and shouldn't be in C (or C++ for that matter).
However, "low in the stack" is different from "low level" like "I refuse to use modern practices" or "I get to omit half the letters from my function names." You can be coding an on-the-metal kind of thing and still think about it in a high level way.
1. Building static libraries. 2. Building so and dll. 3. Binding those so/dll with higher level languages like Ruby, Python, Perl and PHP.
That is my main use of C these days. I always write the IO, user interface code in higher level languages and put all processing code in a so or dll written in C.
Having not to deal with exceptions, STL libs, makes life a bit easier.
Really though, there's no content here. The fact it's been voted to #1 on hackernews shows just how bad things are.
Items should be upvoted on their merit, rather than who did them. Someone writing another book about programming C isn't newsworthy.
1. The post is about Zed's draft; is it really so far of a leap to interpret hp's remarks as a criticism of the book?
2. Is this a fair paraphrase? "Everyone here isn't attacking you, also, your draft sucks and shouldn't have been posted on HN."
No, it's not. speckledjim is complaining that someone writing another book on C is not news, he doesn't say anything about the draft sucking.
However, "someone who wrote a ground-breaking introductory book in xyz modern-popular-hip language is turning his attention to C, of all things" is newsworthy.
eg. 'python for programmers' would not need the first half of it dedicated to explaining strings, loops etc. and could get straight into it from a programmers perspective - a bit like K&R. You could then dedicate more content to explaining philosophy, design decisions, internals, history, politics (learn all the in-jokes ;)), etc.
this would also be a good format to learn new paradigms, eg. 'functional programming in scheme, for programmers'
If you however dig a little deeper you can usually find stuff that isn't too tutorialish. Like http://www.c-faq.com/top.html which has been an extremely solid resource for me.
For example, a conversant Python programmer's idea of string, and strings in C are 2 different things. A conversant C programmers idea of looping is different from idiomatic looping in Racket(and other lisps).
I don't understand what "practicing the syntax" means (language barrier?), but I'm not suggesting that I've stopped learning, which I never will.
Rich Hickey, recently in an interview, said he doesn't do programming exercises - he is not interested in programs which don't make the computer do something useful, interesting, or both.
I, for one, type the trivial examples. Or else, I simply forget how to use them. For example, you have list comprehensions in Python, Racket, F# and if you ask me to do a list comprehension right now without looking up the reference, the only one I will get right is Python. I recently started learning F# and didn't type much code; so the concept is known, but since I didn't practice the syntax, I will have to learn it again.
That's what he means by practicing syntax.
I wouldn't however say that the complexities of C is it's syntax (unless you're playing around with macros : ) ). But I guess the concept can be applied broadly, like remembering functions, headers and other language specifics.
Or you could have books dedicated to a paradigm, such as functional programming, and have examples in many languages that all serve to drive home a specific point (this is how you make something tail-recursive, perhaps), pointing out how each language is both similar and different.
"The Practice of Programming" by Kernighan and Pike is somewhat close to what I want, but is not it.
Mark Pilgrim's http://diveintopython.org/ is what you describe.
Here is a tutorial for using SDL for 'old school' graphics programming:
Of course this would require the student to install SDL and write a makefile to link it all together. I feel that is something worth covering though, since most books just do a bunch of hand waving when it comes to linking external libraries.
Cheers, Salvatore
Frankly, I consider hackernews kind of irritating and not useful as a promotional tool. It amounts to about 5% of my traffic for 1 day at best, and I only have to be here because if I'm not answering the petty little idiots who troll here then the rumors spread through my professional circles.
In other words: I could live without this kind of promotion, so it's definitely not "self-promotion".
I am a bad ass of course. But I thought it was relevant to the comment that I've written a lot of C.
It is relevant - people not long out of school and talking at length about everything that's wrong with XYZ are tedious both to experts and experienced people generally. hp's comments on C use experience, clear reasoning, and absolutely no personal commentary whatsoever. This contrasts with the negative, personal and poorly reasoned attacks of people such as you.
I hope for your sake that you're asserting someone is "extremely" petty based on a mere self-deprecating aside because you're only a few years out of school and haven't learnt to communicate yet, and not just an old hand trapped an inability to relate to people with different opinions or role models. The world can always do with more talent and fewer martyrs.
This isn't the only possible interpretation of that comment. Sometimes words can have more than one meaning, and misinterpretation isn't always a sign of idiocy or acting in bad faith
Also, when you say, "This contrasts with the negative, personal and poorly reasoned attacks of people such as you," do you mean, "I hope for your sake that ... you're only a few years out of school and haven't learnt to communicate yet, and not just an old hand trapped an inability to relate to people with different opinions or role models"?
Genuine curiosity... what other interpretation could there be? Is it an ESL issue with literal translation causing offense? Or is there another interpretation in a non North American culture?
In N.A. at least, this is an extremely common expression.
I don't understand why other commenters feel it was unreasonable of Shaw to interpret it as a criticism.
I think it's been on the order of a year since I last heard this particular violin phrase. It might be common in your community, but that's not exactly North America, yeah?
I don't understand why other commenters feel it was
unreasonable of Shaw to interpret it as a criticism.
To speak for myself: because the other interpretation makes much more sense, if you take into account that people are generally trying to be nice and helpful.I was reading that comment, thought "hmmm, that criticism sounds a bit premature and more based on what others have written than on what Zed is writing... oh wait, that's because he isn't saying anything about Zed's writing at all. Yeah, now it makes perfect sense."
So admittedly, you get confused for a moment, but then there's a perfectly obvious solution: he's not commenting on the linked article, but is offering advice in the form of criticism on what generally goes wrong with these projects. If you don't see that solution and don't go for that interpretation, I think it means you fail to consider the option that someone is just being clumsy at being helpful. So it's another instance of Hanlon's razor: don't attribute to malice what could equally well be explained by stupidity (which is of course much too strong a term for an awkwardly phrased anwer).
BTW, it's not right that your comments are being downvoted.
You're always defending yourself against dreamed up conspiracies. You have a serious martyr complex.
May I clue you in? You create interesting quality projects like Mongrel, Mongrel2, Tir, and LPTHW. You are also entertaining.
If that's self-promotion, let's have more of it.
-- Zed Shaw, backpedaling when his angry schtick gets called out by the community
Here you denigrate the HN community, saying it doesn't generate a significant amount of traffic for your site or whatever, and then go on to say you won't respond to the petty "idiots". I see. You're too "professional" for that. Obviously HN is so meaningless to you that you needn't bother with it. Yet you're here, commenting, making an ass out of yourself.
Woe is poor Zed Shaw; always the victim.
So Im going to go with you having a selection bias.
OTOH I do appreciate your vigorous defense of the HN community, too many people just dont care about implied insults to their self identified tribal group.
Go llambda!
Nonetheless it's disheartening to see self-proclaimed "professionals" (I use this term loosely for while a person may be employed professionally this is not a guarantee a given individual will behave professionally: they are distinct things) who are well known throughout the community, behave like trolls. And even more so when they decide it's appropriate to take a shit on this community.
Well I will always try my best to preserve what I value. It's in my nature I guess. So thank you. :)
I, for instance, have been genuinely impressed with his productivity and his output, his choice of projects and his code and am moved by that aspect of myself that values quality to think very highly of him.
You, on the other hand, appear to feel that the most important aspect of his output is the way he achieves the standards that you have set for him - judged by how he expresses himself towards those who he feels are being disrespectful or rude at him on the internet.
Just out of interest, do you believe that making a strongly negative personal judgment based on a 'fact' that is clearly incorrect and could be easily checked with just a few clicks is the act of a professional? do you meet your own high standards?
I readily admit I'm a hypocrite. But also note, I don't claim nor ever claimed to be a professional; I'm not. I'm not at the helm of projects that are useful to and used by many people. I haven't written books on the topic of computer science. Nor have I given talks or do I run a website that receives a large volume of traffic. Maybe that qualifies [him] as working in the domain of a professional, yes?
Although he may have achieved great things, this doesn't give him a license to troll HN and it doesn't excuse him of baseless personal attacks. The comments he made that I replied to were no less than that, at best. At worst, they were insight into his character. Certainly we can hope the latter is not true. Nonetheless, there's no place for that kind of bullshit; it's inexcusable, I don't care who you are. I made such a judgement because he clearly attacked the user "hp" who had done nothing more than make an observation about texts that introduce C in general. Zed's reaction was puerile, irrational, and even paranoid. I hope you aren't standing up for such behavior?
Either you think he is a professional, or you do not. You appear to be holding up a standard, claiming that you believe he fits the criteria needed to be judged by it, and then lambasting him for not living up to the standard you have set. Overall, I find this somewhat confusing, but possibly I dont have sufficient context to judge.
TBH Im not particularly motivated to 'stand up' for the behavior of anyone I dont know.
I am interested in what makes you so interested in casting judgment on Zed, as opposed to on yourself?
Whatever puerile, irrational and paranoid responses Zed may have produced, you seem to be matching them with a self admitted hypocrisy, a large amount of self righteous vitriol and a bewildering statement of tribal affiliation to an anonymous internet discussion forum that apparently requires your outraged protection lest it collapses completely under the weight of a misunderstanding between Zed and another participant.
Zed has, to my knowledge, done you no harm of any kind. He certainly represents absolutely no realistic threat to HN.
why do you feel justified abusing him in this fashion?
It doesn't apply in this case. I will leave it as an exercise for the reader to discern why.
I don't know much about Zed. I've seen his name and articles here on HN, of course. I read you comment and was curious. I found nothing but a URL in his HN profile and zero submissions attributed to him here. From my very brief scan of his comments on HN posts, it seems like they are on topic. He's not hijacking threads.
So here on HN, at least, it appears that others are doing the promoting.
Zed seems to write original essays and he seems to have strong feelings and opinions on topics that interest the hacker crowd. And dude's got a memorable name.
It reminds me of SEO strategy: Write original, compelling articles and you won't need SEO.
Good luck, I'm all for more programmers understanding C but I wonder if the wonderful days of programming close to the hardware are ancient history. "[P]eople are deathly afraid of C thanks to other language inventor's excellent marketing against it." Maybe, but I think the raison d'être for C is not apparent to programmers who started with Java, Python, or Ruby.
The examples and exercises K&R uses will be very hard for beginners. When it builds a recursive descent top down parser to read the declarations in English(and vice-versa), that will be totally lost on the beginners.
K&R wasn't written for beginners and I doubt it will work well for a programming beginner - should be fine for someone who already knows programming in some other language.
Chapter 1, A Tutorial Introduction starts at the same place Zed's book starts: hello world. Looking at it now I don't see any reason my 13-year-old son (who knows a very little bit of Python) couldn't learn C from K&R with me explaining things here and there -- the same way I learned C. The recursive descent parser doesn't come up until the end of chapter 5 (out of 8), right after quicksort is implemented with pointers to functions. By then the reader has built up some skills and presumably developed enough curiosity to refer to other books for more explanation of sorting and parsers. For me those were Knuth's "The Art of Computer Programming," especially the third volume with all of the great example code which I busily translated to C.
If you think K&R is too hard for beginners consider how many programmers learned programming from the badly-typeset and somewhat inscrutable "Pascal User Manual and Report" (1974). That was the programming 101 text used at the local university when I was there.
"To many programmers, this makes C scary and evil."
Is this actually true for people? I find C code generally very easy and straightforward to understand; there's not any magic behind the scenes, like there is in any language that's more "high level" than C.
I heard them talking about passing arguments to functions as "call by reference" many times, and it was obvious they had no idea what they were talking about, just regurgitating what the instructor said.
Before that point, I had never clearly separated the concept of a variable and its value. It took a huge conceptual leap to think about a variable that didn't hold a value, but rather, the location of a value. It took some serious mental gymnastics to deal with pointers n-levels deep.
Mind you, this was actually Perl references, not pointers, so I didn't even have to try to comprehend doing math on them.
After a while, of course, it became second nature.
My favorite aspect of CS is that every so often, I run into a wall of conceptual understanding that requires completely changing how I think in order to move forward. Have you never had moments like this, or were they just different topics?
I remember reading those three pages over and over and experimenting with the code until I got it. I probably spent several days studying the pointer chapter back when I was a teenager. The light bulb eventually went on and C pointers made sense -- until learning C my programming experience was mainly with time-shared BASIC. The stepwise refinement of an indexed array version of strcpy() to a pointer-based version is a masterpiece of programming writing -- I still refer programmers I am mentoring to that chapter.
Finishing the strcpy() example with the one-liner
while (*s++ = *t++) ;
the authors write "Although this may seem cryptic at first sight, the notational convenience is considerable, and the idiom should be mastered, because you will see it frequently in C programs." I've found that programmers either get that line of code or they don't, and those who don't haven't mastered their craft.The "notational convenience" is more obnoxious than anything. But, then, I'm not particularly a fan of terseness for the sake of being terse.
"That sort of thing" has been in production in the C libraries, the UNIX kernel, and all of the brilliant utilities that make up UNIX for over 30 years. It's also very much in production code at Google.
You might want to read Paul Graham's thoughts on succinct code at http://www.paulgraham.com/power.html and Rob Pike's "Notes On Programming in C" at http://doc.cat-v.org/bell_labs/pikestyle.
Code should be pleasant to read wherever possible. That simply is not, to my mind, pleasant to read. It's one step away from Perl line noise (which I avoid, too). You may disagree with this, and that's fine--different strokes for different folks. I have no interest in seeing it in code I have to maintain; you might, and that's OK by me.
(To be fair, however, I have little interest in working with or, god forbid, maintaining C or C++ code under any circumstances. They press the buttons of a group of developers to which I don't belong.)
I myself never got much into the terseness game, since apart from a brief stint using gwbasic and later QuickBasic at 80x25, I learned programming using DJGPP in DOS with the RHIDE IDE. The IDE could trick the VGA into displaying something like 132x60, leaving plenty of room for descriptive code.
[1] I don't remember the exact gain, but it was at least 10%.
The idiomatic C style is so natural to me now I don't see it as a defect to fix or a game I'm playing. It's how I learned to program in C because I learned from K&R, and if they aren't the authorities I don't know who is. When I see wordy and bloated Java-style code it reminds me of the years I spent writing COBOL. Ultimately, though, it's not efficiency or a desire to show off or confuse other programmers that influences my style. To me code that does what it needs to with no extra fluff is beautiful.
When I see
while (*s++ = *t++) ;
I know what it does -- I don't need any comments or "descriptive" variable name or an explicit test for a null byte to make it clearer. If a programmer comes across a line of code that he or she doesn't understand the fault may be with the author, but it may be with the reader. In my experience there are a lot of unskilled programmers who quickly decide that any code they look at is badly written and should be thrown out. I don't think I should dumb down my code just so programmers less fluent with C can understand it. I have to consult a dictionary sometimes when I read Cormac McCarthy or Nabokov but I don't think they should write with easier words for my sake.Java is considered verbose for a number of reasons. It has a number of variable modifiers, it uses long names(which are generally good but can be stupid, especially in some cases of identifying multiple abstraction levels), and because it lacks type inference for generics and collection literals. c frequently makes the opposite mistakes with tiny cryptic or non whole word names. C lacks generic programming, it does do a decent job with initializers but unfortunately those can only be used at initialization.
Understanding the precedence and order of evaluation of operators is fundamental to mastery of any programming language (chapter 2.12 in K&R). All C code must "rely on" the order of operations, and if an indirect assignment through a pointer with post-increment is confusing... well that's the point I was making in the first place.
I never really understood why they didn't grasp pointers
The root of the problem is the language designers' loose use of star. "star something" is contained in a phrase that means one thing at declaration, "star something" has a different meaning the rest of the time. #include <stdio.h>
void eg(int i) {
int *j = &i; // "huh? Put the address of i into *j?"
*j = *j + 1;
printf("%d\n", *j);
}
int main() {
eg(4);
}
With more detail. In the line.. int *var = something;
.. the system assigns to the pointer. Yet in.. *var = 6;
.. it assigns to the contents of the pointer.Common usage creates further room for confusion:
int *var; // <- this is what people write
int* var; // <- instead of this
If they'd made the syntax ".int var", and then used * solely for dereferencing, people wouldn't have these problems learning pointers. Consider #include <stdio.h>
void eg(int i) {
.int j = &i;
*j = *j + 1;
printf("%d\n", *j);
}
int main() {
eg(4);
}
Further confusion comes from (1) special arrangements around string declaration and (2) printf use of %s to expect a string pointer when %d and %f expects (non-pointer) simple int and simple float. char* something = "huh? so now this does goe into *something?"; #include <stdio.h>
void eg(int i) {
int *j; // "j is a pointer and (hence) *j is an int"
j = &i; // "Put the address of i into *j? Yup"
*j = *j + 1;
printf("%d\n", *j);
}
int main() {
eg(4);
}
Perhaps this is because I am just used to it, but I really see very little room for confusion here. The common usage avoids confusion, if you do not insist on assignment at the time of declaration.To address your second confusion, just keep in mind that strings are char arrays and an array's name is actually a pointer. Again, I find this very straightforward.
I really see very little room for confusion here.
I don't understand how you reach that conclusion. You might understand it, I don't see how you can say there is very little room for confusion.Yes, if you know about declaration follows use, it makes sense.
Yes, if you "keep in mind that strings are char arrays and an array's name is actually a pointer" then it makes sense.
You can get by by knowing to avoid some constructs.
If you know how the c compiler works, pointers make sense.
If you know C then you know C.
But when you're new to the language, you don't know these things and that's what this part of the thread is discussing.
Another responder wrote:
The key is that * is part of the type of the
declaration, not of the variable; an int* is
not an int.
The grammar is structured as though it's not. Consider this: int* c, d;
Since star is part of the type, if the language was designed well then both of them would be int pointers. But in that case, only c is. d is an int. Awful.Do you mean to reinforce that the language is not beginner-friendly, or are you really asserting that makes the language poorly-designed? If it's the latter, you should really at least explore some other factors before making the conclusion. It seems to me that it's a relatively minor distinction once you know it, so from a design perspective that may simply be a tradeoff for some other advantage.
Do you mean to reinforce that the language is not
beginner-friendly, or are you really asserting that
makes the language poorly-designed?
I was seeking to account to ramidarigaz why his CS classmates didn't understand pointers.I think poor grammar is poor design - i.e. part of the type affects both variables (int), the other part doesn't (star).
The other issue is use of star to in one place to mean create reference, in others to mean dereference. That was the focus of my first post.
from a design perspective that may simply be a
tradeoff for some other advantage.
I've yet to see evidence of any. What sort of things were you thinking about?> I don't understand how you reach that conclusion. You might understand it, I don't see how you can say there is very little room for confusion.
Please do not put words in my mouth. I wrote "I see very little room for confusion" not "There is very little room for confusion". I tried to make it clear that I was talking about my personal experience; and I was talking about my personal experience since I was hoping it would help, not to defend the syntax of C.
Sometimes a particular point of view allows you to understand something; in some cases it makes the previously mystifying point "trivial" or "obvious". I am sorry that the POV that helped me so much does not help you at all.
The point of my earlier post was to describe strong reasons for people to have trouble with pointers.
int *j = &i;
is more correctly expressed and easier to understand when written like int* j = &i;
The only reason to put the * in front of the variable name is when declaring several pointers in one line. So the solution is to only use it that context, or not doing it at all.Stroustrup wrote something somewhere where he explained that int* a; is more appropriate for use in C++ because C++ is supposed to be more focused on types, and int * a; is more appropriate for C because of something about C's philosophy, but I can't remember what. I wish they would have changed the syntax for C++, but I guess he couldn't have while still keeping C++ a superset of C.
edit: found it: http://www2.research.att.com/~bs/bs_faq2.html#whitespace
"A ``typical C programmer'' writes ``int p;'' and explains it ``p is what is the int'' emphasizing syntax, and may point to the C (and C++) declaration grammar to argue for the correctness of the style. Indeed, the * binds to the name p in the grammar.
A ``typical C++ programmer'' writes ``int* p;'' and explains it ``p is a pointer to an int'' emphasizing type. Indeed the type of p is int*. I clearly prefer that emphasis and see it as important for using the more advanced parts of C++ well."
The main issue with C is it takes some time before you are ready to take it head on. An experienced C programmers would have his repertoire of generic data structures library with time complexity guarantees (programs without hashes, expandable lists and operations on them are a pain), will know how to properly use function pointers to do that dependency injection thing other programmers are raving about, separate interface from implementation, know the build environment, know how to use structures and function pointer to build abstractions etc.
But before that, C takes much work to produce little. For someone starting programming, the learning curve is steep. Or more like, the gratification is really, really delayed. It takes some time before he can take on a real world project(it does in the high level languages as well, but the initial progress is faster).
What happens when you don't get a clean segfault is what got C the "dangerous" reputation.
Buffer overflows do not produce cores, they just sit there until a determined cracker makes use of them. And a dangling pointer might still access memory that looks valid both to OS and memcheck.
Seeing that you are the guy(or onof the guys) behind redis which is written in C, do you have any recommendations for generic data structures and operations on them?
I personally have a trivial vector implementation which resizes when full, and a red-black tree implementation for associative arrays. Both of them work fine for my purpose - does the job, good locality of reference, generic over void*.
I have seen glib but largely neglected it because I only need a very small part of it.
See also Rusty Russell's rules about "easy to use versus hard to misuse". I find C easy enough to use, but also easy to misuse.
Yes, I do. But it's not because of the language itself. It's trying to figure out what you can do with it after you grasp the fundamentals. I taught myself C using K&R a long time ago, but I never did anything with it. At the time I figured there were two paths I could progress along - UI related (e.g., a Windows app) or systems related (something Unixy). Both paths presented large hurdles. Nothing insurmountable, but I wasn't a programmer at the time - just doing it for my own edification. I always thought it would be nice if there was a second level book that took you from post-basics to writing something useful.
FWIW, I enjoy programming very much, and have built some non-trivial stuff in several languages. Still, I found C extremely taxing. Not in the sense that I felt it was beyond me, but in that I was fighting or recoiling from the language at practically every turn. In what follows, I am acutely aware that I am nowhere near fully fluent in C yet, and am writing only to offer the first impressions of a student. Nevertheless, in that time I have been able to draw on the advice of several "experienced colleagues", as K&R urge. In that sense, if I misstate the facts, I will be repeating the misconceptions of people who have objectively spent an awful lot of time writing C. I would be glad to be corrected on any point, but I'd also find that kind of symptomatic of the issues I have with C.
Firstly, C is incredibly stateful. I would be glad to learn that I am just Doing It Wrong, but there seem to be no obvious way around stateful manipulations as a way of life. Take, for instance, the fact that arrays are essentially second-class citizens and must be stuffed into functions as (pointers to) extra parameters in order to capture the "function's" "output". I am breaking out the scare quotes here since it is an abuse of vocabulary to refer to a procedure that communicates with the world by modifying its inputs as a function. If you have not cut your teeth on it, it seems almost obscurantist. If I want to multiply A by x and store the result in b, I want to write
b = matrixProduct(A,x);
not
matrixProduct(b, A, x);
as I must in C. I go back and forth over whether this is a deliberate and worthwhile performance tradeoff or just myopic design[1], but I don't want to have to settle for this in code I read every day.
There is also a kind of bureaucratic spirit in much C code, resulting from the fact that one must attend so closely to the how of computation rather than the what. In some contexts, like when you're doing distributed simulations that may run for days and performance is at an absolute premium, this emphasis may be appropriate. In most contexts, however, one finds oneself implementing and reimplementing standard operations by hand. Why should it take four lines to sum an array?
One of the most alarming symptoms of this style can be witnessed by watching an experienced C programmer read code. Old hands don't read lines, they scan an entire section of a page at a time. I was awed by this ability until I realized that it is possible only because each line of C does so little. I recall someone (probably in an HN comment) describing the "rhythm" of reading C code. That, to me, is not an encouraging sign.
Then there is, of course, the penance of debugging segfaults and space leaks. What happens when you declare 10 pointers and allocate 9 of them? No one knows, because C doesn't know either. It's up to the compiler. Towers of Hanoi, nasal demons, &c. Even with the help of with smart and experienced people, I've never spent so much time diagnosing such trivial, silent runtime errors. Yes, I know it's that way for a reason, but that reason is not legibility or ease of understanding.
At end, C has performance and a relatively (but not exceptionally) compact semantics. The importance of cycle-squeezing is becoming less of a consideration daily, for reasons well-rehearsed on this site and elsewhere. As for C's semantics, I'd much rather spend my time thinking about the transformations I want to map over my data than orchestrating von Neumann machines to carry those transformations out. If your model of computation is something other than register machines, it's not quite cricket to describe the implementation as magic.
-----------------------------
[1]I'd like to qualify that remark in two ways. Firstly, I certainly recognize the brilliance of K&R for working C from the raw conceptual materials of the time. Secondly language designers in 1969 did not have the benefit of the last 40 years' of object lessons in readability. There are many potential languages--points in language space, if you will--that are semantically identical or near-identical but much more readable than the ansi standard. For an example given by Kernighan himself, postfixing the dereference operator would have done miracles for legibility. Ultimately, a lot of more or less arbitrary choices were made early on and now it's too late to correct them.
Yes, "under the hood" it might be the same, but that's true of many languages if you dig deep enough. C is just a little closer. Yes, at one level a char declaration is just a smaller minimum memory allocation than int, but C will check both those types and if you want to use C properly you'll want to understand its typing and casting rules.
I spent many years believing it when people made exactly the assertion you have(see my other posts on this article), but it wasn't until I tried to build a C compiler myself how wrong it is to think of arrays this way. Yes, C gives you the power to reference memory in a more or less arbitrary way. That does not mean that the arrays you declare are not arrays.
If you are writing firmware, for example, the exact sequence of memory writes is critical. Program the hardware registers in the wrong order, and the device doesn't work. Access the FIFO the wrong way, and your ISR has a data race. Hiding memory access from the programmer is useless when accessing memory correctly is the problem. C is pretty much the only usable language for this kind of work.
The guys writing kernels and system-level libraries face similar issues.
So, C is stateful because the hardware is stateful. C has raw pointers because the hardware has raw pointers. C doesn't manage memory for you because you don't want C to manage memory for you. This all makes sense when you realize that C was invented for writing operating systems.
So, while I do understand your criticism, I think you are looking at this from the wrong angle. The electrical engineering guys build the hardware, and the C guys make the hardware boot up. Fancy functional programming languages are useless without real machines to run on, and C makes those machines go. It's part of the plumbing, just like transistors. Plumbing may be messy and unpleasant, but even the architects designing skyscrapers need to know how it works.
[] malleable in the sense of bottom-up programming and defining your own "vocabularies". When I develop in C (and C++) I usually solve the problems 75% bottom-up and 25% top-down. With bottom-up approach, it is easy to stay focused at the problem at hand and write mostly bug-free code. The "top-down" bits just put pieces of the puzzle together.
I won't defend C: it's all that you say. I think one of the acknowledged reasons for C's longevity is that it is impedance-matched for UNIX because they grew up together, and for various reasons UNIX is popular with hackers; therefore C is popular. C is deeply rooted within UNIX, partly because the ABI has been so stable for so long. If you want to write a library for UNIX, you generally target C because anything else will run into a quagmire of cross-compiler incompatibility issues. That means all the good libraries on UNIX are written in C or present a C ABI/API. The easiest language to use a C library from is... C. Or C++. So application writers (and tool writers) have often favoured C as well, although the rise of interpreted languages such as Python, Perl, and Ruby has changed that a bit. Also, the C/C++ toolchains have tended to be more advanced than those of other languages.
So, C is still with us, but for reasons that don't have as much to do with its merits as a language as with its merits as a platform (when coupled with UNIX).
I greatly appreciate an 'opinionated' programming book. I've probably heard more debates on formatting and style for C than any other language.
Plus even if you do primarily program in higher level languages, it's a great tool to have in your belt when you need to fix a bug in a library whose bindings you're using in higher level language, or when you legitimately do need to eak a little more performance from a particular piece of code.
Also, love that the book starts by teaching you how to use make as well, so many C books gloss over the tools.
If you've got others I'd love to hear them.
The hardest part for me was learning the patterns necessary to do anything, and specifically to do it remotely well.
For example maybe one of the exercises could be some kind of board game, where you demonstrate how you translate "the ideas" into "the C code"? For example just because you know that a Knight can move two-plus-one spaces doesn't mean you have any idea where that code should be inserted into the program's structure, or how to write it in such a way where you avoid overrunning an array. (The naive solution would crash when the Knight tries to move off the board.)
Translating ideas -> C code was easily my hardest task when I was first starting out.
there are a lot of C devs who don't understand what it is their code is producing, and how an application and memory are managed
I understand all of it now, but had no idea about the stack or the x86 instructions, etc for years. I was still productive, and wasn't hindered.
The details can come later. Even big details like "how it works at the low level".
The most important part is to keep things fun; LPTHW was fun. For me, mucking about in assembly wasn't, and that sentiment seems like it might be common among new C programmers. Assembly has a way of slowly steamrolling your motivation.
This most often comes up when people are asked to do some binary manipulation of numbers, i.e.:
unsigned u = 19;
unsigned v = u >> 1;
"v" is now 9, and to really understand it one must grasp how numbers are represented in binary under the hood.People also fail to understand strings:
char* s = "string literal";
To some, it's absolutely opaque that the first byte "s" points to contains 0x73, and that represents "s" in ASCII.Correct me if I am wrong, but I don't think C guarantees anything about the binary representation. Depending on the architecture, `v` can have different value.
You might be thinking of character representation for the later example.
Is there anything that says, for example, it can't be BCD?
The C99 standard defines the >> operator in terms of division by powers of 2, so one can determine the result of 19 >> 1 without needing to know anything about how 19 is actually represented by the machine.
Be careful when explaining the compiler and how a program actually runs. I've found that a lot of student problems come from "the compiler is magic" when it really isn't (maybe related to my other surprised comment below, about how people "don't get C"--they attribute too much magic to the compiler.) Maybe even emphasize that every piece of C code can be translated in a fairly easy fashion to a pretty small amount of assembly.
For the preprocessor, emphasize that it is a solely textual replacement, with no symbolic evaluation. Explain why:
#define FOO BAR + BAZ
or #define MAX(x,y) (((x) < (y)) ? (y) : (x))
will go horribly wrong (the first in 5*FOO, the second in MAX(x++, y++)).For example...
const char* current = "ohai thar";
const char* end = current + strlen( current );
assert( end >= current );
size_t bytecount = (end - current);
(size_t is an unsigned type, so if 'end' is less than 'current', it will overflow. If you want to allow for that, use ptrdiff_t.)I always thought that "all identifiers with a leading underscore were reserved". I just consciously ignored it, and have never had a problem in years. But I was also always using member variable names like "_children", "_childCount", etc, not "_Children".
This is the big selling point for C++ : the convenience of STL. In all the projects that I worked, it was one of the important reasons to choose C++ over C.
Yes, this focuses less on the language itself and more on the ecosystem around it, but I figured if you have an entire chapter dedicated to "make", then this makes sense as well.
- The concept of a variable, and variables changing over time, seems quite difficult for people to grasp even when explained a few different ways. "x = 5" followed by "x = 12" proves quite mystifying, and "x = x + 1" even more so. People seem to have the most success with the idea that declaring a variable "int x" creates a location x which can hold an int, and you can put things in that location.
- Pointers actually don't seem to trip that many people up initially, once they get to that point. However, I don't think people actually understand exactly what they do, so much as memorize the rules for dealing with them. The same idea of a location to put something applies here too.
- Any case where the same function gets called more than once often ends up tripping people up; this applies particularly to recursion, but it can happen even when just calling the same function several times. In particular, this often interacts badly with people's understandings of variables. People need some understanding of scope.
- Combining several of the above, it would help to have clear explanations of the interactions between pointers, locations, and functions. Bonus for explaining what goes horribly wrong if a pointer refers to something that goes out of scope. That concept requires understanding several different pieces of C and putting them together.
If you set out to only teach the language syntax and paradigms you are leaving a beginner with a lot of extra work before they can start or contribute to meaningful projects. This is the reason that people learn so much from reading other people's code in open-source projects: most writers skim over trying to teach the most fundamental skill of programming.
I do believe it is possible to teach practicalities in addition to theory and if you attempt to do this, you will be doing a lot more than most writers have done in the past.
y=1+x=3;
Why and how this works is foreign to most beginners. Likewise (ok, this is C++ but the point remains)
cout<<1+x;
Isn't a command to display 1+x. The output is the result of inserting the computed value into couture. There is a semantic difference.
Groking operators early is key to understanding C well.
For people who already know programming, the official tutorial is the fastest and the best way to start: http://docs.python.org/tutorial/index.html
The critical factor, IMHO, is clarity and conciseness of writing. Zed does a good job with this (I didn't know his name until this thread, BTW).
With that said, when I first saw Learn Python The Hard Way, it immediately reminded me of the time I spent reading K&R2. Its exercise rich style is what programming books are lacking these days. K&R2 taught me how to program.
I hope Zed actually covers how dangerous format strings can be if not handled properly. Format strings are still (hilariously) one of the major exploitation vectors in C-based applications today.
Edit: According to Wikipedia, %i and %d are synonymous. Sorry for the confusion.
As for how "dangerous" it is, yeah it's not going to rape your family, it'll just crash. So I'll be showing people how to prevent it.
Well, if it hits undefined behavior it could in fact rape your family, the standard allows for that.
Intelligent people support Tau: http://tauday.com/ This is the least Zed could do.
Thanks for the explanation. I've never seen/had to use %d with integers. I've only used it with %.Nd for floats/doubles.
Mentioning code execution and shell code isn't really in the scope. If he mentions code execution, then it sort of warrants mentioning modern architecture prevent executing data as code, or the code segment is not writable on many architectures...and shell code is basically the op code that your machine executes, and injecting shell code in absence of any protecting mechanism will execute arbitrary code, and in presence of protection mechanism, it will crash.
I think just mentioning shell codes and code execution aren't doing a beginner any good. And explaining it is well out of scope. It's not that this is going to be the end all C book, and as long as the reader sticks to using proper format strings, he is good. If he doesn't, knowing about what might happen isn't doing much good either.
Another important point to mention is that the format string itself should not be user coercible! XCode/LLVM/whatever Apple is using nowadays actually treats non-constant format strings as a compilation error, which is pretty cool.
You should probably read this book. :-)
I'll pass on reading it since I don't intend to write C ever again, but I hope the book works out well for you. :)
Edit: Just dug up some of my old C code. It is indeed %.2f that I was thinking of. Sorry for the confusion once again.
Edit: Actually, I finally remembered why I always use %d, and why I find seeing %i odd. While %d and %i may be the same for output with printf, they are different for input with scanf. (Also I'm not sure that %i was part of the C89/90 standard for printf, it may have been a compiler specific thing (another reason to avoid); it's defined in C99 but I don't have a copy of the 89/90 pdf.)
At least I didn't fully appreciate C until I understood some of the underlying concepts.
Looking forward to reading this!
One of the best things about LPTHW was the context it was written in, and if LCTHW is written in the same way, it should be a really awesome read!
I'll definitely be waiting anxiously for you to finish.
Good luck!
This guy loves to program
That is, the ecosystem of, "What can I include without dicking around with compiler and linker settings, which I do not care to learn very well because I am just starting?", and the ecosystem of, "Why are all these standard libraries full of functions that all the documentation tells me not to use?", and most importantly, the ecosystem of, "Oh, this looks like a nice library that would make my life easy, (and later), wait, why isn't my program working on this other machine? I copied the binary over? Wait, what's this about a missing so? Oh, that's the library I installed, wait, how do I put it the same directory? Oh my god so many configuration settings! Wait, why can it still not find the library? It's right there now! What's LD_LIBRARY_PATH? LD_LIBRARY_PATH is bad? Why doesn't -R work? Oh, that's only for Solaris? What's the equivalent for friggin' Linux?! Ah, rpath, wait... it's trying to find the ABSOLUTE path? UGGARRHGHGHAAHHHH! Okay, finally, $ORIGIN... now let me just put that in the make file like they said I should.... AHGHGHGHGHGHGHHHHHHHHHHHHHHH!!!!!!!"
Which is to say, the ecosystem of fucking ratholes that have built up over 40 years of poor tool design that cannot be corrected now due to historical precedent.
To add another example, binary only distributions (e.g. some commercial software) bring their dependencies with them, meaning that you have to use roughly the same environment (compiler version, stl) to use them.
Libraries work a lot better in VM languages.
With C and C++, you'll have to use your operating system's package management to get all the important libraries (or build them by hand from Git sources, etc), then have a build system that configures the build environment and searches for all the libraries and other dependencies. It's not as nice as using a dedicated tool for this, like Gem and Bundler in Ruby but usually you get the job done - unless you work on Windows and don't have a package manager.
There are many alternatives for you that offer 1click build/deployment.
Just google for "msvcrt_win2000.obj" and see the madness (yes, it's about using MSVCRT.DLL instead of later MSVC libs, and still get your shit working on 2000 or XP). I did that just last few hours :)
There are way too many subtle details (and more complicated with Windows's manifests, side by side assemblies and crap like that).
There is one that I can recommend, for the record, at least for C++:
Qt's QMake.
It's sad, but the best user experiences I've had were either hand-coded makefiles, or autotools (ick).
http://labs.qt.nokia.com/2009/10/14/to-make-or-not-to-make-q...
http://lists.qt.nokia.com/pipermail/qt5-feedback/2011-May/00...
The language C is a big problem for beginners, though. Pointer syntax is not just a tricky optional feature, it's necessary for a number of common tasks including defining functions and passing parameters. The type system is also important, not intuitive and rarely taught effectively. Countless times I read or was told that the syntax for referencing arrays was the same as referencing a pointer in memory, but nobody ever bothered to clarify or reinforce the idea that arrays are still a distinct static type.
Pointers and types are fundamentals and anyone who was lucky enough to learn them early on might not understand how hard it can be to figure this stuff out on your own and how difficult it is to actually use C before you do.
Why a string must end with a 0 byte.
Binary shifting numbers, especially signed numbers.
Memory management; it's fun for the whole family, and more than just knowing how to allocate memory and store a pointer to it.
Some other points that I'm sure I'm missing.
Unless it's a pstring or you're working with assembly, other language, etc.
But yes there is a lot to learn with C you don't necessarily get with higher level languages in most cases.
Learn how the language works, follow good patterns, and expect some bumps in the road.
Most of the time, though, you can use a different higher level language, and be a happier and less stressed programmer
But ooh, Columbia, so you know it's hardkore.