C for Python programmers (2011)
toves.org
toves.org
From my experience, the best way for learning C has been [0] Build Your Own Lisp and Zed Shaw's [1] Learn C The Hard Way.
That and of course spending countless hours debugging segfaults and memory leaks.
Build Your Own Lisp implicitly claims you don't need to know Lisp to learn from the book:
> We will be covering many new concepts, and essentially learning two new programming languages at once.
Do you think that true? Did you know Lisp before reading it?
(I've done the first half or so of SICP so I know some Lisp.)
Right now I'm working on designing and implementing a virtual 16 bit CPU and an assembler, I guess the logical next step would be an OS class, then something like Build Your Own Lisp.
(For context, I write Python applications at work, but the work done by our embedded system C guys seems more interesting; I'm trying to learn enough C that I could be useful on one of their projects.)
In my case, I did know some lisp (Clojure) before starting out, but I'd strongly recommend that you don't hold back based only on that requirement.
It's a book written in a literate programming style that describes how to build a flexible and modular library of data structures.
// declare and initialize x
@uint32_t x = 42;
// dereference x
uint32_t y = ~x;
I.e. have different symbols for the type and the dereferencer ('@' and '~' as an example).//edit: thanks for the help, fixed snippets - HN's parser does not get pointers either apparently :)
EDIT: you can print one by leaving a space after it: * aa
C insists that int * , char * and int are totally different types that you should not mix. It makes sense most of the time but it can be confusing when you do not realize that internally they typically are the same thing <insert here the disclaimer about 32/64 bits systems>.
http://www.freepascal.org/docs-html/ref/refse15.html
It's quite similar to c - yet I find it a little bit clearer.
Which the declaration syntax hints at:
int i, *p;My point is that the concepts "having a datatype, that stores an address" and "reading (typed) data stored at a specific address" are different, but share the same keyword (which is in this case just one symbol - maybe because C programmers have to use it often).
E.g. assuming C99's bool (0|1) datatype, you might define the "negate" functionality, which is often done by '!'. So we have something like that:
// init a new bool
bool success = 1;
...
// negate the bool
if (!success) {
...
}
For me that makes sense. However if we take the pointer approach (as it is implemented), it would look like this: // init a new bool
bool success = 1;
...
// negate the bool
if (bool success) {
...
}
I hope this makes sense. For an experienced programmer * 's role is obvious from the context, but for me that was the most troubling concept when starting with C's pointers.The reason the same symbol is used is because of the way C types are read:
int x; // (x) is an int
int *x; // (*x) is an int
int **x; // (**x) is an int
int x(double); // (x(1.)) is an intHowever, I am still believing a different operator would make more sense.
They are different things, but C's syntax was deliberately designed to imply the former from the latter. "Declaration reflects use" is the term they used for it.
The idea is that in something like:
int *i;
You don't read it like, "Declare a variable 'i' whose type is 'int star', which is a pointer to an int". You read it like "declare variable 'i' whose type is such that taking 'star i' (i.e. doing a dereference) would give you an int". The type in this case is a pointer to an int.This is always why function pointer syntax is so totally bizarre in C.
"Declaration reflects use" was a neat idea, but I think it practice it ended up causing more confusion than it solved. At the time, maybe they thought users would be tripped up by compound type expressions and thought it would help if they focused on the operations performed on the variable being declared.
In practice, it turns out that composed types don't seem to be that hard.
I definitely think of (int * i) as (i :: Ptr Integer) and (* i) in an expression like 42 + * i as (* :: Ptr Integer -> Integer).
*i :: Integer
From which you can conclude i :: Ptr Integer.And the dereference operator does have the type you specified, but not in an lvalue—there the operator is really a mixfix one:
(*_ = _) :: Ptr a -> a -> a
This isn’t specified directly by the standard, but follows from the rules about how an assignment operator must examine the structure of its first operand.Depends on what you mean by hard. Hard to read or hard to reason about?
* -> * and * -> * -> * are easy to read and reason about, i.e. Int -> Char -> Int. But what about types of types?
(* -> * ) -> * is confusing to reason about IMO. i.e. Applicative Functors:
(<$>) :: Functor f => (a -> b) -> f a -> f b
Not really intuitive at all, the only reason I can understand it is that I know what functors and Monads are.
List<int>
(Some, Tuple, Type)
Matrix[4][4]
(Para, Meter) -> ReturnType> I don't recall that from the K&R
I don't recall either, but this Wikipedia mention of "declaration reflects use" cites K&R:
https://en.wikipedia.org/wiki/C_%28programming_language%29#C...
Yes I agree about the 'declaration reflects use'; what I meant was: does that imply that the declaration should consider the asterisk (the 'make this a pointer' part) to be 'part of', and thus right next to, the variable name or the type?
I'm strongly in the 'type' camp myself, and therefore I think that int *p; is nonsense; only to be used out of necessity when declaring multiple pointer variables on one line. So I'm wondering if 'declaration reflects use' reaffirms that, or contradicts it.
int i, *p;
Tells me that i is an int. And so is *p. int i;
int* p; int *p; /* What is the type of p? Pointer to int */
int a = *p; /* What is the type of *p? int */ struct Foo {
int a;
};
Foo* f;
, the following two are equivalent: int b = f->a;
and int b = (*f).a;
Never had it spelled out for me like this, probably because it's so obvious once you get it - thought I'd share in case it gives anyone else an 'aha' moment.for what its worth, your "aha" scenario would have just confused me more. even today, i have to think a bit before your second example makes sense to me.
Notational Machine for the win!
https://usborne.com/browse-books/features/computer-and-codin...
I confused that termination with the idea that my pointer has never been dereferenced in the first place. Now, I don't know if the kernel intervenes before the invalid address is going to be dereferenced or just after it. The thing is, I understood the concept, but getting this error was giving me doubt since I made an error of accessing valid memory.
Once I understood that I tried to force a normal int to become a pointer and eventually it worked (needed two ints and an 8-bit variable to cast it to a 64 bit address), proving to me that I understood the concept.
Accessing (valid) memory and pointers should be explained together.
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.34....
(for example being much more conscious about pass by reference, the cost of various functions, the overhead of calling functions and doing data conversions etc).
Every once in a while, you'll see a C program crash, with a message like “segmentation fault” or "bus error.” It won't helpfully include any indication of what part of the program is at fault: all you get is those those two words. Such errors usually mean that the program attempted to access an invalid memory location. This may indicate an attempt to access an invalid array index, but typically the index needs to be pretty far out of bounds for this to occur. (It often instead indicates an attempt to reference an uninitialized pointer or a NULL pointer, which we'll discuss later.)"
$ gcc -Wall -g prog.c -o prog
$ valgrind ./prog
...
==25494== Process terminating with default action of signal 11 (SIGSEGV)
==25494== Access not within mapped region at address 0x0
==25494== at 0x400532: main (prog.c:5)
==25494== If you believe this happened [...]
...
$
This helps a lot while debugging. Note that it is also helpful when debugging memory leaks and can give you the exact line in your code where you issued alloc() that you did not free later.Are there other such tools a casual C programmer should be aware of?
Scanbuild (LLVM based, Linux and macOS)
Perf (recent Linux) or DTrace (macOS, SmartOS, Solaris and FreeBSD)
Valgrind (Linux, macOS) with the KCacheGrind/QCacheGrind (CPU usage), Valkyrie (leaks and undefined behavior) and Massif-visualizer (memory usage) GUIs
The compiler sanitizers (GCC, LLVM, Linux and macOS)
strace (Linux)
HeapTrack (Linux, a true hidden gem, best memory profiler ever)
Binutils (Linux, macOS)
And really, get to know your debugger. Both GDB and the Visual studio debugger are extremely powerful. If you think a debugger does beakpoints and nothing else, you really, really need to get to know your debugger better. LLDB is getting there too, it will catch up soon (surprising given how long it took the other 2 to mature).
For people from a Python background, know that both GDB and LLDB are available as a Python shell with full access to the C program internal. You can add triggers to execute Python callbacks, conditional breakpoints and even gather stat and have them display in MathPlotLib or iPython Notebook.
https://gist.github.com/Elv13/48b43ead347faba1b59378267d2364...
export CFLAGS="-ggdb -fsanitize=address"
And few other "sanitizers". They add a runtime to the binary so you get the backtrace, detect runtime "silent" errors (undefined behaviors) and overall help develop C/C++ programs.
Syntax and data structures are usually the easiest part of learning a new language. Leveraging implicit conventions and trying to build anything useful is much harder.
This is still a useful reference for anyone looking to jump into C.
https://drewdevault.com/2016/05/28/Understanding-pointers.ht...
2 / 3 + 4 * 3
is not parsed left-to-right but as (2 / 3) + (4 * 3), int **foo[10];
is parsed as int *(*(foo[10]))
. What "declaration follows use" means is that the ultimate type of the object is exactly as it says after all the operators are applied (in the correct order): when the array subscript is applied to foo, and then that object dereferenced twice, the type of the expression is int. Thus foo is an array of 10 pointers to pointer to int, foo[i] a pointer to pointer to int, * foo[i] a pointer to int, and * * foo[i] an int.I mean, I've been writing C for 20 years and would still prefer to break down a declaration with intermediate typedefs rather than write "pointer to function taking array of pointers to int-returning functions taking a void pointer, which returns a pointer to int-returning function taking a void pointer". If you can write that as a declaration right first time you can give yourself a serious pat on the back.
int (*(*p)(int (*[])(void *)))(void *);
The process: ?which-returns (*p)( ?taking )
? (*p)( int (*[])(void *) )
int (*(*p)( int (*[])(void *) ))(void *)
Testing with cdecl.org [1]: declare p as pointer to function [taking] (array of pointer to function (pointer to void)
returning int) [, which] returning pointer to function (pointer to void) returning int
F-yeah. Can I has my pat? Seriously, it is not so hard if you used each component at least 100 times. That's few years of practice of non-hiding behind typedefs. You can't learn if you don't practice. (I'm not arguing that this should not be decomposed.)[1] http://cdecl.ridiculousfish.com/?q=int+%28*%28*p%29%28int+%2...
What confuses me about pointers in C is simply the syntax, and I'm sure it's because I just don't write enough C to be able to read it with any kind of fluency. I find C programs to be harder to read than any other language that I'm regularly in contact with. Of course this could be because of what C is used for.
Python is a high-level language: it provides tools to manipulate abstractions easily. C is a low-level one: it gives you access to chunks of memory (and far enough ropes to hang yourself while shooting yourself in the foot while drowning)
Assembly is not hard to learn. It has a reputation of being hard because it is tedious to use and that optimizations can be very intricate but learning basic assembly does not take long. Learn about MOV, ADD, JMP, INT, a few conditional jumps, learn how negative numbers are represented, what a memory mapping is, what a program counter or a stack counter are.
Then, go back to C. Read a bit about how function calls are made, memory allocated, and you will see that C is actually a high-level assembly language. All the hard parts of C will become obvious: a pointer is a variable that stores an address, a stack overflow, a segfault, a memory leak, all these will make sense very easily within that framework.
In the past, I had always avoided C because I didn't understand why many aspects of why it is the way that it is (eg pointers etc). Then I took my university's CPU architecture courses. That sorted that fear right out since I had to go from the ground up, learning everything I'd previously largely avoided or ignored - everything from transistors and the basic logic gates (AND, OR, XOR, NOT, etc) they consist of to full adders, all the way up to pipelined CPUs. Naturally, we learned assembly as part of this.
C made an awful lot more sense after all of this! I still don't like using it, but that's a lot more to do with my understanding of it's dangers (I prefer to let the compiler do the hard work of verifying my programs make sense, a la Rust &c), rather than a fear born of ignorance.
[0]: http://download.savannah.gnu.org/releases/pgubook/
[1]: http://asmtutor.com/
Doing simple reverse engineering challenges is also a fun and easy way to get a patter matching kind of feeling for assembly.
I agree. I also sort of learned it that way. The difference was that I learned Turbo Pascal first, before C, and BASIC before TP, but then either while learning TP or C, I also learned some assembly and something about the fundamentals of computer hardware and microprocessors (as related to programming, not electronics) in an interleaved manner. Was lucky to have access to some very good books on all these (and many other topics) from the British Council Library in the city I was in at the time. Gave me a solid grounding in many topics related to computers.
Coincidently, I was making small talk at the office today with an undergrad (EE major) intern who's unenthusiastically taking a course in C programming to satisfy requirements. I told her that although those two points you brought up were precisely why traditional CS types generally hate/avoid C, it's imperative that she grasp these concepts as soon as possible (along with picking up an HDL) if a career in hardware floats her boat.
Coming from a formal hardware background myself--and ever so envious of the traditional CS types who were always hacking the wee hours away--I ended up picking up a used copy of Cormen and took some CS baby self-steps over a few semesters, using both Facebook Puzzles[1] and Project Euler[2] as a condom so it'd be much harder to become impregnated with my own bullshit. And yet, to this day, I still feel like a bag of suckass compared to our resident graybeards; last time I hacked some C was only a few months ago to make a JEDEC[3] interpreter that metaprograms source in an obscure, obsolete proprietary language, but I'm admittedly ashamed to share the source with anyone out of fear of affirmation that I still suck hind tit at C after all these years. ><
[1] https://web.archive.org/web/20091130184215/http://www.facebo...
Maybe that's because you think that:
> the best way for learning C has been [0] Build Your Own Lisp and Zed Shaw's [1] Learn C The Hard Way.
Really, how can you get proficient at programming C without spending countless hours debugging segfaults and memory leaks?!
You get proficient at C programming by properly understanding what you learn rather than spend countless hours in trial and error cycles.
There isn't much magic to it and the concepts are quite simple. In my opinion if you're looking to get up to speed quickly instead of carefully writing code then C is not for you.
In your opinion, what would be:
1. the modern equivalent of K&R?
2. a solid reference to keep around and consult afterwards?
(with 1 and only 1 answer to the questions above, please, not a list, 'cause every time I search "learn c" I get over lists upon lists, of high quality comments - which is bad because you can't actually dismiss them as rubish, you actually have to read them -, that only throw you into paralysis by analysis, so you just say, "fuck it, just grab K&R or LCTHW, then grab a C open source project you feel like hacking on and start banging you head on it"...)
2. The C standard draft, your compiler(s') manual and man pages if you're on POSIX. Also pencil and paper for drawing your arrays.
Clarification about 1: Haskell has good solutions to a lot of problems that have plagued programming, learn them because they are not language-specific but are mindset-specific. The advantage is that once your mindset changes, you start paying attention to off-by-one errors and potential overflows and overruns. You also start paying attention to using the correct types in a C program and you start taking advantage of enums to reduce the space of possible states in each function in your program. Once this clicks, memory management is a breeze.
Also learn the mindset of solving the problem at hand and nothing but the problem at hand. C is not a language for creating abstractions everywhere, it's a language for writing down as little steps as possible to solve that specific problem.
I've always discouraged learning Python first. Python disconnects you so much from what is really going on and does so much for you that you don't develop algorithm and conceptual skills and computer knowledge.
All the developers where I'm at, from SQL to Python, Node to Angular, everyone, all say the same thing about Python and other "hyper-package-assisted" languages.
> #define forever while(1)
> Expert C programmers consider this very poor style, since it quickly leads to unreadable programs.
So why even mention it??? There are more important subjects which could have been introduced in this space.
Edit: And thinking on it our computer graphics course was in C++ (yea yea C != C++).
I wrote c wrappers for ada, and it was a pain. Much easier to use C (or C++) to get to OS functionality.
As long as OS are written in C, it will be with us. Also useful for embedded.
see: /usr/include
My OS class (a while ago) used assembly language for a simple VM, with one or two assignments requiring us to modify the VM itself to implement new instructions required for new OS features (task switching, virtual memory, etc.).
Personally I did study C via a couple of electrical engineering courses I took.
While I imagine it was good to get familiar with writing C, even assignments for operating systems classes didn't have all of the performance concerns or strict checking that production code would, right? (but C still is probably the best choice for other reasons in an OS class)
I asked my professors why and they argued that Python "doesn't get in the way" (static typing, compilation etc.) allowing us to focus more on learning the algorithms and theory.
I still preferred C to Java. I'm sure Python would have made that class way easier, though I do love pointers.
"A pointer is a variable that contains the address of another variable." [1st edition, 1978]
Two of the biggest stumbling blocks in C (and their ilk) are pointers and memory/garbage collection... which this brief intro barely mentions.
It actually reminded me of many programming courses I've taken over the years where the instructor spends an inordinate amount of time on the easiest concepts and quickly glosses over the meaty bits.
The book is small and pretty good. I found it useful. C is a small language thus the book is small (I think they say something to that effect in the forward).
Its book famous enough to have its own wiki entry:
If you want to write efficient software, which is certainly not always necessary, you're looking at a ~50x factor difference between the top and the bottom of that list, which even in this age of "let's just spin up another couple dozen instances" is non-trivial.
For example, look at the difference between the gemini-postgres and dropwizard-mongodb: https://www.techempower.com/benchmarks/#section=data-r12&hw=...
IF, and this is a big IF, you're doing something weird enough that you have to do the heavy lifting yourself, choosing an efficient language will matter a lot. At least for the heavy-lifting parts - you could still implement most of the the app in an easy inefficient language, and only do the heavy parts in C.
That's what happens in Python, and that's why my Python apps are only 3-10 times slower than pure C ("only" considering that Python is likes 100x slower than C).
And even then often a better algorithm is going to get bigger gains than a faster language.
And the same holds for memory. Your system has only so much of it, and for applications with many small objects most high level languages are very memory-inefficient if written in an idiomatic style. This can easily amount to a size factor of 20x in what's processable at all (disregarding execution time). Add garbage collection (which can "drown" processing) and the factor might even be higher.
Try adding a few Integers to a Java map and you will be surprised how inefficient it is.
The problem with writing things in C or C++ is in the time it takes you to finish the first working cut of a nontrivial program you could have done the same thing in Java and finished four or five tuning cycles. If you stop there the Java will most likely be faster.
How often do you get to write, say, a driver or a utility where "fast enough" won't do and there is enough of a business case to wring out those last few cycles?
Depending on your university's CS program, it might cover the equivalent of 2-3 semesters' worth of material without feeling like drinking from a firehose.
If you already know how to program, then you could just pick up K&R and work through it.
I totally disagree. I am not sure if your comment is satire or not, I would do it the other way around. Learn C as it is fundamental to modern OSes and just learn C++ if you have to (e.g. maybe later at a job, or because you need a specific library, like OpenCV). However I do not think that fully grasping C++ is a worthwile endeavor, as there are many things implemented due to historical and/or compatibility reasons (to C) and not because they provide a real benefit. Not that C is perfect but I consider it less fucked-up than C++.
C++ is a very high level language compared to C. C++ has low level constructs which closely resemble C but should only be used for high performance implementations of high level abstractions.
Not that I would discourage anyone from learning C++ (ok, probably, I might) I just consider it a bad PL for people starting out with programming.
More details here: http://scikit-learn.org/stable/developers/performance.html
CppCon 2015: Kate Gregory “Stop Teaching C"
A lot of developers still write in a C-esque fashion because that's what they were taught in school or just haven't learned a new way of doing things. Developers still use the C standard library for file IO instead of C++ fstream. Developers use new and delete with raw pointers and manage the memory themselves instead of using a vector. I was recently on a project where all the code was written in an old C style- variables declared at the top of each function, raw pointers everywhere, usage of C strings, etc. Memory leaks and segfaults were abundant. Uninitialized and unused variables a'plenty. And they were using a C++ compiler.
C++ and C, while both using the letter C, are for all purposes, entirely differently languages with different goals. Use the one that meets your goals best. While I can agree that K&R is outdated in some areas, there are still plenty of reasons to learn and even use C. If for no other reason, some of the most popular, active, and most used code bases in the world are written in C.
C is essential if you want to understand a lot of the ecosystems around you, fix them, interact with them properly, and maybe one day contribute to or patch them if necessary. I know some people will argue that if they know C++ (or programming in general) they can read C, but hands-on experience is the only real way to learn. You will not know why things are done or recognize when they are done poorly/incorrectly. Moreover, C is a great teacher of how things work, and sometimes also how not to do things. To be ignorant of C is generally to be woefully unaware of how a huge part of a lot of the software you probably use works.
I worked/work in game programming, distributed computing, AI, and many other fields. I find C to be extremely valuable in all fields I have touched and at home and at work. If nothing else, learning about memory, byte ordering, alignment, packing, bit manipulation, etc is a huge solid base. C++ can teach you a lot of that too, but it deals with certain issues on fundamentally different levels (especially modern C++) and has just as many potholes and here be dragons areas. It is true that I don't start as many new projects in C as I used to, but it doesn't stop me if lets say the main thing I need to do is interact with something already in C/embedded, or have a certain level of control that C affords me while being able to also hire people that can work in it (factors when evaluating project constraints).
One argument I consistently hear with regard to C that you hint at is that if you learn C first, before C++, whenever, it will somehow taint you. I would argue that with every language, you are not learning the language if you do not how to write programs in it idiomatically, or even when to counter-balance idiomatic code with code that does what you need to do (ex: boost performance or other tradeoffs). At that point, you are not a programmer but a parrot or like someone who memorizes math formulas but can't apply them.
Programming is not regurgitating countless lines of syntax, it is critical thinking, problem solving, creativity, and many more things. Learn C, and learn other languages. Do not listen to anyone who tells you to learn any particular languages. Match the language with your task, problems, and other constraints. When in C, do C. When in C++, do C++. When in Python, don't do C. And so on. It's really not that hard for anyone experienced.
Python like syntax, statically typed, garbage collected, C like perf.
C++ ("c w/ classes", i thought to myself):
- https://en.wikipedia.org/wiki/C++
- https://learnxinyminutes.com/docs/c++/
Python 2, 3:
- https://en.wikipedia.org/wiki/Python_(programming_language)
- https://learnxinyminutes.com/docs/python/
- https://learnxinyminutes.com/docs/python3/
C:
- https://en.wikipedia.org/wiki/C_(programming_language)
- https://learnxinyminutes.com/docs/c/
And then I learned Java. But before that, I learned (q)BASIC, so
SICP also uses that style throughout and I love that. Wish I could explain things that well.
Well, there is sizeof(array)/sizeof(array[0])
https://magic.io/blog/uvloop-blazing-fast-python-networking/
https://github.com/MagicStack/uvloop
Note how GitHub claims it's all mostly Python code ;) That's because Cython like I said looks Pythonic.
There's other examples, but I think this is one of the one's that come to mind the most to me.
There's also D which is called Native Python by some (unlike projects like Go and Rust, you can have your Object Oriented Programming (optional like in C++), and concurrency / parallelism too and other goodies like built-in unit testing, when you compile your code your unit tests are ran but not included in the final binary):
http://bitbashing.io/2015/01/26/d-is-like-native-python.html
https://blog.experimentalworks.net/2015/01/the-d-language-a-...
If it's been more than a few years since you've evaluated D you might want to check it out again, it may be worth your time. D is a language I knew about for years, and recently is where I've come to appreciate it for it's many features.
D has things like array slicing, [OPTIONAL] Garbage Collection, an amazing Web Framework called Vibe.d with it's own template engine called Diet-:
https://vibed.org/blog/posts/introducing-diet-ng
Things I like is Vibe.d is not just a Web Framework but a full networking stack too, also it supports Pug / Jade like templates (see Diet-NG) you make and compiles them when you compile your project, so your website runs off a native executable using fibers instead of threads. Vibe.d is undergoing a period where the core is being rewritten to where it is more compartmentalized so that you pick and choose which parts you need, MongoDB, PostgreSQL, layout engine and other goodies, there's even a templating library whose syntax is based around eRuby called Temple (though the syntax can be tweaked) that supports Vibe.d: