Armin Ronacher: Collections in C
lucumr.pocoo.org
lucumr.pocoo.org
Having once had the pleasure of inheriting a codebase animated by CPP-macro collections, allow me to be a voice in favor of running, not walking, from programmers who embrace them. They're tricky, they blow up, they're extremely noisy in the code, and (worst of all) they don't encapsulate, offering new devs a myriad of ways to write tangly hard-to-understand subtly broken code that works directly with the collection structure.
Type safety is entirely overrated. C requires so much deliberation to deploy even the simplest collection that the likelihood of you picking up a foo where a bar was what was provided is minimal, and easily diagnosed. Just use voidstar.
One of the first things I did at that job was to port STLport's Red-Black tree (from <map>) to C code, specialized on voidstar. It worked beautifully. If there are times when you don't want to use a void* collection (or a gossamer-thin wrapper around one), those are also times when you don't want a generic collection library to begin with.
Why, there's even a whole book on that: http://www.amazon.com/Interfaces-Implementations-Techniques-...
Like Norvig's Paradigms of Artificial Intelligence: Case Studies in Common Lisp and Joshua Bloch's Effective Java, it's one of those books which despite having a specific programming language in the title, is really about programming in general.
I've written entire macro-based systems similar to those recommended by the author here, and find them very useful. Void pointers are great where they can be used, but I love the efficiency and clarity of the generated code approach.
For instance, in a chapter on string atoms (symbols) --- a concept which virtually no C program in the world takes advantage of, despite the centrality of the concept to Lisp, Python, and Ruby --- and this is the first chapter in the book --- you were shocked by his use of a 43 byte string to hold a number string... because someone might be using that code on a 192 bit machine?
You missed the forest for the trees. If you don't want your code tainted by the number 43, don't write that code. The point of the book is how you structure your code, divide it into subsystems, and present coherent interfaces to the rest of your program.
Anyway, I appreciate your response. It's good try to appreciate what others see that I do not. Yes, if you ignore the details of the code and the explanations, there are some good parts. If you look _only_ at the big picture, it's probably a fine book. And I love his clearly prefixed naming conventions. But I think you'd do a lot better reading something like the SQLite source code rather than this book if you want to see examples of good C.
I guess I have to ask: do you feel that chapter 4 on using setjmp() and longjmp() plus some brittle macros for error handling is also good for learners? I thought it was technically very clear but about 40 years out of date as to good practice. Is this a forest or a tree?
I wished there was a language (not C++) which could help for this kind of things. Unfortunately, it becomes difficult very fast. One interesting approach is K as suggested by some FreeBSD hackers, but it never went into production AFAIK (http://wiki.freebsd.org/K)
I am certainly not advocating doing this in general - I think the need for atomic support in generic collections is quite low (I have been investigating the issue recently to add fast and generic support for sparse matrices in scipy). I am pretty sure the macro, specialized ones used in freebsd (tree/queue.h) and linux (rbtree, list) have been benchmarked to hell, though, and would trust them more than most STL implementations.
When those are not issues, C++ is appropriate. Otherwise, it is a pain.
He could have literally written C-style code except used C++ just to create a few generic data stores instead of using all of this preprocessor magic which I guarantee is more fragile/harder to debug than basic C++ templates.
What does this mean?
As a simple example, suppose I use malloc to allocate memory (because I don't want to allocate with new which would oblige me to handle exceptions and prevent me from handing ownership over to a C client or using an allocator provided by a C client (without extra work)). I have to cast the return value as in
struct Foo *x = (struct Foo*)malloc(n*sizeof(*x));
instead of struct Foo *x = malloc(n*sizeof(*x));
This looks purely cosmetic, but in C, I can write the macro #define MallocA(n,p) (!(*(p) = malloc((n)*sizeof(**(p))))
which then works with a standard error checking convention struct Foo *arr;
err = MallocA(n,&arr);CHK(err);
This doesn't work in C++ without non-portable typeof or evil and less safe (not conforming for function pointers) *(void**)(p) = malloc((n)*sizeof(**(p)))
Similarly, if a client registers a callback with a context, I store their context in a void* and pass it back to them int UserCallback(void *ctx,...) {
struct User *user = (struct User*)ctx;
instead of int UserCallback(void *ctx,...) {
struct User *user = ctx;
I understand that this is just cosmetic. I don't see how "going generic full speed" helps with this. Also note that aggregate returns are slower for all but trivially small structures, and downright bad for big structures.It's got nearly every data structure you can imagine, all implemented as CPP meta hackery. Brilliant.
I went to go check out the code to see if this issue had been fixed, but the download link has gone bad.
I mentioned this library on HN some years back, when I was still excited about it, and the reply I got was something like "No! Not that macro boneyard!" He was right.
"I normally like to avoid tools that generate new C files for the very simple reason that these usually generate ugly looking code I then have to look at which is annoying or at least require yet another tool in my toolchain which causes headaches when compiling code on more than on operating system. A simple Python script that generates a C file sounds simple, but it stops being simple if you also want that thing to be part of your windows installation where Python is usually not available or works differently."
I have moved most of my macro code generation to using another language, with a proper template engine, to generate code (in my case C++ but the same holds for C). If he doesn't trust Python across platforms, he can take a fixed version with known properties and code around it, or take another language (which presumably will have the same issues). I use PHP, I'm a bit careful in cross-platform features and it works great.
Apart from this, he can still write his code generator in C, so that on a new platforms it can be bootstrapped with a regular C compiler, then process his templates, then compile his actual code. It's painful to do string processing in C, but this generator only needs to be written once anyway, and it's not much work. Plus if he uses a small template engine, that'll take most of the pain away (most of the work will be in modifying the templates, not the code generation engine).
http://library.gnome.org/devel/glib/
It does many other things, provides abstractions for threads, files, etc. It is a general-purpose utility library originally written for GTK+. I never really understood why GLib is not more often used.
Had he never heard of unions?
Even without knowledge of unions, you could treat a pointer as a, depending on the architecture, sequence of 32 bits. That sequence can be cast to whatever you want. So long as you have a clear way of describing what you've stored in that 32 bit sequence, you should never get confused.
Most of the time it actually works well.
This proves that you can do magic in C but it also proves that C is really a low level assembler :-)
For me, stuff like this has been one of the prime reasons to prefer C++ over C. You can totally get the same performance and compiled code when using templates. But you gain compile-time checking and much more robust code.
(These preprocessor tricks are also popular in the BSD kernel)
""" None of these macro names, nor the identifier `defined`, shall be the subject of a #define or a #undef preprocessing directive. Any other predefined macro names shall begin with a leading underscore followed by an uppercase letter or a second underscore. """ (s6.10.8)
https://github.com/lukesandberg/Regex/blob/master/src/util/f...
when you construct it you just tell it how big each item is then it copies the value into the stack element for storage.
Sometimes it a few extra copy/cast statements but at least its very clear what you are doing.
Variable Queue