Macros on Steroids: How pure C can benefit from metaprogramming
hirrolot.github.io
hirrolot.github.io
libcello has quite different philosophy than Metalang99 and the accompanying libraries like Datatype99. The thing is that the former comes with its own runtime environment, but the latter ones even don't require the standard library. I believe the philosophy of C is closer to Metalang99/Datatype99/Interface99.
P.S. I also did metaprogramming in C by writing C programs to generate C code, such as converting an array of structs to multiple arrays, one for each field (for efficiency). This worked out fairly cleanly, but with D's ability to run functions at compile time, I was able to put this back into the main program.
Starting small with a small utility such as this one or creating one of your own is the best aproach in my opinion.
It is basically impossible to provide a "kind" development environment using deep macro trickery. The error messages you get from this kind of library are a major handicap. This failing is intrinsic to the landscape and not something you can really fix.
Historically, macro processors in major compilers were inconsistent, limited, and arbitrary, making heavy macro use frustrating, limited, non-portable. I think this is mostly historical now.
IMO people should prefer boost::preprocessor for purposes like this (even in pure C), since it is battle-tested and more polished than "little" metaprogramming libraries will likely ever be. (This is advice to possible users, and not a knock on the author. Props for exploring the space and building something useful.)
In fact, I used to program in boost::preprocessor before, trying to implement ADTs [1]. At the end of the game, I just realised I didn't have enough patience and whatever else to continue using it. I could accidentally see that some macro got blocked and then debugging it for a half of a day to eventually come up with ugly workarounds.
Then I made Metalang99. It has many neat features but really the very motivation for it was to get rid of macro recursion blocking. For example, if you call FOO, then FOO calls BAR, BAR calls JAR, and JAR calls FOO, everything should work as expected. This was a complete nightmare with boost::preprocessor.
So in conclusion, I agree that boost::preprocessor is more polished and can work on ancient compilers too. However, since nowadays the situation with preprocessors is better, I would endeavour to use more modern tools which offer greater convenience and flexibility.
On the boost front, I should have noted the difference between VMD [1] and boost::preprocessor [2]. I think VMD [1] trades some backwards compatibility for coherence and consistency and is a better starting point for new code.
[1]: https://www.boost.org/doc/libs/1_70_0/libs/vmd/doc/html/inde...
[2]: https://www.boost.org/doc/libs/1_76_0/libs/preprocessor/doc/...
C++ effectively lets you do this too. And Zig will let you import and compile C files seamlessly. I suppose you have to switch compilers, but that's not so different to adding the additional pre-processor compiler.
The thing with Zig is that, if I understand it correctly, Zig can't be interleaved with other C code in the same file, which forces you to separate things, loosing convenience. On the other hand, to use Datatype99, you can just #include <datatype99.h> and you're ready to go.
Even in c++20, there are things you can do with macros you can't do in the language. This may change when reflection lands in c++2x, but I won't be able to use it for years and widespread adoption won't be possible for another decade.
Back when I was working in C#, there were quite a few times when I'd encounter things being done with reflection that I wish had been done with less code and a more type safe manner by just... not using reflection. I always just assumed a lot of that was just bad habits left over from the .NET 1.0 days, back before generics. So I'm curious if there's much overlap there? If it's only a subset, I don't know C as well as I'd like, so it would be interesting to see something that the C preprocessor handles better.
The typical implementation is something like this:
public string CustomerName
{
get
{
return this.customerNameValue;
}
set
{
if (value != this.customerNameValue)
{
this.customerNameValue = value;
NotifyPropertyChanged();
}
}
}
When you have 20 properties you have a ton of repetitive code. With a preprocessor I could reduce each property to a one liner. You can also do tricks with reflection but that's way more complex to implement.Or were you really just meaning to advocate for metaprogramming in general?
In this system, I write a given function in C and surround it with a macro wrapper. This wrapper generates a JSON-RPC equivalent of the C function using the Jansson JSON library [1], complete with argument typechecking and error management. My C code can call the function normally, and an embedded webserver [2] can export it across the network where it's used fairly transparently by a Python stack on a control PC.
All three of these "embedded software" elements (the embedded webserver, the JSON library, and the FPGA interface) are very well suited to C. I have recently dabbled in moving parts of it to C++, and have found a few surprising impedance mismatches. (As I mentioned above, C++ still lacks reflection.)
The macros, here, are not for "elegance" or any other abstract purpose. They have allowed a ~80% reduction in boilerplate and an equally impressive reduction in the bugs and hassles associated with this kind of low-density, high-sprawl code.
Might be nice to bring it to their attention, if you haven't already, and if you have the time and inclination.
Unless it's in a language actually designed for meta-programming. Obviously if you are already in a C project this might not be practical, but if you need a lot of it you may just be using the wrong tool for the job.
I don't have a ton of experience w/ C, but the large projects I've seen appear to have a tendency to avoid heavy macro usage (e.g. using `void *` for hashmap implementations instead of reaching for macros, for example). I see macros appear more frequently among those that prioritize machiavellian cleverness, often under the pretext of squeezing performance: code written by (g|x)ooglers, for example.
I'm curious what more experienced C developers think of jumping into projects with heavy usage of macro-based DSLs, especially the syntax-bending variety being suggested here. Does it get in the way of contributing? Or of reasoning about performance? Etc.
What can I say? Everything goes well, nothing critical has appeared soon, my coworkers are able to understand the code.
Regarding macro-based eDSLs, I would not resort to them for needs other than abstractions to be considered as a part of the language (ADTs, interfaces, etc.), since they are like languages on their own; when you try to understand/contribute to eDSL-based code, you have to understand its syntax and semantics, aside from how does it actually solve a particular problem in the problem domain.
In fact, as for now, I don't see much need in something like Metalang99 except for Datatype99 & Interface99 (and probably my feature request to libmprompt [1]).
Early on in my programming career, I implemented all sorts of beautiful magic, but I was alone. Later, I realized that beautiful magic was keeping me alone. Now, much later in my career, I program mostly like a novice. I don't write clever code, I unroll powerful one liners into boring for loops, and I use all sorts of temporary variables to make intent perfectly clear, because I want and need as much help as I can get. Confusing contributors for some wanky obsession with "terseness" or "elegance" is a good indication of either a great team, or someone that works alone.
So, I guess "simplicity in understanding".
Most of them just were not conscious of the consequences of their behavior. E.g this behavior is typical of someone that programs for a living but does not debug her own code, because someone else in the company does. So she is isolated from the consequences of her actions.
For making them understand it is as easy as making the debugger people to program and the programming people to debug for a while.
Nothing like making people miserable suffering from bad source code they have themselves created for making them understand.
Then you give them solutions for their misery and they learn pretty fast.
Experienced C programmers tend to avoid the trap of using heavy macros because they have learned. But at the same time they look for alternatives that make them more productive with some king of real metaprogramming.
For example, I think x-macros can be super useful and a consistent way to define certain data structures.
But using macros to define your own special DSL and make C code not look like C code is pretty nasty.
[1]: https://hirrolot.github.io/posts/extend-your-language-dont-a...
github.com/glouw/ctl
Check CTL/vec.h, for instance
Do you have some code to show bloated binaries with Datatype99/Interface99? The last time I checked the generated assembly code, it was nearly of the same size as hand-written code [2].
[1]: https://github.com/Hirrolot/metalang99#q-why-not-third-party... [2]: https://godbolt.org/z/ns6Ma7csd
The traditional way to do any kind of meaningful metaprogramming in C is just a printf() to a .h or .c file which then is included to your build.
There are a lot of projects doing that. Bison is supposed to be used this way, other projects are doing build configuration like that - by emitting a header with a ton of #define’s, and there’s a ton of languages which use C as a compilation target - and you can see what they are doing and get inspiration from that.
In my opinion it’s an extremely powerful model, much better than anything you can do with the preprocessor.
[1]: https://github.com/Hirrolot/metalang99#q-why-not-third-party...
Compared to that, debugging generated code is a breeze.
Also, there’s no “third-party” generators - everything just lives in your own source tree. If I ever need to go meta, it’s just a printf away; I can even commit the generated files to my VCS and be able to see what had changed in them between commits in a simple and understandable diff.
Regarding the integration, I’ll take setting up an additional build phase (once) over having to debug C macros any day.
~95% of errors from Datatype99 can be observed from the console, I hardly ever run my compiler with -E. What I mean by IDE support is that you invoke macros in the same files in which you write ordinary C code, you can't do that with printf. Imagine that you write your tagged unions (datatype(...)) inside separate files, it's clearly less convenient than embedded definitions.
> Regarding the integration, I’ll take setting up an additional build phase (once) over having to debug C macros any day.
I can't remember the time when I debugged already written and tested macros from Datatype99/Interface99, to be honest.
If you're asking about debugging code generated by macros, Interface99 has no problems with it since the generated code is trivial, pretty much as if you wrote by hand. Datatype99 can introduce a little inconvenience into it as it generates a single-step for-loop for each variable binding [2], but this should not be a big problem too.
[1]: https://hirrolot.gitbook.io/metalang99/testing-debugging-and... [2]: https://hirrolot.github.io/posts/compiling-algebraic-data-ty...
Never had such a nightmare.
For metaprogramming C you should use real macros like Lisp. We have been doing that for a long time.
Using the C preprocesor is a terrible idea. With the preprocessor you could create code, but a professional environment requires things like being able to go backwards, not just from Macro source code to executable but from the executable to source code.
What do I mean with that?
In a professional environment, when something happens, for example your program goes too slow for a customer you need to understand what is happening as fast as possible. C MACROS and preprocessor are evil for that.
The c preprocessor replaces something by something else and the process is completely opaque. If you have multiple layers of macros, codes becomes impossible to follow, isolate and understand with a debugger or a profiler.
All our code has C MACROS of any type forbidden, only permitted in external libraries. Our build process detects C Macros and stops compilation if it finds them.
Usually the way things work someone creates a easy C macro to automate some small thing, then a month later someone else creates another macro that uses the macro in a two layer system, then someone else creates another macro over the macros and leaves the old macros there.
That makes code extremely hard to understand, isolate, modularize, trace or debug. C programmers could have the temptation not to learn different tools for metaprogramming. We don't let you do that, if you want to do metaprogramming you are forced to learn the proper tools for the job.
That has made our codebase extremely robust. We use real metaprogramming with our own tools, not hacks, and that gives us a tremendous competitive advantage because problems takes 1/10th or 1/100th the time for being solved.
It's not true for Datatype99 & Interface99. Their code generation semantics are completely transparent to a user of these macros [1] [2].
> That has made our codebase extremely robust. We use real metaprogramming with our own tools, not hacks, and that gives us a tremendous competitive advantage because problems takes 1/10th or 1/100th the time for being solved.
If you use third-party tools and you're okay with that, I'm not saying you should stop using them, I'm saying that there is another solution with advantages over third-party tools [3].
If native macros haven't worked for your codebase, it doesn't mean they don't work for others. I would not say that the preprocessor is a thing to always avoid -- there are many examples why it is helpful, and even more helpful than any kind of third-party tools you can come up with.
> Usually the way things work someone creates a easy C macro to automate some small thing, then a month later someone else creates another macro that uses the macro in a two layer system, then someone else creates another macro over the macros and leaves the old macros there.
How third-party tools are different from native macros in this case?
[1]: https://github.com/Hirrolot/datatype99#semantics [2]: https://github.com/Hirrolot/interface99#semantics [3]: https://github.com/Hirrolot/metalang99#q-why-not-third-party...