Is the preprocessor still needed in C++? (2017)
foonathan.net
foonathan.net
#if LIBVERSION >= 2
draw_point (2, 3, RED);
#else
set_color (RED);
draw_point (2, 3);
#endif
This is actually a case where the C preprocessor would be useful in many more languages. OCaml has cppo which is like a better cpp and is very useful for solving these sorts of problems. (https://github.com/ocaml-community/cppo)https://www.godbolt.org/z/4KK997
PS: interesting to note that Zig does the "right thing":
Andrei Alexandrescu mentions it in a talk, if constexpr doesn't really do much of anything useful because it introduces a scope.
Every single "fix" probably requires more lines of code under the hood than the entire preprocessor and in the end you have added tons of additional features to the language to fix problems that (often) don't need fixing.
The preprocessor being a simple text replacement tool is a feature, not a bug, but like every universal tool it requires some common sense to not abuse it.
Someone said once that the truly portable thing between C++ compilers is only preprocessor :)
That must be why the implementation of D's templates, which are designed to be easier to implement than C++'s is at least 8337 lines?
edit: Clang's clocks in at about 11k lines (.cpp alone), I'm too scared to find out for GCC.
This comment is short-sighted.
Take for example C++'s use of include guards to use translation units instead of modules to just compile a damn file. No one in their right mind would argue in favour of a preprocessor with #include instead of proper modules if they were to develop a new programming language.
Using #define to specify constant values is also absolutely awful.
Same with #define, no harm in adding actual constants to the language since it's a simple, straightforward and expected feature. But that's no reason getting rid of #define because that's also useful as a catch-all text replacement which doesn't work on the language-level (and that's useful in many situations).
#define ZERO 1
The cost isn't the cost of parsing. It's the cost of compiling something you've compiled before. If you change a header that is included in N translation units, you compile N translation units, even if you didn't fundamentally change the header contents in a way that would effect the final object files.
Even with a "module" system as handwaved in TFA, there's still the possibility that you (or someone else) changed a "module". C++ makes it almost impossible to decide if the change requires recompilation (without effectively doing the compilation to decide).
And certainly doesn't help with compile times.
Java tried so hard to "do the right thing" by abolishing the preprocessor, and we ended up with another preprocessor called IDE, unnecessary code patterns, and (oh my) Maven profiles for conditional compilation (among other things).
So that means you really need to name your preprocessor symbols (and any other all-caps names, because that's the convention) in ways that probably won't collide. Like MYLIB_OK.
So it starts off dumb, but then you have to start layering on convention and defensive programming immediately. And it complicates entire other features of the language, naming constants and enumerated values especially.
Also, why is std::experimental::source_location loc = std::experimental::source_location::current(); loc.line better than __LINE__? what an unreadable monster that is!
s/^/\/\// (and the reverse) work well for me. It nests.
you may use multiline string literals
(void) R"long-comment(
/* C-style comment */
std::cout << "hello" << std::endl;
)long-comment"; auto loc = std::source_location::current();
at some point, which seems fair enough to me.Fortunately you only need to write this at the utility function, not at every call site, so it's ok.
In general I find various new features in C++ suffer from verbosity, but of course without experimental this one gets better.
It's tidier for the compiler though, but the change does not seem to make it easier for the reader of the code to comprehend.
D just uses __LINE__, because it hasn't got a preprocessor so the compiler can resolve the token properly.
The implementation of that is at https://github.com/dlang/dmd/blob/v2.094.2/src/dmd/expressio...
- it has file name, line number, and char number! That already makes the number of characters more similar if that’s your metric
- it can be forwarded/passed around. It’s much harder to pass macros around
- it can easily capture the caller’s location rather than the location of the macro
#define __WHERE__ std::source_location::current()
:) #define MEMBER(C, M) { offsetof(C, M), sizeof(C::M) }
I couldn't figure out nice a way to do this without the preprocessor. The best I came up with was to use a lambda: [] (const C& c) { return std::cref(c.m); }
But these are stored in a std::map which means I have to use function pointers or accept the overhead of std::function struct M {
std::byte M::*p;
std::size_t l;
template<class C, class T>
M(T C::*e) : p{reinterpret_cast<std::byte M::*>(e)}, l{sizeof(T)} {}
};
Also there are ways (not necessarily legal) to convert a pointer to data member to an offset; see the proposal http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p090...Something akin to the https://www.python.org/dev/peps/pep-0638/
Code generators/transformers are rare only because it's so hard to actually start.
Too bad https://www.circle-lang.org never took off.
For example, let's say I have a bunch of structs:
struct GeoCoordinate {
int lat, long;
};
struct GeoArea {
std::vector<GeoCoordinate> perimeter;
};
struct Place {
std::string name;
std::string contact_number;
GeoArea area;
};
Now I need to serialize these structs into a format to be sent over the wire. Currently, I have a few choices:1. Use an off-the-shelf library like protobuf (disclaimer: I work for Google). Then I have to convert my code to a protobuf definition and rely on its code generator to perform [de]serialization. I also have to hope that my library supports all the field definitions I need.
2. Write macros to define each field in each structure. These macros perform some arcane magicks that somehow create the necessary [de]serialization functions. These macros are difficult to write and maintain (or I find a library).
3. Manually define the methods myself. This is tedious, hard to maintain, and error prone.
What if I could write some code in C++ which could read the structure and generate the appropriate serialization code? Something like (syntax hypothetical):
Serializable(Class) {
std::string serialize() {
std::string output;
for (auto member : Class.members()) { // loop unrolled at compile time
if (member.type == int) {
output.append(std::format("{:10}"), member.get())
} else if (member.type == std::string) {
...
} else if (member.type == std::vector) {
...
} else if (std::has_metaclass_v<member.type, Serializable>) {
output.append(member.get().serialize());
}
}
};
};
Then I could annotate my classes with Serializable instead.It basicaly allows you full C++ reflection, but only for classes that you mark with the UE4 macros.
"With current C++(17), most of the preprocessor use can’t be replaced easily."
"And even then: I think that proper macros, which are part of the compiler and very powerful tools for AST generation, are a useful thing to have. Something like Herb Sutter’s metaclasses, for example. However, I definitely don’t want the primitive text replacement of #define."
First, and perhaps controversially, the preprocessor is type-safe; it just isn’t the same type system that C and C++ use. The syntactic elements that make up the preprocessor language like parentheses, commas, whitespace, hash signs and alphanumeric characters have their own unique types, and can only be used in contexts where those types are expected. You’ll receive an error if your preprocessor program tries to token-paste parentheses, or end function-like macro invocations with whitespace instead of parentheses, or skip commas in macro arguments when they’re expected. It’s important that people stop thinking of the preprocessor as “the thing that turns BIG_ALL_CAPS_CONSTANTS into C code”; the preprocessor it’s its own distinct language, and its purely by coincidence and some nudging by people involved in the early days of C 50 years ago that it happens to have its language interpreter run during the C compilation process.
As far as performance goes, the implementations used by the big three compilers are horrific in terms of memory usage (reaching tens of gigabytes in larger preprocessor programs, nothing ever gets freed) and processing speed (exponential algorithms galore). Clang’s preprocessor still isn’t fully standard-compliant even today. Heck, it took until 2020 for MSVC to get the /Zc:preprocessor flag to enable correct functionality. Twenty years after the last major addition! There’s a lot to be desired with the tools we use, even taking into account the complex macro expansion rules that some faster preprocessors (see: Warp) break to trade functionality for speed. It could be argued that any language that takes that long to get correct (let alone performant) implementations built is worth replacing to get rid of that complexity alone, but it’s worth keeping in mind that what we’re working with today could be much, much better than it is.
Lastly, the crappiness of the preprocessor as a general-purpose code generation language is greatly exaggerated, mostly because it isn’t Turing-complete. Yes, there’s no such thing as direct recursion with macros. But, there is such thing as indirect recursion, where each scan applied by the preprocessor can evaluate a macro again even if it was just evaluated. So, if you can set up a chain of macros that is capable of applying some huge number or scans (2^32, 2^64, whatever), even if that number is finite, it’s enough to do any conceivable code generation task. https://github.com/rofl0r/order-pp/blob/master/doc/notes.txt is the poster child of where that idea gets you; a functional programming language built on the preprocessor that can output any sequence of preprocessing tokens, with high-level language features like closures, lexical scoping, first-class functions, arbitrary precision arithmetic, eval, call/cc, etc.
The preprocessor is still the most powerful metaprogramming and language extension tool available in C++, since it’s the only tool we have to just.. generate code. No necessary reliance on compiler optimization to translate our recursive pattern-matching sfinae’d templates and constexpr functions into the code we expect. Just plain, simple text. I think that’s beautiful, and it’s not something that’s easy to replace.
Despite the author clearly disliking the preprocessor, for justified reasons, most of the article is about how essential it still is.