Speeding up C++ build times
figma.com
figma.com
The idea is the same: reduce the duplicate parsing of .h files.
I don't use any tools, just a hard-core discipline of only #include'ing .h in .cpp files.
The problem is that if you start #include'ing .h in .h, you quickly start introducing duplication that is intractable, for a human, to avoid.
On another note: C++ compiler should by default keep statistics about the chain of #include's / parsing during compilation and dump it to a file at the end and also summarize how badly you're re-parsing the same .h files during build.
That info would help people remove redundant #include's.
But of course even if they do have such options, you have to turn on some flags and they'll spam your build output instead of writing to a file.
Clang does offer something very close to this, and if you use it you'll find parsing the same/duplicate header files contributes on the order of microseconds to the overall compile time. Your article is from 1989, which is likely before compilers implemented duplicate header file elimination [1], but nowadays all C++ compilers optimize the common #ifndef/#define header guard as well as #pragma once so that they entirely ignore parsing the same file over and over again.
[1] https://gcc.gnu.org/onlinedocs/cpp/Once-Only-Headers.html
Also, given your experience in compilers - keen to see if you agree that because modern compilers optimize away re-scanning of the same header file within a compilation unit anyway (in the presence of include guards), the strategy of only #including header files within .c files is close to useless.
> the strategy of only #including header files within .c files is close to useless
It probably is. It also means the user of the .h file has to manage the .h file's dependencies, which is not the best practice. .h files should be self-contained.
struct foo { int bar; };
or typedef int xyzzy_t;
This is why you have include guards: #ifndef FOO_H_3DF0_755A
#define FOO_H_3DF0_755A
struct foo { int bar; }
#endif
GCC optimized the handling of headers with include guards 30 years ago already.If you think most of your compile time is spent in preprocessing, benchmark a clean build with the optimization set to -O0 versus your full optimization like -O2 or whatever you are using.
Both builds perform preprocessing; thus the preprocessing time is bounded by the total time spent in the -O0 build: if the actual semantic analysis of tokens and code generation at -O0 were to take next to no time at all, then we would have to attribute the time to tokenizing and preprocessing. But even then, all the additional time observed under -O2 is not tokenizing and preprocessing.
gcc -c a.c
gcc -c b.c
requires reparsing of the .h files used by both.1. phases of translation
2. constant rescanning and reparsing of .h files
3. cannot parse without doing semantic analysis
4. the preprocessor has its own tokens - so you gotta tokenize the .h file, do the preprocessing, convert it back to text, then tokenize the text again with the C/C++ compiler. This is madness. (Although with the C compiler I did manage to merge the preprocessor lexer with the compiler lexer, this made it the speed champ.)
This experience fed into D which:
1. uses modules instead of .h files. No matter how many times a module is imported, it is lexed/parsed/semanticed exactly once.
2. module semantics are independent of who/what imports them
3. no phases of translation
4. lexing and parsing is independent of semantic analysis
It used to be a common wisdom that the character-level processing of code took the most time. Just like the old floating-point is slow; always use integer when possible.
Also note that the ccache tool greatly speeds up C and C++ builds. Yet, the input to ccache is the preprocessed translation unit! When you're using ccache, none of the preprocessing is skipped. ccache hashes the preprocesed translation unit (plus compiler command line options and the path of the compiler executable and such) in an intelligent way and the checks its cache. If there is a hit, it pulls the .o out of its cache, otherwise it invokes the compiler on the preprocessed translation unit.
If most of the time were spent in preprocessing, a much more modest speedup would be observed with ccache.
Back in the Bronze Age (1990s) I endeavored to speed up compilation in a manner that you describe ccache as doing. After the .h files were taken care of, the compiler would roll out to disk the state of the compiler. (It could also do this with individual .h files.) Then, instead of doing all the .h files again, it would just memory map in the precompiled .h file.
And yes, it resulted in a dramatic improvement in compile times, as you describe.
The downside was one had to be extremely careful about compiling the .h files the same way each time. One difference could affect the path through the .h files, and invalidate the precompiled version.
It was quite a lot of careful work to make that work, and I expect ccache is also a complex piece of work.
What I learned from that is it's easier to just fix the language so none of that is necessary. C/C++ can be so fixed, the proof is ImportC, a C compiler that can use imports instead of .h files, and can compile multiple .c files in one invocation and merge them into a single .o file.
> unoptimized build is used, as that is the edit-compile-debug loop
is no longer true.
Modern C++ has a lot of metaprogramming abstractions in it and they are no cost only in optimized builds.
In my years of gamedev work I have not met a sizeable project that was working in unoptimized builds even for debug purposes. Unoptimized only worked in unit tests or small tools.
(I'm sure you've been there, though; gamedev is one of the areas I would expect people to be more sensible about their C++ feature usage in.)
Now the elephant in the room of build times that no one wants to talk about is 'the optimized' build with PGO+LTO. I think none of the projects I worked that got used to it ever did a local pipeline to do it xD. But if you ask people if they want to ship a build without it the answer is clear 'no'.
I will totally understand if authors of the linked article also do not like to talk about it. What I am trying to do here is to clear confusion about importance of it. Pretending that IWYU is more than polishing of last 5% of build times helps almost no one. YMMV ofc.
I know this was only an aside - but it took me the longest time to properly internalize that floats were fast these days. I'm still getting used to the idea that double precision isn't a preposterous extravagance. :D
Also doing semantic analysis during parsing should save time compared to having additional tree walking later (and does in my experience).
it replaces the prior pattern of including code predicated on a unique definition defined only within that same block to avoid double parsing.
/* if it isn't unique, you're going to have a bad time */
#ifndef SOME_HOPEFULLY_UNIQUE_DEFINITION
#define SOME_HOPEFULLY_UNIQUE_DEFINITION
...code...
#endif /* SOME_HOPEFULLY_UNIQUE_DEFINITION */
This, however, doesn't stop those same headers from needing to be reread and reparsed and reread and reparsed for every single cpp file in the project.If the header files are properly protected with inclusion guards, then at worst the contents are tokenized by the preprocessor, and not seen by anything else. But for decades now, compilers have been smart enough to recognize include guards and not actually open the file. So that is to say, when the compiler has seen a file of this form once:
#ifndef FOO
#define FOO
...
#endif
and is asked to #include that same file again, it knows that this file is guarded by FOO. If FOO exists, it won't even open that file.There are reasons to avoid including .h files in .h files, but you're not going to get compile time gains out of it, other than perhaps through secondary effects.
By secondary effects I mean this: when in a project you forbid .h files including other .h files, it hurts to add new dependencies into header files. When you change a header such that it needs some definition in another, you have to go into every .cpp file and add the #include. So, you find ways to reduce the dependencies to avoid doing that.
When .h files are allowed to include others, the dependencies grow rampant. You want to compile a little banana, but it depends on the gorilla, which needs the whole jungle, just to be defined. Juggling the jungle definition takes a bit of time. C++ definitions are complicated, requiring lots of processing to develop. Not as much as optimizing the banana that is the actual subject being compiled to code, but not as little as skipping header files.
Why don't include guards solve the problem of duplicate parsing of .h files within one compilation unit? I believe modern compilers can entirely optimize away even the opening and scanning of the file. And even without that, modern NVMe disks are so fast I would imagine the file opens would be negligible.
Curious to hear if anyone has data on whether duplicate header file parsing is still an actual performance issue even with modern compilers, modern SSDs, and #pragma once.
Not exactly that, but do you know clang's -ftime-trace and tools like https://github.com/aras-p/ClangBuildAnalyzer which help analyzing where time is actually spent? (In small repeated headers I don't see much of a problem, but they of course may contain not so small things ...)
In the meantime there is -ftime-report
You can immediately see the include complexity of each source file by looking at the top of the file, all required includes are there as a flat list. Otherwise a single #include may hide dozens, hundreds or even thousands of other includes hidden in a deep dependency tree.
It also nudges the programmer away from "every class in a tiny header/source pair" nonsense.
The old include guards are error prone, requiring a unique key per header. Getting that wrong can cause anything from a compile error, an obtuse linker error, or a runtime bug depending on what’s in the header.
The chances of someone copy pasting a header and breaking the include guard are significantly higher than the chances of someone deciding that we need to support an esoteric embedded target overnight.
I did come across some esoteric path-related problems with pragma once though which wouldn't occur with include guards. IIRC it was related to filesystem links.
Most .c files had ended up including both versions, but the headers had the same guards so only the first included definition would be in effect.
The project maintainer had been fixing obscure memory bugs for years until someone noticed the real issue.
Wouldn't have happened with #pragma once I suppose. With pragma once you don't have to painfully choose an identifier for the header guard, and you can't get that wrong.
Anyway false negative identity tests is not a problem like false positive identity tests are.
GNU Compiler Collection (GCC): Supported since GCC 3.4 (2004).
Clang (based on LLVM): Supported since Clang 2.9 (2009).
Intel C++ Compiler: Supported since version 8.0 (2003).
For my own personal projects, I just use ccache and precompiled headers. It's good enough for me. I don't want to have to apply "hard-core discipline" to my projects.
If build times are a wide spread problem with a language, that’s a language problem not a user problem.
Rob Pike’s rule is that header files should not include other header files; they should only document which other header files they depend on so programmers can directly include those other header files in their .c files.
https://bytes.com/topic/c/answers/217074-rob-pikes-simple-in...
I have the exact opposite rule -- it should always be OK to include a header file on its own. As a corollary, headers must include or forward-declare everything they need.
#if !defined(SOKOL_GFX_INCLUDED)
#error "Please include sokol_gfx.h before sokol_shape.h"
#endifUnfortunately, the project still remains heavy to compile because of our use of Eigen throughout the entire codebase. The analysis with Clang's "-ftime-trace" show that 75-80% of the compilation time is spent in the optimisation stage, but not really sure what to do about that.
But that feature was pulled by the chrome team with the stated justification being that since C++ guaranteed different things (iirc around variable scopes outside of a namespace) for one file vs multiple files, supporting the jumbo build option meant writing some language that was "not C++".
Unfortunate.
It's was built and is regularly maintained by a Riot Games principal C++ architect, and automatically compiles files in large "unity" chunks, distributes builds across all machines in an organization, and creates convenient Visual Studio .sln files, and XCode projects. It's also all command line driven and open source.
This is industrial strength C++ builds for very large rapidly changing code bases. It works.
You can also specify a `-DUNITY_BUILD_BATCH_SIZE` to control how many get grouped, so you can still get some parallelism. However, I think it'd be more natural to be able to specify number of batches (e.g. `nproc`) than their size.
Code bases may need some updating to work.
dmd -c a.d
dmd -c b.d
dmd a.o b.o
or do it all in one go: dmd a.d b.d
Over time, the latter became the preferred method. With it, the compiler generates one large .o file for a.d and b.d, more or less creating a "pre-linked" object file. This also means lots of inlining opportunities are present without needing linker support for intermodule inlining.https://sqlite.org/amalgamation.html
Over 100 separate source files are concatenated into a single large file of C-code named "sqlite3.c" and referred to as "the amalgamation". The amalgamation contains everything an application needs to embed SQLite.
Combining all the code for SQLite into one big file makes SQLite easier to deploy — there is just one file to keep track of. And because all code is in a single translation unit, compilers can do better inter-procedure and inlining optimization resulting in machine code that is between 5% and 10% faster.Merging everything into one big .cc file reduces the compilation job back to an O(n) task, since each header only needs to be parsed once.
Its stupid that any of this is necessary, but I suppose its easier to hack around the problem than fix the problem in the language.
dmd a.c b.c
and it will compile and link the C files together. ImportC also supports modules (to solve the .h problems).It's all quite doable.
I don't have a good name for it but it would force the compiler to ignore previous definitions. With an 'undo' pragma as well.
edit: ok, you meant that each header is included once in each translation unit.
(Technically big-O notation specifically refers to worst case performance - but thats not how most people use the notation.)
The whole C++ build model is terrible and broken. Everyone knows n^2 algorithms are bad and yet here we are.
Also everyone: "Just do the stupidest thing in the shortest amount of time possible. We'll fix it later."
Consider the following theoretically simple change:
A definition in a file may not affect headers included after it. If you want global configuration, define them at the project level, or in a header included by all files that need it.
i.e. we need to break this construct:
#define MY_CONFIG=1
#include "header_using_MY_CONFIG.h"
Thats really all we need to do to completely eliminate the nonsense that is constant re-parsing of headers and turn the build process into a task graph where each file is processed exactly one time, and each template is instantiated exactly one time, the intermediate outputs of which can be fully cached.Most real-world large projects already practice IWYU meaning they are already fully compatible with this.
There are some videos by Jonathan Blow on how this is exactly how the Jai compiler is so fast. Why must we still suffer with these outdated design decisions from 50 years ago in C++? Why can't the tech evolve?
/end rant
The real issue is just that C++ compilers are horrendously slow. They have been designed with the intention of producing fast executables rather than compiling quickly. Think of Rust, which has a high degree of structure in its compilation process and so a high degree of opimisability, yet it still suffers slow compilation due to its usage of a C++ compiler backend.
I think this really because C++ builds are fundamentally unstructured. Rather than invoking a compiler on the entire build directory and letting it handle each file, it is invoked once for every file in a way that might be non-trivial. Improving the build process almost always comes at the
Beyond that, C++ developers simply do not care about slow compilations times. If they did, they wouldn't be using C++. It's my personal theory that C++ as a language has effectively self-selected a user-base that is immensely tolerant of this kind of thing by driving off anyone who isn't.
True, this would also need to be fixed. Compilation would need to become a single process that can effectively manage concurrency and share data.
Edit: Saw after posting this was already posted in a top level comment.
I doubt we'll see Unreal Engine get any benefit from that in a long time for example. It could be so much better, working fully automatically with almost all existing code so long as you use IWYU, which is already standard for large projects where this is needed the most.
You should really post this as a separate Show HN story!
"Estimated finish by the year 5134" made me chuckle
I know that it's unpopular in a world of header-only libraries, but it really does make a dramatic difference to build-time. Especially if your code is heavy on templates.
What the article is about, though, is changing the source code so that it is intrinsically faster to compile. At some point you say "this program isn't complicated, why does it take so long to compile?" Then you start looking at unnecessary includes, transitive includes, forward declarations, excessive inlining, etc.
> After trying a few stopgap solutions—like purchasing M1 Maxs for our team—build times gradually reverted to their original pace; Ccache and remote caching weren’t enough either.
Oh, the irony
The blog post authors would do well if they got up to speed on the basics of working with C++ projects. Books such as "Large scale C++ vol1" by John Lakos already cover this and much more.
All are useful tools but they are very poor in eliminating unknown unknowns like a book would.
Iirc, I think this was how most compilers did. The downside is that transitive deps can easily explode. Thus, compiling a super small main function can takes seconds.
I did suggest a solution to just lazily parse/check symbols if they are encountered in source. Instead of when including a type, you have to parse all the transitive header files of the file that defines the type.
I feel like this kills my compile time but I'm not sure how to fix it. Precompiled headers?
Of course the downside is that every tiny code change triggers a full rebuild then, but it's quite likely that the most time is spent in the linker anyway, so maybe worth a try.
Cmake has a feature to do that automatically during the build: https://cmake.org/cmake/help/latest/prop_tgt/UNITY_BUILD.htm...
...haven't tinkered with it yet though.
And each system should only have a single 'public interface header' to keep the number of cross-system include dependencies low.
Currently mostly usable on VC++ vlatest, and clang 17, with clang 18 bringing in support for c++23 import std (VC++ already does it).
Sadly GCC is still far behind, not to mention all the other ones still catching up with C++17.
I haven’t used C++ in quite a while, but aren’t templates a big part of this issue?
https://www.godbolt.org/z/G18WGdET5
...add <string> and <algorithm> and you're at 45kloc:
https://www.godbolt.org/z/Whv73YPYh
...and those numbers have been growing steadily by a couple thousand lines in each new C++ version.
Multiply this with a few thousand source files (not atypical with the old 'clean code' rule to prefer small source files, e.g. one file per class), and that's already dozens to hundreds of million lines of code the compiler needs to process on a full rebuild, all spent on compiling <vector> over and over again.
TL;DR: the most effective way to improve build times in C++ is to split your project into few big source files instead of many small files (either manually, e.g. one big source file per 'system', or let the build system take care of it via 'unity' or 'jumbo' builds).
I also don't believe in tiny instantiation units, but when compiling real code, parsing the headers themselves is not necessarily the bottleneck.
> which are the only ones that really matter to software development.
Debatable in this age of cloud CI builds ;)
A lot of time is spent in the linker because the linker needs to parse and deduplicate any monomorphized C++ classes (like vector). This takes time proportional to the number of compiled copies of the class / function that are kicking around.
So I'd expect linking times to also decrease if you're compiling fewer, larger source files.
Linking is pretty much just I/O-bound unless you're using LTO. This is assuming you're using a modern linker like mold.
Also, many C++ projects I've seen indirectly include almost anything into anything under the hood, so a header change on one end of the project may trigger a rebuild of seemingly unrelated source files.
Such solution introduces more funny failure modes into build though. People get quite irritated when no change edit in a single file breaks build :)
https://github.com/mozilla/sccache/blob/main/docs/Distribute...
Unfortunately linking becomes the bottleneck.
The OP was seeing build times increase faster than loc. Probably someone on the team likes small .cc files.
The article emphasizes a common issue about headers.. X-Code and Visual Studio work around this to some extent with pre-compiled headers, something that can be really hard to set up in ccache. If Figma's whole team is using macs (they mention getting everybody macbooks?) then I wonder if they could just switch to X-Code and use built-in pch support. While that introduces a dependency on X-Code :( maybe their whole C++ stack will get effectively re-written in the next couple of years anyways?
The caching of compilation fails if you touch a central .h file. When you work on projects like that you start dreading a change to the .h files because development slows to a crawl as the compile times explode.
I worked on V8 and over time the .cc files got smaller and the build times got much worse. Some people felt this was neater, but if you are not in an office with 20 beefy workstations using distcc the effect is brutal.
Though I was an early adopter of STL at the time, it still hadn't enjoyed widespread use yet. Are templates now the problem with pre-compiled headers? If so, then that should be the problem we tackle.
This requires extremely bleeding-edge toolchains, though: VS 2022 17.10, Clang 18, and GCC 14.
[0] https://github.com/clangd/clangd/issues/1293 [1] https://github.com/snu-sf-class/swpp202401/issues/21
[1]: https://github.com/llvm/llvm-project/releases/tag/llvmorg-18...
[2]: https://learn.microsoft.com/en-us/visualstudio/releases/2022...
Also, modules is probably one of the biggest changes in terms of work required in build systems, compilers and tooling since at least C++11.
I didn't realize how big the impact of these little changes were (only include what you use, forward declare as much as you can, PIMPL etc..) until I worked on a large codebase in that company.
Cfront with tcc would sound nice enough to try
No, that's long build times.
More serious - I moved to CMake presets and with that came a lot of cache optimization - including parallel builds. MacOS is now almost as fast as Windows for build, and Linux/gcc not far behind. Windows C++ seems to have the lowest modern feature compatibility, followed by MacOS/Clang, with Linux/recent GCC being the most complex. A lot of the newer features seem to add a lot to the build time..
... mind I've been working with C++ only for the last few months, and C for many years before, so consider it a beginner post in a lot of ways. Still, it was interesting to explore, and I'll be continuing to explore - I haven't yet enabled ccache for instance which I suspect will improve a lot.
Chrome got even better speedups I believe by building with clang/ninja on Windows.
Bazel is where the real benefits lie by reusing other people's (or CI machine's) partial build artifacts via a centralized cache and by avoiding to run tests that are not affected by code changes.
Obviously both fille the same purpose of being a build system, though Bazel is also a build executor not just a generator. Integration would mean either adding BUILD language support to CMake or vice-versa, but you wouldn't get the particular benefits of either this way.
CLION also highlights unused includes, nothing new here. Use a good IDE. A networked ccache also does wonders if your org allows it.
Slow build otherwise stem from a combination of: a) lack of proper modules in C++ (until recently) and b) unidiomatic or just terrible code bases. To help with the latter, hide physical implementations (PIMPL for class state, forward declarations for imports), avoid OOP-style C++ above all, minimize use of templates, design sound and minimal modules. No rocket science.
This article, while not the nerdy deep dive I'd like, does touch on what happens when you try to do that. You realize that the C++ standard library is really complicated, that your existing code is really fucked up, and that libclang is too limited a tool. You end up writing a XSLT engine in hacked up python, but by a different name.
[LibTooling][1] is probably The Right Thing ("in C++", as the article says), but I never spent the time to get it working.
Somebody write a DSL for C++ inspection and transformations that uses LibTooling as a backend. I bet there are many, but none close at hand.
edit: [this][2] is close...
[1]: https://clang.llvm.org/docs/LibTooling.html
[2]: https://clang.llvm.org/docs/LibASTMatchersTutorial.html#inte...
With the right techniques, you can absolutely forward-declare basically all of a class's functionality. Then you can put it into its own translation unit.
Members and function signatures have to be declared in the header, but details about member values/initialization and function implementations can absolutely be placed in a single translation unit.
Not at all! If your private member variables use any interesting types (and why shouldn't they?) you need to include all the headers for those types too. That's exactly why you get an explosion of header inclusion.
However, consider that in C# if you add a field to a struct, which is a value type and hence no indirection, then you do need to recompile all assemblies that make use of that struct. It's no different than C++ in this regard.
The C++ committee could address this, but instead they seem to want to pretend separate compilation doesn't exist. (Why are there no official headers to forward-declare STL types, except for whatever happens to be in <iosfwd>?) Then they complain about how annoying it is to preserve ABI stability for the standard library, blaming the very concept of a stable ABI [1] [2], all while there are simple language tweaks that could make it infinitely more tractable! But now I'm ranting.
[1] https://cor3ntin.github.io/posts/abi/
[2] https://thephd.dev/binary-banshees-digital-demons-abi-c-c++-...
> You just forward-declare a struct, and declare functions that accept pointers to it: the header doesn't need to contain the struct fields, and you don't need to define any wrapper functions.
Those functions would be more verbose because they must contain an explicit `this` equivalent pointer. This would have to be repeated at every single call site. So it's not really helping.
You don't need wrapper functions for PIMPL. You can have them if you think it's worthwhile, of course.
>Technically you can do the same in C++. But in C++ to make an idiomatic API you need methods instead of free functions, and you can't declare methods on forward-declared classes. Why not?
There are good technical reasons why you can't tack member functions into the interface of a forward-declared function. There would be nowhere for that information to go, if nothing else. I think I heard a talk about adding new metaprogramming features to C++ that might address this in like C++26, but anyway it's not a significant problem to simply avoid the problem
I think you can probably make some template-based thing that would automate implementing the wrappers for you. But it would be a convoluted solution to what I consider a non-problem.
>The C++ committee could address this, but instead they seem to want to pretend separate compilation doesn't exist. (Why are there no official headers to forward-declare STL types, except for whatever happens to be in <iosfwd>?)
Most of the STL types that people need are based on templates. It does not make sense to forward-declare those. I just don't see a use case for forward-declaring much besides io stuff and maybe strings.
>Then they complain about how annoying it is to preserve ABI stability for the standard library, blaming the very concept of a stable ABI [1] [2], all while there are simple language tweaks that could make it infinitely more tractable! But now I'm ranting.
There seems to be a faction of the C++ committee that does not share the traditional commitment to backward compatibility. They have gone so far as to lobby for a rolling release language, which is guaranteed to be a disaster if implemented. I think wanting to break ABI might be a sign of that. Let's hope they use good judgement and not turn the language into an ever-shifting code rot generator.
Keep in mind, there may be ABI breakage coming from your library provider anyway, on top of what the committee wants. So it's not necessarily such a cataclysmic surprise as something you're supposed to plan around anyway. ABI stability between language standards is mostly a concern for people who link code built with different C++ standards (probably, a lot of code). It wouldn't be the end of the world if you had to recompile old code to a newer standard, but it might generate significant work.
Are there? You could have a class decorator to mark a definition as incomplete and only allow member functions, types and constant definitions:
// in Foo.h
incomplete struct Foo {
public:
Foo();
Foo(const& Foo);
void frobnicate();
};
// in foo.cc
struct Foo { // redefinition of incomplete structs is allowed
public:
Foo(){...}
Foo(const& Foo) {...}
void frobnicate(){...}
private:
void bar() {...}
int baz;
string bax;
};
edit: there are also very good reasons to fwd declare templates. You might want to add support in your interface for an std template without imposing it to all users of your header. In most companies I have worked, we had technically illegal fwd headers for standard templates.C++ already has a way to do this via inheritance, even multiple inheritance. I suppose some of the same machinery used for inheritance could be repurposed for partial class definitions but it is unnecessary.
Edit: I think I overlooked something here at first glance. Yes it might be nice to have a public partial definition of a class and a private full definition. But the technical reason you can't have this is that using a class in C++ requires knowing its memory characteristics. If that information does not come from the code, then it must come from somewhere else like a binary. Maybe the partial definition could be shorthand for "use PIMPL" but I haven't thought through all the ways it could go wrong, such as with inheritance.
>edit: there are also very good reasons to fwd declare templates. You might want to add support in your interface for an std template without imposing it to all users of your header. In most companies I have worked, we had technically illegal fwd headers for standard templates.
I have never seen illegal forward declaration headers. Not at any company I've worked at, nor in any open-source project. I don't think there is a reasonable value proposition to doing that. What kind of speedup are you expecting from that?
>You might want to add support in your interface for an std template without imposing it to all users of your header.
This sounds good in theory but in practice, most interfaces I've seen use the same handful of types or std headers, so it can't be avoided and furthermore you'd be forcing everyone to bring their own std headers every time (and probably forget why they ever included them in their own code, in the first place). You'd be talking about a lot of trouble to maybe save one simple include, and introducing a lot of potential for unused and noisy includes elsewhere.
Base classes almost work, but you either need all your functions to be virtual or you need to play nasty casts in your member functions.
re std fwds, typically the forwarding is needed when specializing traits while metaprogramming.
Ada has piqued my curiosity before but I think if it was as good as you make it sound, it might have at least 1% market share after 40 years. It doesn't. I can't justify the time investment to learn it unless I get a job that demands it.
Trying to make the best out of an improved langauge, while not touching the broken tools of the existing kingdom they were trying to build upon.
Naturally such decision goes both ways, helps gain adoption, and becomes a huge weight to carry around when backwards compatibility matters for staying relevant.
In C, idiomatic C, you can forward-declare a struct and the functions that operate on it, and you don't need any indirection at runtime. C++ has plenty of nice features, and in general I'd reach for it rather than C, but for some reason it can't do that!
Sure you can forward declare a struct and functions that operate on it, but you can't call the function or instantiate the struct without the definition. That's no different than in C++.
The purpose of PIMPL is that you can actually call the function with a complete instantiation of the struct in such a way that changes to the struct do not require a rebuild of anything that uses the struct.
It's not about just declaring things, it's about actually being able to use them.
typedef struct foo Foo;
extern Foo* foo_create();
extern const char* foo_getStringOrSomething(Foo* foo);
That's a fully opaque type, and it's reasonably efficient. The one thing you can't do store a Foo on the stack, because you don't know its internal size and layout. So it's always a heap pointer, but there's only one level of indirection.In idiomatic C++ I think you'd have something like:
class Foo {
struct impl;
std::unique_ptr<impl> _impl;
std::string getStringOrSomething();
}
If I have a pointer to a Foo, that's two pointer indirections to get to the opaque _impl. So, okay, I can store my Foo on the stack and then I'm back to one pointer. But if I want shared access from multiple places, I use shared_ptr<Foo>, and then I'm back to two indirections again.The idiomatic C++ way to avoid those indirections is to declare the implementation inline, but it make it private so it's still opaque. But then you get the exploding compile times that people are complaining about on this thread.
The C approach is a nice balance between convenience, compilation speed and runtime performance. But you can't do this in idiomatic C++! It's an OO approach, but C++'s classes and templates don't help you implement it. C++ gives you RAII, which is very nice and a big advantage over C, but in other respects it just gets in the way.
Edit to add: now I look at this, Foo::getStringOrSomething() will always be called with a pointer (Foo&) so it will always need a double-dereference to access the impl. Unless, again, you inline the definition so the compiler has enough information to flatten that indirection.
I don't see how that pImpl approach can ever be as performant as the basic C approach. Am I missing something?
class Foo {
public:
static std::unique_ptr<Foo> create();
std::string getStringOrSomething();
private:
struct FooImp;
Foo() = delete;
Foo(const Foo&) = delete;
auto& self() {
return static_cast<FooImp&>(*this);
}
};
And then you use inheritance to hide your implementation in a single translation unit/ie. a .cpp file source file. struct Foo::FooImp : Foo {
std::string something;
};
std::string Foo::getStringOrSomething() {
return self().something;
}
With this, there is a single indirection to access the object just as in the C example.But PIMPL is used when you want to preserve value semantics, things like a copy constructor, move semantics, assignment operations, RAII, etc...
It still seems like an awful lot of boilerplate just to reproduce the C approach, albeit with the addition of method call syntax and scoped destructors.
I feel like there must be an easier way to do it. Hmm, maybe I’m at risk of becoming a Go fan...!
If you want the same type safety as C, which is basically none, then you can write it as:
// foo.hpp
struct Foo {
static Foo* make();
void method();
};
// foo.cpp
struct FooImp : Foo {
std::string something;
};
Foo* Foo::make() {
return new FooImp();
}
void Foo::something() {
...
}
And yes it's rare because nowadays most C++ developers stick as much to value semantics as possible rather than reference semantics, but this approach was very common in the early 2000s, especially when writing Windows COM components.Nowadays if you want ABI stability, you'd use PIMPL. Qt is probably the biggest library that uses this approach to preserve ABI.
typedef struct foo Foo;
extern Foo* foo_create();
extern const char* foo_getStringOrSomething(Foo* foo);
The only valid way to get a Foo (or a Foo ptr) is by calling foo_create(). Inside foo_getStringOrSomething(), the pointer is definitely the correct type unless the caller has done something naughty.Of course there are a few caveats. First, the Foo could have been deleted, so you have a use-after-free. That's a biggie for sure! Likewise the caller could pass NULL, but that's easily checked at runtime. Those are part of the nature of C, but they're not "no type safety".
You can also cast an arbitrary pointer to Foo*, but that's equally possibly in C++.
The C++ approach of sticking to value semantics doesn't involve any of the issues you get working with pointers, like lifetime issues, null pointer checks, invalid casts, forgetting to properly deallocate it, for example you have a foo_create but you didn't provide the corresponding foo_delete. Do I delete it using free, which could potentially lead to memory leaks? The type system gives me no indication of how I am supposed to properly deallocate your foo.
You don't like boilerplate, fair enough it's annoying to write, but is boilerplate in the implementation worse than having to burden every user of your class by prefixing every single function name with foo_?
The C++ approach allows you to treat the class like any other built in type, so you can give it a copy and assignment operator, or give it move semantics.
So no it's not literally true that C has absolutely zero type safety. It is true that compared to the C++ approach it is incredibly error prone.
While older C++ code is rampant with pointers, references, and runtime polymorphism, best practices when writing modern C++ code is to stick to value types, abstain from exposing pointers in your APIs, prefer type-checked parametric polymorphism over object oriented programming.
If anything, the worst parts of C++, including your point about being able to perform an invalid cast, is inherited from C. C++ casts, for example, do not allow arbitrary pointers to be cast to each other.
No, they need know only the size and alignment of "class Foo". Unfortunately, in C++, either a client has to see every type definition recursively down to the primitives, or you give them an opaque pointer and hide all of the internals in the implementation (all they know about "sizeof(Foo)" is "it contains a pointer to something on the heap, probably").
edit: Ok, there's also the copy constructors and other possibly auto-generated value semantic operations, but pretend you've defined those explicitly, too.
I like to see things simple: I need something, so I include something. Compile times are really orthogonal to that, and mostly a job for compiler devs, hardware people, modules or whatever. Changing my code because of compile times seems pretty harsh
Of course there is also the relative succintness of Python and other advantages, too.
You don't want to recompile, re-parse a 12GB parquet file, etc. everytime you try a new parameter in a model.