C++ Modules: Packaging Story
blog.conan.io
blog.conan.io
But in working with c++ day to day I feel the most annoying things are indeed dealing with dependencies and headers. It can be a pain to set up a complex c++ project on a new machine. Even with Conan it can still be a pain because configuration management is a mess in c++.
As for headers I can’t wait to see them rot in hell. Some people say that they can help you reason about your Program, but this is a very small positive point. Too many times I had problems with includes which only worked because a Translation Unit included some other header beforehand. I once read a blogpost of the developer of SumatraPDF where they describe that they never include anything directly[1]. This can be an improvement for compile time. But if someone else is working on your code or has to refactor it, it can be impossible to read yourself into it or track any compile time error. If one would only use modules these errors would simply not exist.
[1]https://blog.kowalczyk.info/article/96a4706ec8e44bc4b0bafda2...
So it's not a story of "abandoning", more like an insane C idiosyncrasy that wasn't ever used anywhere else.
import x
class Foo:
def bar(self):
import y # break circular import
y.something()
Would be nice to be able to import `Foo` without pulling in also `y`, or moving `y` inline.Can be solved in different ways, but you see inline imports everywhere.
Almost every other module-based language does not have issues with circular dependencies. Python, in theory, could follow their lead, but they won't.
EDIT: I wasn't as clear as I could be. The issue again isn't modules, but scripting languages that allow top-level statements at all. Intermixing types and method declarations with executable code makes circular decencies an issue. Compiled languages don't allow top-level statements, so they don't have the dependency resolution problem.
That still happens.
It takes a lot of magic trickery to make cyclical require/imports work for JavaScript and a lot of times they don't/can't.
Realistically, that's just what happens when a language allows top-level statements, as they are executed when a file is loaded. As such, scripting languages tend to fall victim to the problem, but compiled ones don't.
I'll update my comment.
So if you just need to call that binding within, say, a function or a method call, its typically fine to do so assuming that by the time that function or method call is executed, both modules will have completed their initializations.
const foo = require("foo"); // starts out as empty object
exports.bar = function() {
return foo.fn();
}let x = require('foo').x
of course, cannot do that. It's a small difference but it does make it easier to fall into the happy path in more cases.
Lots of languages allow initialization code (e.g. Java static initializer).
But in the case of scripting languages, usually the classes and functions themselves are initialization code....i.e. everything is initialization code.
I recall having seen a circular dependency compiler error in Go.
They do not work well C++ though because it puts the implementation into the header for some reason.
ELF files based on the C type system can't provide such information on object files, and binary libraries.
In the modules world it can read it from the symbol table on the BMI, like in any other module based language.
It is simply a language design flaw of C++ that class implementation (not just templates) ended up in the headers.
Hence name mangling, the only way to add some linking type safety, while using a bare bones UNIX linker that only knows C and Assembly.
But from a developer user experience standpoint, what is the difference between having to write things twice for this reason:
// header.h
int foo();
// library.c
int foo() { return 42; }
Versus writing things twice for this reason:
// interface.cs
interface IBar { int foo(); }
// library.cs
class Bar : IBar { int foo() { return 42; } }
Granted in C it was for dumb reasons whereas in a modern language it's for better (?) reasons, but: You're still writing things twice!
Not in Oberon.
Expose the types, and the public procedures. Users only need to see the spec during compilation.
In exchange for reducing parallelism because you are not using forward declerations to break dependency chains up.
I once refactored part of my code for headers to only #include forward declarations (and few generics, of course), except where really needed (implementation)
It made compile times 2x faster, probably could get 3x if fully done.
The big difference is that those communities embrace compiler and tooling are part of the same story.
Thankfully now we have a tools working group trying to bring C++ community into the modern world of module based compiler toolchains.
Go, compiling code in parallel since Go 1.19. One is able to compile the whole toolchain from scratch faster than many C++ codebases.
Active Oberon, compiling in parallel via the Paco toolchain since 2003, a full graphical workstation OS, built in a couple of minutes.
Ada, parallel compilation available in most toolchains since 1989. Also available in GNAT.
Delphi, yet another example.
It is C++ that needs to get up to date with modern toolchains and away from hacks like unity builds.
Otherwise, I agree that LLVM and GCC are not an end all and be all, and have significant issues with compilation performance.
While I can pick other examples, I can't be bothered providing an exaustive list of every single language with compiler toolchains doing parallel compilation.
$ gcc -E hello.c | cloc --force-lang=c -
419
$ g++ -E -std=c++11 hello_iostream.cpp | cloc --force-lang=c++ -
20707
$ g++ -E -std=c++20 hello_iostream.cpp | cloc --force-lang=c++ -
30876
$ g++ -E -std=c++20 hello_std_format.cpp | cloc --force-lang=c++ -
42757
Just bumping the C++ std version can add thousands of lines of code.It's a shame they add so much header bloat, but there's a lot of functionality in there and it's hard to see a way it could have been avoided if you want type safety.
Oh, sure, you can manipulate this by inserting magic formatting objects, but that's not different from saying all functions take one argument a la Category Theory or Haskell without syntactic sugar. Or you could write a class or function that allows you to tag your object or value by changing its type to a different one so a different operator<< overload is called. In the iostreams design, you're only indirectly selecting which print function to use by manipulating either the stream or the printed object. But in actuality the types of your two arguments are not different, so either you must munge global settings on the stream or "hack" which operator<< overload is selected.
And that's the crux of the issue; what you "really want" is a stream API with three arguments: the stream, the object or value, and a pointer to a function (or function object) that takes the value and transforms it to a stream of characters. Why not just pass that function directly and explicitly? Of course that requires something that's not just a single chained infix operator; my thesis being that starting with that quirky choice leads to down a path that ends in a bad design.
Or maybe a format object that returns a string in the way described by some directive, the way printf() worked in C or format conversions work in Python? You could call it something like std::format and make it available through a standard header like <format>.
As for the header bloat, fear not, import std; for the complete C++ standard library, is faster than a plain #include <iostream>, as per VC++ team measurements.
And yet none of this heavily influences compile times. At least I have not found so on MM LoC projects I've been working so far. Not sure what the fuss is around modules, I honestly don't think it will solve C++ build times. Happy to be proved otherwise.
Given that modules are PCH in disguise, I don't believe modules are going to have a dramatically bigger, and better, effect on C++ build times. Whether or not they are going to cut the build time from 45 seconds to 30 seconds I also couldn't care less - small projects don't even have the build time problem worth solving. Mid and big projects is where it will really count and if modules are going to have a substantial effect in the range of at least 20-30% in build time cut, I will be all ears.
On the flip side, I think there are still a lot of developers who don't use language servers or use the equivalent of notepad who do rely on separate header files as a means of documentation.
That said, it's always been painful to structure inter-related headers and source files to avoid circular dependencies, and if modules truly resolves that, I'm happy to see it.
compiler error: you have to match the flags of that one dependency you're not familiar with and you didn't even compile yourself, rookie!If you link against a static library that's compiled with -D_GLIBCXX_ASSERTIONS=1 it needs to be defined in your code too. Same for the CRT flags with MSVC
Also, CPPFLAGS ensure the shared code in headers is the same, whereas with modules it feels as if you have to manually bring the compiler itself in the right state.
CPPFLAGS ensures not such thing when using binary libraries.
edit: incidentally, _GLIBCXX_ASSERTIONS is not ABI breaking.
Of course, compiling everything yourself is the only way to get LTO to work well, and you probably want that.
I'm not sure if a lot of people are using pkg-config manually but you can do it easily. If you have a pkg-config file for the library it will handle the flags for you.
I guess a build system would handle this for you also.
Even for "hello world" I've found meson to be useful, and it's only two lines of meson to get "hello world" to compile. I would prefer this to manually compiling by calling "c++ main.c++ -o hello" from the shell.
With Fortran mod files I believe it's not even backwards compatible for a single compiler: gfortran.
If it's not standardized, you end up recompiling the world, which defies the point.
Only Microsoft implements that so far. I gather EDG (which powers IntelliSense) has been at least researching support for consuming IFC files.
I'm not aware of anyone sponsoring work in Clang or GCC to add IFC support.
It would enable C++ to have an Edition concept like Rust, and be able to start removing warts and poor defaults from the language, without breaking old code, and without needing the whole world to upgrade at once (well, except the modules adoption).
Hm, build2 was able to do this back in 2021: https://build2.org/blog/build2-cxx20-modules-gcc.xhtml And it was able to do it for both named modules and header units (the mentioned CMake release can only handle named modules).
How does build2 understand module dependencies without that?
> given communication APIs for GCC just landed in its main branch a month ago
Can you elaborate on what are these "communication APIs"?
Well, the daemon-based approach won't really work with certain kinds of distributed build and caching systems. I suspect it will work for a lot of use cases though. It's also problematic if you want to parse code without a build system as such. Does every analysis tool and IDE need to add support for that GCC specific API?
But the p1689 approach is implemented the same way for all three major compilers. And it should be reasonably adoptable by anything else that might want to work with CMake and other build systems that choose to support p1689.
Ah, that. It's just a new -M output. It's still cannot handle header units though without a major extension to the preprocessor semantics, which so far only Clang managed to implement (for details, see https://developercommunity.visualstudio.com/t/scanDependenci... ; in particular notice how it was reported 1.5 years ago but is still unfixed).
> Does every analysis tool and IDE need to add support for that GCC specific API?
That besides the original point, which was that I claimed build2 supported this since 2021 to which you replied that it couldn't have.
Viable designs to support header units was still up in the air until six months ago or so. I don't see significant investment for that happening in GCC or Clang still, especially with respect to how build systems are supposed to understand how to invalidate BMIs and such. Current trajectory is that we'll consider header units a niche technology or even an unfinished idea.
Thank you.
> I still think the p1689 is a more robust implementation
I agree prescan has advantages, like being easier to integrate into existing build systems/analyzers/IDEs, but robustness is definitely not one of them. What can be more robust than the compiler asking the build system directly during compilation for the information it needs? Compared to the prescan, where the build system first scans the world with one set of command lines, digests the dependencies, and then starts invoking the compiler with another set of command lines.
Also, the mapper approach could be used to address other long-standing issues, like proper generated header support: https://wg21.link/P1842R0
Even in open source it's extremely important to use proper options (arch, flags, deps in which variant, ...). E.g. with something Debian or Redhat they decide on these options. With conan you can decide by yourself. E.g. use better hardening flags or a better libc or a more stable dependency package.
└-- lib
|-- cxx
| └-- foo.cppm ---> this is a module interface (does `export module foo`)
└-- libfoo.a
Hm, .../lib/ normally contains architecture-specific files so I wonder what was the rationale behind installing architecture-independent source code (foo.cppm is a source file) there instead of something architecture-independent like .../include/?/usr/include on the other hand only is used for C and C++ header files. Modules are not header files.
While this is definitely true, the presence in /usr/share/ of many files from library packages on my system (Debian) still suggests that architecture-independent files should not go into /usr/lib/.
> See Python files at /usr/lib/python<version>
Aren't the compiled (.pyc) files in there architecture-dependent (or could be; genuine question, I have no idea)?
> /usr/include on the other hand only is used for C and C++ header files. Modules are not header files.
Yes, but it doesn't follow they cannot be installed there if that location is the most suitable from the distribution packaging point of view.
I’d like to have the option to have a bmi section in the module file itself; that of course would mean that most people would either have to have a writable /lib &. /usr/lib etc or constrain the compiler flags available. I wonder how many people would actually be inconvenienced by the latter.
The lib is still the lib. The header files are replaced with module interfaces. It's just that those are probably going to be shipped with some parsing instructions, like required preprocessor flags, the C++ version used, and that sort of thing.
Quick, someone make a mascot!
- Conan and vcpkg are probably the closest equivalent to other "language package managers" (e.g. Rust's cargo, or NodeJS's NPM).
- Spack is typically used in HPC/Scientific Computing domain.
- Nix (and Guix) are very powerful (system) package managers; these have enough nifty features that I'd call them "next-generation".
Probably also for sufficiently complicated projects, it may be worth mentioning sophisticated build tools like Bazel (which I think manages dependencies in its own way).
There is windows support now (in addition to linux and macOS), which I think was a roadblock for many, and we have gotten some very interest from folks outside the HPC community, e.g. Replica.one, who hope to use it for a software-defined OS for embedded devices: https://www.youtube.com/watch?v=sMxNafpDhng (skip to ~12min for some discussion of Spack, Nix, Guix, and portage)
Nobody wants a package manager for C in the C world. We want a package manager for whatever language we write our code in. A lot of C code is not a single C ecosystem, but coexists with other languages, even where it is all C we often have various code generation tools that look a lot like different languages that just happen to generate C, and sometimes we want to package the into for those tools not the C output.