Killing the C++ Modules Technical Specification
izzys.casa
izzys.casa
...
> "We'll make the build systems understand".
This seems to be a critical point that I agree with: you can't divorce these two, their fates must be bound to one another. Hopefully the build systems can be easily modified to accommodate.
Rust community folks might be wrong about c++ but they're right about Rust. Rust had the luxury of not being bound to any legacy specifications and it wasn't unreasonable for them to write their own build+dependency/package management tool. The net effect of them having done that means that independent of all the other great things Rust brings, much of the tedium of developing with C/C++ is gone.
Make no mistake, C++ standard committee: the bar is high and Rust has set the bar.
Modules were introduced in the 70's as concept, Mesa being one of the first languages to use them, followed by CLU and ML....
The issue with lack of modules in C++, was that Bjarne had the compromise C++ should fit into C toolchains used by AT&T, and later all C shops.
EDIT: I get Rust is quite cool, but maybe a bit of programming language research is also welcomed.
All used in the context of systems programming.
Rust modules don't add anything new to them.
Those ideas are also present in Maven, Gradle, NuGet. All older than cargo.
And all trace back to CPAN as general concept.
I am still waiting for cargo to support binary libraries.
Which in my mind matters more than all this pedantry.
I'm not sure if 'shares' is a typo or not, but it fits great with my impression of why companies choose the languages they do sometimes: everyone else uses it (it's the local culture/custom), and nothing/no one has forced them to change (by lowering their profits/share value).
There are YouTube videos....
> A definition like this, where the choice of an implementation is dependent on dynamics, is entirely natural in object-oriented languages. Yet, it is not expressible with ordinary ML modules.
Here, ML modules are compared to classes (not packages or modules has defined in Java, for example). This makes me wonder how often people talk past each other without noticing, when they discuss modules.
https://en.wikipedia.org/wiki/Modular_programming
There are plenty CS books which offer a similar definition to Wikipedia.
We might have consensus about the definition, but it does not seem to be a useful precise one. The C++ working group seems to have a similar problem. Things like "compilation unit" have a precise meaning, but they cannot find consensus if it should relate to their "module" or not.
I think the issue is more with some C++ devs that never used other languages, trying to grasp how modules fit into translation unit + PCH model they know.
Nothing in Rust is original from PL theory perspective. It's deliberately not a research language, but a practical one. Rust set the bar here merely by showing up and delivering a decent implementation (in a language that you can actually use today).
Everything in terms of safety and modules was present in languages like Ada and Modula-3.
Google improved that implementation for C++ and has a private implementation of it in use.
Gabriel dos Reis picked up some ideas from a former work he did together with Bjarne for MSVC++.
The ongoing C++ Modules TS is to merge the ideas from both sides into what will become C++ Modules.
The history of modules you provided is helpful though, thanks.
I don't think Rust deserves much praise for its module system. It is widely regarded as being extremely confusing even for pros. There are efforts underway to make it saner:
https://withoutboats.github.io/blog/rust/2017/01/04/the-rust...
No, it's mostly considered to be unintuitive, not confusing for people who know it. You're unlikely to "just get" how it works, but once you learn it it's perfectly reasonable.
Edit: Note that the RFCs don't actually change how the module system works, just the syntax
Anything else is mission creep at this point and risk delaying the effort further. Let's fix a problem at a time.
I would say for the C++ standard committee the bar is high and D has already set the bar.
My perspective as a C++ developer has been that modules have one primary purpose, which is to improve build times by eliminating the need to parse huge amounts of C++ in recursively included header files when compiling a compilation unit. If that's all a modules proposal accomplishes I'm fine with that. A modules proposal that doesn't accomplish that is useless.
Build tools generally already parse .cpp files for includes to figure out header file dependencies for a .cpp file and have to combine the filenames extracted with information from the build system on include search paths to turn those into dependencies on specific concrete files. That involves information that lives in the build system and not in the code. Modules will require something similar from build tools but as with parsing includes the rules about imports (must be before actual code, must be alone on a line) parsing out the imports is trivial and cheap compared to recursively parsing the full code in header dependencies.
So I guess I don't quite understand the problem. Modules are an improvement over headers for build times but they are not going to do everything that module systems perhaps do in some other languages because they have to be adopted incrementally into the many different existing build systems used in C++ code bases (which live outside the standard). Perhaps my ignorance of Rust means I'm missing what the author wants in C++ that's different from the current proposal.
I don't think he trys that and i don't think it's true. The committee doesn't want to that, but I don't see why it would be possible by extending the language.
Introducte a module - keyword, that works pretty much like a namespace keyword. Force the file name and path to match the name a la java. Introduce a mechanism a la "using namespace" that does import the namespace (like #include header.h) of the module and tells the build system what to build. Neither needs to become the other to do this if we forbid preprocessor commands and if constexprs around that. Allow attributes to specify visibility.
Of course it's going to be a way more complicated, but i can't see how it would be impossible.
I think c++ devs just don't want that.
The compiler is usually passed include paths to resolve includes when invoked. The build tool need to track this just to enable minimal incremental builds.
This design replaced an earlier one where the user had to manually run ‘make depend’, which would pre-scan for dependencies using a separate tool, makedepend(1). First ‘gcc -M’ was added as a straight replacement of makedepend, then -MD and variants were added to write dependency files while compiling each object file rather than requiring a separate invocation (which is slower).
But it sounds like modules will complicate things, since it could be mandatory to gather dependency info upfront in order to compile modules before files that depend on them...
You will need a similar mode were the compiler does only minimal parsing of the source file and generate a dep file containing the list of modules defined in this file and the list of modules imported by this file.
The build system can then collet all these dep files and reconstruct the full dependency graph. Given the right format, you could probably feed them directly to make.
extern template class std::basic_string<char>;
extern template class std::vector<int>;
// etc...
and a corresponding .cc file like this: template class std::basic_string<char>;
template class std::vector<int>;
// etc...
The template instantiations you list will be compiled in that .cc's object file. Anyone who includes the header will assume those templates were instantiated elsewhere. Net effect is the listed templates are only compiled once.(Of course this doesn't alleviate the problem of parsing the template header files over and over.)
See the section on explicit instantiation here: http://en.cppreference.com/w/cpp/language/class_template
find build/ -name \*.o -exec nm -C \{} \; | egrep ' W .*<' | sort | uniq -c | sort -n
is what I use. This lists all templated, implicitly instantiated, non-inlined symbols in increasing order of number of uses. Note that templates internal to libstdc++ will also show up and you may want to instantiate these too.Note that explicit instantiation won't decrease the size of your object files, only the compilation time. (By default, the linker throws away duplicate instantiations.)
(Whether inlining of non-trivial methods is desirable is - and in fact recently was - a discussion for another HN story.)
And you actually want some tools to do a version of this. In particular, IDEs and smarter text editors. And, at scale, maybe even automated build systems. Though all these (not compilers and not build scripts) tools will likely want to do things incrementally, if it's possible and beneficial enough.
The issue is, "how do you find these dependent modules?". As of right now, there is no way to find a module. It is not tied to a location in a subdirectory, nor is the name of a module tied to the file itself. I'm fine if we can get the dependency information, but how can we do that if there is no guarantee for finding the actual dependencies? Leaving these decisions unspecified by the C++ standard is dangerous. Having an important thing such as "how does the name of a module map to the name of the translation unit it contains?" be none of "undefined", "ill-formed", or "implementation defined" is asking for trouble. And to be quite honest, I don't see how running the compiler in a minimal parsing mode (once for each module), followed by running the compiler a second time for its full IFC (or whatever file format is used) and possible object file generation can be a good idea.
The standard currently imposes very little requirements on the actual build process, headers and source files do not even need to exist as actual files on a system, so it would be a lot of work to actually introduce these concepts in the standard and risk delaying modules further.
That doesn't mean that there might not valuable to standardize that, but it is another battle and many (I, for example) will argue that a strict mapping from file names to module names is wrong.
What percentage of the Rust community is a professional C++ developer, either now or recently in the past? I was under the impression that "nearly all" was a fair estimate.
If that's true, does that mean that the C++ community itself is "typically wrong" about their language?
There is no "right" or "wrong" C++, there's also no single "C++ community". I'd say that C++ isn't even a programming language, it's more like a meta-language to build your own language.
Whether that's good or bad is up for discussion, but (a) it gives a lot of freedom, (b) it invites tinkering, (c) it wastes a massive amount of time when trying to communicate with C++ coders from other confessions, and (c) it makes it hard to integrate C++ libraries into C++ projects.
edit: confession => 'denomination' seems to be the correct English term
All (good) programming languages are like that. In general, any API, any library that you develop is a language. I agree that C++ provides powerful tools of abstraction that make these languages easier to use (compared to what can be accomplished in C or Go, for example).
I think the only other language with a similar feel is Scala. It's the programming language equivalent to Magic The Gathering: half the fun is just in the mechanics of it all
I think complexity of C++ is merely a (partial) reflection of complexity of programming. For example, move semantics that you have mentioned is not specific to C++, and I find it nice when a language offers an explicit formalism for it.
Well, where they're right is in having a modest and non-condescending attitude, which is something that C++ pundits should consider adopting.
There are C++ experts who display modesty, just as there are members of the Rust community who display modesty. There is no need to try to retaliate against one person's words by hurling generalized insults against an entire group.
But we have both, and most of the folks doing comparisons with C++ do tend to be actual C++ devs (because if you're coming from Python or Ruby and haven't done C++ you really ... can't? compare with C++?). At least within the community folks pretty strongly care about misrepresenting other languages so if stuff is inaccurate it gets called out; and for most "rust vs foo" blog posts I've only rarely seen glaring mistakes a few times. Small mistakes are common but those are common to any kind of post.
I took the comment in the post as a bit of hyperbole and probably referring to subjective bits that Rust+C++ & only-C++ programmers tend to disagree on.
In general most of the vocal Rust folks who talk about C++ talk about C++ because they have lived it.
shrug
I still chuckle when randos make arguments like "C++ can't be parsed with an LALR parser", or "virtual functions are more expensive than function pointers", as if the former is the end of the world, and the latter is a checkmate that will suddenly make my knees week, arms heavy. There's code in my editor, and it's just spaghetti.
I'm not an expert on parsers but I don't think LALR parser can evaluate constexpr C++ functions which is required for parsing C++.
E.g. in the following code, expression `A<f()>::a * u` will parse either as a variable declaration (int* u) or a multiplication (5 * 5) depending on the value returned by constexpr function f:
template<bool b> class A {};
template<> struct A<true> { typedef int a; };
template<> struct A<false> { static const int a = 5; };
constexpr bool f() { return true; }
const int u = 5;
int main() { A<f()>::a * u; }
When people talk about "C++", they may mean:
- "C++ as it theoretically could be in when C++20 or so is used pervasively"
- "C++ as it can be when certain C++17 features are used pervasively"
- "C++ as it is often used in practice today"
(and this is still oversimplifying more nuanced realities).
It's very common for people to make observations about "C++" which are wrong on one sense and right in another.
Coming from the C++ world and being fairly knowledgeable about the language (not a full language lawyer but someone people usually turn to) and learning Rust, I find it rare for there to be bad comments from the Rust folks. I don't understand where this is coming from.
Now for my own judgemental comment. I see a trend of people in /r/cpp and other places ignoring the differences (like safety) or misrepresent Rust.
To be clear, I fully accept people having differences in priority (how much safety needs to be compiler-enforced). Its people who don't recognize the difference that I'm bothered by.
Rust is also foreign enough that a trivial glance can be misleading like assuming exceptions are the equivalent of panics because someone got caught up on them having a similar implementation (stack unwinding by default) when in reality, Result is more what people should look at (transport information, ?-operator for unwinding, etc).
[1] https://www.reddit.com/r/cpp/comments/75eqal/millennials_are...
I feel like the only way to have truly useful, simple and intuitive modules in C++ would involve breaking a certain amount of backward compatibility and force a few constraints (project layout, file naming scheme, build system, ...) on the user. But clearly that's not really in C++'s DNA and you risk ending up with something that's not quite C++ while at the same time not having the simplicity and elegance of modern languages that have been designed from scratch without all the baggage C++ carries.
C++ won't go away any time soon, that's for sure. But when I read articles like TFA or for instance https://bitbashing.io/std-visit.html my gut reaction is always that maybe if you get frustrated with C++'s extreme complexity, historical baggage and user-unfriendliness you should consider moving to something else instead of trying to bolt on even more exotic features on a language which is already crumbling under the weight of all the features and programing styles it already supports. I know I did.
As an example I stumbled upon this std::visit link by perusing the top posts of reddit's /c/cpp (linked by TFA), here's the discussion: https://www.reddit.com/r/cpp/comments/703k9k/stdvisit_is_eve...
There's an interesting thread about how:
variant<string, int, bool> mySetting = "Hello!";
Doesn't do what one might thing it should do. The variant ends up holding a `true` boolean instead of a string. Why?>char const* to bool is a standard conversion, but to std::string is a user-defined conversion. Standard conversion wins.
If you know C++'s history and its C heritage it makes sense but it sure as hell doesn't make me want to use C++ for my next project.
There is already a std::string liberal. Better:
MyVariant mySetting = "Hello!"s;
But std::string_view more closely models a string literal. std::string is more of a string buffer. So even better: MyVariant mySetting = "Hello!"sv;The fact that pointers automatically coerce into bools makes some sense in C but is an aberration for modern C++. It's just an example of a legacy "feature" pointing its ugly head to sabotage a modern C++ construct.
If you look at C++ guidelines out there a lot of it is about what par of C++ not to use. Don't use raw pointers, don't use raw arrays, use exceptions, don't use exceptions, maybe use exceptions but only in some cases...
And a lot of the time the most intuitive notation, the one that goes all the way back to C, is also the one you don't want to use. Don't use "Hello, world!", use "Hello, world!"sv. What does the sv do exactly? Uh, we'll talk about that in chapter 27 when we talk about user-defined literals. This is right between the chapter about suffix return types and the one about variadic templates.
Haskell? Python 2?
(Also, C/C++ string literals are perfectly safe to use — they're guaranteed-null-terminated immutable arrays that implicitly convert to `string` and `string_view`; it's more that its safer in modern C++ for function parameters to be `string_view`/`string`/`const char(&)[N]` instead of `const char*`).
{-# LANGUAGE OverloadedStrings #-}
plus lots of other ceremony.(There's only one form of string literal in C++ too; the "s" and "sv" suffixes shown above are operators, not part of the literal).
This is a rather odd complaint; I can't think of any programming language which has first-class support for nullable types, and where they aren't truthy/falsy, at least in conditional contexts. It's a pretty straightforward boilerplate-reducing idiom.
Did you mean to say builtin arrays coercing to pointers? (I'd agree that's probably the biggest problem with C and, by extension, C++).
Or did you mean `NULL` being an `intptr_t` in C++? That's the very rare C++ misfeature that doesn't come from C (where it's a `void*`), but at least C++11 `nullptr` fixes that.
String foo = null;
if (foo) { // error: incompatible types: String cannot be converted to boolean
System.out.println("was true");
} else {
System.out.println("was false");
}
If one changes `if (foo)` to `if (foo != null)`, then the code will compile.It was his solution for not having to touch such a low level language, after being forced to exchange Simula for BCPL.
So yeah C++ is plagued with C compatibility, but it does provide enough tooling for anyone that cares about strong typing and safety.
Better alternatives are needed, but they will only get wide scale adoption like Apple is doing with Swift, pushing full speed ahead regardless of what devs think.
Hence why for me, even if it isn't ideal, Java / .NET languages + C++ for low level stuff is already kind of sweet spot.
I am not deeply familiar with the Modules TS, but couldn't each environment solve the module name <-> translation unit mapping problem in its own way? For systems that can't keep reliable track of changes to the underlying representation (i.e., files) that would mean a fairly strict naming scheme.
The mapping doesn't have to be maintained by the compiler. C/C++ already depend on a suite of more-or-less independent tools to produce runnable software. I don't see how this problem conceptually is much worse than what a linker has to do.
I'm asking the same question to myself. There's a bit in the article about it being slower than the status quo, but I'd like to see actual benchmarks before making pronouncements. If that hypothesis is true, then surely the TS will fail before it makes it into C++20 proper.
#include just results in a tremendous amount of code in each and every source file. Even looking at a 6 line program, we see that after the preprocessor runs, the compiler is given 18,162 lines to process.
$ cat > main.cpp <<EOF
#include <iostream>
int main() {
std::cout << "Hello World!" << std::endl;
return 0;
}
EOF
$ gcc -E main.cpp | wc -l
18162
That is still nearly instantaneous, but it only gets worse from there and it happens to every source file in the project. I just took at look at a 1,848 line source file I'm working with, and it expands out to 147,054 lines after preprocessing. The current state of things is awful and modules are by far the most important feature for the future of C++. We need module support in our build tools and we need them to be fast.Should the TS be delayed to see if the Clang idea is better? I'm curious to know what they're doing, but that's a lot to ask for. The TS is effectively a beta, and pushing back its start means less time in beta before the C++20 standard is finalized. There's a lot of work that's ultimately going to be built on top of this so getting it right is vital, but the clock is ticking.
"Deploying C++ modules to 100s of millions of lines of code"
https://www.youtube.com/watch?v=dHFNpBfemDI
"There and Back Again: An Incremental C++ Modules Design"
1) In the source file
2) In the makefile
3) In both
It seems to me that the only sane option is (1) with a project-wide or system-wide set of "roots". Why is anyone in favor of (2) or (3)?
In general, modules should not import their dependencies, but have them provided to them by their common parent. See Newspeak's module system as an example.
Alternatively, infer the interface and type check that in the composing module.
As the other commenter wrote, the answer is that modules can never "import" concrete modules, only interfaces. And that means some extra effort, though having type-inference assist with the construction of that module should help. Definitely something that needs figuring out.
IMHO, explicit composition is something we really need to figure out in general, and it should solve the weird mix of linguistic and extra-linguistic mechanisms we have now.
I am dreaming of a programming system that takes advantage of the modern developer workstations which are very powerful and thus can support more sophisticated design and build tools than a text editor and a make. (IDEs provide only an incremental improvement on top of these.)
The only counterexample that I can think of is LabVIEW, which is pretty successful with (or rather, despite) its visual editing. (Its main selling point is device compatibility, AFAIK.)
So if you have a design for a visual (or, more generally, "beyond-text") programming system in your head that's any good, please let it out. I'd like to see it. (And until then, I'm sticking to vim.)
That said, nothing prevents anyone from having (admittedly rather vague) ideas about using a computer as something more than a glorified typewriter.
I know, it's hard - people could not imagine a car looking different from a horseless carriage back in the day...
(Speaking of the "dream", it is something along the lines of modern environments for Smalltalk. When you model a system, you do not necessarily want to be repeatedly going through the rigid edit/compile/debug cycle; the environment could in theory give you more "immediate" experience modeling the system you are trying to develop.)
If X needs one of (Y1, Y2, ...) to run - I could understand some kind of external configuration.
It's especially useful for frameworks and game engines. Kinda strange to edit sourcecode of tomcat or unreal engine to configure it to use particular regexp implementation.
So, currently we have four ways to resolve dependencies:
1. Specific at compile time (import in source file via compiler)
2. Unspecific at compile time (via build system)
3. Specific at run time (shared library via OS)
4. Unspecific at compile time (via plugin system)
You’d be crazy to create a final release of something without starting from a “clean” state, which means they DO NOT need a module implementation with 100% accuracy in all of C++’s asinine corner cases. I wouldn’t trust that accuracy even if they claimed it, I’d “clean” anyway.
This means that compilation speedup just has to ensure that MOST incremental rebuilds perform better without adding insane development overhead (such as having to respecify things).
If 60% of the time my incremental rebuilds are faster, I’d say that is more than enough to justify a C++1x release. They need to just move forward.
The drawback of OCaml modules is that all dependencies must form an acyclic graph, which can be a pain sometimes although there are well understood workarounds.
Now I didn't say that OCaml has a perfect module system, although some of the first-class module features are stunning, but my point is that it solves a lot of the problems that this C++ guy is complaining about.
Can you elaborate on that?
It's hard to imagine a sane build system where the dependency graph contains cycles.
[1] Of course you can parametrize one of the modules with a signature declared elsewhere and the other module implement that signature, but that's really just regular old dependency-breaking, so I'm not sure that counts.
For classic C++ they have #import directive (1) Recently, for windows store/UWP platforms, C++ can consume windows runtime components (2)
Two things that allow these features to work are standardized ABI, and standardized type info format.
I don’t think high-level modules are possible in C++ before ABI and type info are both standardized in a compiler-neutral way. Better yet, in a language neutral way, like MS did that: you can implement a COM object in a script language as a WSC file, and consume it from C++ using #import.
[1] https://msdn.microsoft.com/en-us/library/8etzzkb6.aspx
[2] https://docs.microsoft.com/en-us/cpp/windows/how-to-activate...
Isn't this is what cmake does? IIRC it already parses the headers.
But the point about "how does it work now?" is fair. I'm confused why a module search path is any worse than an include search path.
The idea is that when you change a module, you might need to recompile all the modules that include it, since its interface might have changed.
What I don't get in this criticism is that you have this exact problem right now with headers, and it has been solved by the preprocessor outputting dependency information (based on #include directives) for the build system to read. Why can't the same be done with modules (of course it might not be the preprocessor that does it, but the compiler or a specific tool)?
So the problem could be solved by saying 'import foobar;' maps to a file (on a search path, maybe) named exactly "foobar". But then we need to work out how subdirectories and module names map to each other. How do you import "foo/bar.ixx"? "foo.bar"? "<foo/bar>"?
Though all that seems technically solvable to me. Maybe there's just difficulty in designing this in the context of a committee.
Anyway, Java dealt with a bunch of this by enforcing that the file naming and directory hierarchy would match the class hierarchy.
When using cmake to generate VS projects it probably does something equivalent, unless msbuild has dependency tracking built in out of the box, I don't know msbuild very well.
What happens when two .cpp files want to do the equivalent of including each other's header files?
nitpicking.
equating innocent, if terse, comments with threats.
Why am I reading what this person has to say?
(I am joking, but the complexity of the language may, in fact, reflect in the relative complexity of the APIs, say.)