Apple's Module proposal to replace headers for C-based languages [pdf]
llvm.org
llvm.org
Sure it's non-standard and no one who cares about portability will use this (yet). But this is exactly the way that good ideas get refined and eventually standardized. You surely wouldn't want to standardize a module system that hadn't been already implemented and tested -- that would just leave you with surprises when theory meets reality.
C and C++ are here to stay -- we should be open to improvements in them.
They don't explicitly mention this, but I'm sure that they have no plans to remove existing #include functionality -- it is a near certainty that someone, somewhere depends on having the preprocessor state affect how an include file is processed. There are probably even cases where you can look at the design rationale for this choice and say "yep, that really is the best solution for what you are trying to do."
Considering it's Apple-designed and Apple works mostly with C (and Obj-C), it's not really surprising that they designed it for C and working with C++ rather than design it for C++.
As for removing #include, at least in Obj-C, #include is still there to allow working with C libraries. I would assume the same for C & C++.
* Not really sure what is Apple's Obj-C versus a standardized Obj-C
Sadly, there's no such thing as "Standardized Obj-C", and Apple's version is about as close as it comes. They basically own the language.
Doug chairs the study group that's evaluating a module system for c++ ("""Sutter announces there is a Study Group for modules and Doug Gregor is the chair.""" http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2012/n338...)
Doug happens to be employed by Apple, but calling this "Apple's proposal" is somewhat misleading I think.
dgregor is clearly heavily involved and the likely the principle architect of this proposal, but he did it on Apple's behalf to solve their problems, by one of their employees, with clear input for a number of Apple's internal stakeholders. Yes, he is the chair of study group, but is participation in WG21 is funded by and at the behest of Apple.
The germ of every idea starts with one person, but many (most?) ideas require the involvement of many people to make them happen. Where credit lies is complicated, but I think it is completely accurate to refer to this as either Doug's proposal or Apple's proposal.
So, I think I have some insight into what's going on. The design and implementation of the modules support in Clang is absolutely being driven by Doug at Apple. Anyone can see that. =] However, there are some other aspects to this effort that were not the focus of an LLVM dev meeting talk. Daveed Vandevoorde has written the proposals to the committee thus far[1], and is continuing to work on the proposal and language-design side[2]. Doug is currently the one driving the implementation forward, but Clang and this implementation is completely open source, and the intent (to my knowledge thus far) is absolutely to converge with the proposed standardized feature.
[1]: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2012/n334... [2]: http://youtu.be/8SOCYQ033K8
I expect lots of others will end up contributing ideas and and implementation effort long-term, even though Doug has charted the course on the implementation side and Daveed on the proposal side thus far. I don't think this is at all likely to become a vendor-specific extension with no standards support. This is something the committee is really actively pursuing with broad interest across organizations and representatives.
As far as I can see clang is just the first compiler implementing this proposal which is a very important step for standardization, because prior implementation experience ends up influencing the process heavily.
[1] : www.open-std.org/jtc1/sc22/wg21/docs/papers/2012/n3347.pdf
The important bit is that the proposal's ideas for making the transition easier are good and make it seem like this may get traction where similar efforts have failed before. That Doug Gregor and other LLVM/Clang/LLDB developers are already working on the Clang implementation is even better. At the very least we may see this in Objective-C.
M x N -> M + N
With M being numbers of included headers and N being the number of files including that header.
I think that would be solving that mangling of stuff not working at cross purposes.
The point is to make it so you go over each file once, rather than multiplicatively many as you do now.
I'm a little wary at such a major change to C, though.
(In the selection called "Selective Import"). We can't save the world from horrible coders, but selective import would drastically drop the constant costs.
On the other hand: if someone would do the equivalent to their browser, people would call it fragmentation.
It will be interesting to see how gcc reacts to this. If this decreases compilation times significantly, I think they will have to follow suit.
I haven’t seen many people calling Dart, Pepper, Native Client or the Chrome Web Store “fragmentation”.
The moment somebody writes the first module, the world sees the first 'best compiled with clang' code. Once that code uses #import to improve compilation speed, that could become 'must be compiled with clang'. That, IMO, is similar to adding an extra tag to your browser's HTML dialect and then promoting its use.
I expect Apple is aware of the risk and wants to to prevent this (the fact that they plan a 100% compatible middle step is an indication of that), but the risk is there.
Also, people have argued that the effort spent on Native Client should be spent on improving JavaScript performance because of the fragmentation issue. Responding to another comment: people also have complained about the gcc-isms present in lots of open source libraries.
The moment somebody writes the first Dart or NaCl based website or the first Pepper plugin, the world sees yet another (not the first) "must be viewed with Chrome".
This makes sense because no widely deployed browser, not even Chrome, has the native Dart VM in it. You can download a build of Chromium with the Dart VM in it, but Chrome itself doesn't have it (though we hope it will when the time is right).
Check out api.dartlang.org. If the site works for you then your browser supports Dart just fine. :)
This happens all of the time in browsers. See: Dart, vendor prefixes, JavaScript, etc, etc.
This is also absolutely nothing new in compilers. GCC has had a bucket load of its own C extensions for years (decades), as have many other compilers from many vendors.
Vendor-specific extensions are par for course. In fact, they're a good thing! The first step in moving a standardized language forward is to have the vendors designing and adding non-standard extensions so that they can experiment with ways to "scratch the itch" they're feeling. Good extensions get taken up in committee, and if they can be made palatable to all involved, they get standardized. Bad extensions die on the table.
Having tried-and-tested features drive standardization allows hindsight and experience to strengthen resulting standards. We call the opposite, where the standards body invents a feature out of whole cloth with no example implementation having been tested in the real world, "Design by Committee". This strategy does not have a high reputation for quality and success.
• M headers with N source files -> M x N build cost
It's only MxN if there is no use of the "#ifndef _HEADER_H" workaround that he mentioned earlier. Wouldn't adding a preprocessor directive like "#include_once <header.h>" solve this? Alternatively, these guards could be added to the headers themselves without changing the preprocessor. This probably should be a parallel set of headers (#include <std/stdio.h>) to avoid breaking the rare cases that depend on multiple inclusions, but creating that set would be a simple mechanical translation. • C++ templates exacerbate the problem
I'm mostly a C programmer, so I have no argument here. • Precompiled headers are a terrible solution
Why is this? It likely would break the ability of headers to be conditional on previous #define statement, but since the proposal does this anyway it doesn't seem insurmountable. Along that lines, how does this proposal handle cases where one needs/wants conditional behavior in the header such as "#ifdef WINDOWS" or the like? And is caching headers during the same compilation also "terrible"?There still is a processing cost for evaluating the instruction
"The preprocessor notices such header files, so that if the header file appears in a subsequent #include directive and FOO is defined, then it is ignored and it doesn't preprocess or even re-open the file a second time. This is referred to as the multiple include optimization."
http://gcc.gnu.org/onlinedocs/cppinternals/Guard-Macros.html
No, that means that for foo.c a given header is only added once. It doesn't prevent bar.c from importing and recompiling that same header.
No -- the point is still that you recompile each header for each source file that uses it. Running the compiler N times (once for each source file) costs M x N, not M + N.
> Why is this?
I agree with you here, I think this dismissal of precompiled headers needs a more thorough explanation (but perhaps it was given orally in the talk).
To get any benefit from the use of precompiled headers at all, you need to include practically every single header in your project in your precompiled header. Doing this is the opposite of modular.... it tightly couples each of your N compilation units with each of the M modules (changing any one of the modules bundled in the pch will force you to recompile all compilation units using the pch.)
It also seems like this would be a win even if only used for rarely-changing system headers. As shown in the slides, for small projects the lines of unchanging standard includes dwarf the project specific code. Wouldn't it be a big win to avoid parsing and compiling all of these?
Is there a fundamental reason the compilation can't be optimized away?
For example, given a header file A.h, couldn't you precompile it and generate a bloom filter of all preprocessor tokens contained therein (except ones defined in files that A.h includes)? Then if B.h (which includes A.h) changes, see if any of the preprocessor macros that are defined at the point of inclusion are in A's bloom filter. If not, you can safely skip the recompilation of A.h. If so, you should probably change A.h anyway to directly include the definition of any macros that are affecting its compilation.
I'm not sure how it gets away with this and works as well as it does. Maybe there is a behind the scenes check that I don't know about? In any case, it might be a good framework to add the caching you describe.
You're right, I was thinking about this wrong. I was thinking that the main cost is not the direct includes (M x N), but that each header further includes other headers. (M x N^O). One-time includes prevent the exponential explosion, but you are still left with M x N. But I tend to use ccache to avoid unnecessary recompilations, so I rarely feel the brunt of this.
Obj-C already has this: #import http://en.wikipedia.org/wiki/Objective-C#.23import
It’s just #include with built-in guards, but headers still have to be compiled once per compilation unit.
It may be possible to do precompiled headers right. However, compiler vendors have been trying to do so since the mid 90s and haven't succeeded yet.
The reason I personally like this proposal better than PCH is that it requires explicit use. You can slowly convert libraries to the import syntax while still #including other headers that require conditional stuff. Seriously, just converting all #include <anything from C++ STL or Boost> to this will save 10s of gigabytes of source from being compiled on some projects I've seen.
I'm not sure what you'd do about the conditional behaviour, though - seems like something of an omission. I've found it handy, and it is used a lot by the Windows headers. (I don't recall seeing it much outside Windows, though - perhaps a case of "Unix doesn't use it, OS X doesn't use it, RISC-OS doesn't use it, AmigaOS doesn't use it, Windows DOES use it - OK, four to one, it's not important"? ;)
Think about it: the header potentially depends on any "define" value defined before the header is processed. If we'd have modules that define interfaces that don't depend on macros, we can process all the definitions that make the module only once for all the obj files that depend on it.
You typically don't need conditionals on if another module is used or not. You need conditionals to compile the whole application or the dynamic library some specific way, but that is also only once per the compilation of the whole application.
At the moment there is really code that can include more times he same header to define different types, but these are extreme use cases, however most of the libraries would gain if they would be defined as "modules" definitions of which can be compiled only once per application build.
module ClangAST {
umbrella header “AST/AST.h”
module * { }
link “-lclangAST”
}
This hardcodes an implementation-specific syntax and yet says nothing meaningful. Drop the "-l" and you're just restating the name of the module.
What value is there in baking a command-line flag into the module definition?P.S. I also find it strange that something intended to blend with C/C++ doesn't use semi-colons. This is just stylistic, though.
I'm more concerned with the developer usability benefits & drawbacks of the feature. As somebody who is a polyglot, but spends a large amount of time writing Objective-C, I have come to absolutely love header files.
I see header files almost as documentation. To me, a header file is a description of everything that's public about an API. My header files tend to be very well commented, and very sparse, only containing public methods and typedefs.
When the need arises to make internally-includable headers (say I'm writing a static library, and have methods that are private to the library, but public to other classes within the library), I will usually write a `MyApi+Internal.h` header for internal use, which doesn't ship with the library.
A developer should never have to dig into implementation files, or into documentation, in order to use a library. Its headers ought to be sufficient. Things like private instance variables or anything private does not belong in a header file.
FWIW, here's the public header for the library I spend most of my time working on:
Not if you're working on a large project. Firefox takes about 15 minutes to build from scratch on my fast Linux box. Getting that down by a factor of 10 or more would be fantastic.
If you have a look at the proposed ".map" files, the headers are still fully enumerated. I assume the IDE/debugger will still be able to find definitions in the original header files via the .map files.
Tell that to my fecking codebase.
You are free to take care of our > 1 hour compile time build.
Compiling Qt or LLVM from source will quickly put a rest to these unwarranted assertions.
I'm aware of the performance issues headers create at compile time but I still like them quite a lot. On the other hand I haven't really been too fond of any module system I've seen so far. Maybe I'm just extremely old fashioned.
Granted there is some compiler overhead for importing large header files but I don't really notice it at all.
Also, we already have an Apple/Next non-standard C extension (objective-C). I don't think we want anything else added without proper standardisation regardless of the motivation. I'd rather they forked the language.
This is confusing. Surely Objective-C, which adds a hell of a lot that C does not address and many syntax and runtime changes to support it, would fall under the definition of "fork of the language", rather as C++ does, rather than simply a "non-standard C extension" (surely that description better applies to the many GNU C extensions in GCC?).
Re: "adding things without proper standardization", the role of standards committees is to reach consensus among vendors so that they can standardize non-standard extensions that they have variously implemented and tested in the real world first. To argue for the opposite, that the vendors must do nothing until the committee hands down the One True Way from On High, untested outside of their heads, is the height of Design By Committee.
Yes we all know where vendors that do that got us:
-moz-gradient:
-ms-gradient:
gradient:
-webkit-gradient:
Oh and Microsoft with their C++ CLR extensions and middle finger to C99.This is daft. Objective-C has been implemented using a full-fledged compiler for decades. Is C++ "just an extension to C" because it was once a pre-processor on top of C? Is Common Lisp "an extension to C" because ECL transforms it into C?
Utter lunacy.
> Yes we all know where vendors that do that got us
Yes; they got us tried and tested ways of implementing things like gradients, so that the standards body had something in the real world to base a successful standard off of. The alternative gets us things like C++98's export templates: unimplementable garbage that no vendors could support because they were invented on paper and any attempt to actually build them in the real world caused more problems than they solved.
And just so we're clear what's going on here:
The linked slides are about LLVM's implementation of Daveed Vandevoorde's proposal to the C++ committee (http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2012/n334...). The C++ Standards Committee held off on adopting that proposal for C++11 because there were no implementations of it, and asked that vendors try implementing it, adapting it as necessary to make it feasible, and provide feedback so that informed decisions based on experience could be made in the final draft of C++17.
Let me repeat that: LLVM is doing exactly what the standards committee asked them to do.
It's not even a fork, it's a different language source-compatible with C. It has its own semantics, its own syntax and its own runtime.
But yeah - the standardisation issue is worrying, although at least it looks like they've thought about the issue with their transitional proposal.
Knowing your shit gets you further than language changes here.
For ref, I've been writing c since 1986 and I've seen proposals like this come and go lots and every one doesn't end up with an improvement.
I also think that Apple wants every compilation speed improvement it can find, so that it can improve syntax coloring and error highlighting in (almost) real time while you are editing.
Finally, I don't think this really is a proposal. Apple shows its intent earlier than they did with WebKit, but read that last slide: "clang implementation underway".
I expect they will listen to feedback and change this if people propose real improvements, but I think it is a given that this will be in clang soon, and also that it will be used in clang, LLVM, and WebKit. From there, chances are it will spread, either via tools that use WebKit or LLVM, or because of superior compilation speed.
How soon that 'soon' will be, I don't know. Implementation may not be as simple as it looks.
Well-defined modules, if standardized, would be a huge boost to C and C++ compilation speed. Remember the proposal does not take anything away, you can always use #include, but adds a new method for those wanting to take advantage of it. Why the resistance there?
So it looks like Apple is working towards making this a standard, not a non-standard extension.
http://dl.rust-lang.org/doc/tutorial.html#modules-and-crates
Browsed golang.org for a bit and didn't find any info on how Go handles modules/packages. All I know is that it has something to do with the folder hierarchy of your project (links appreciated).
with Ada.Strings.Unbounded;
but "withing" ada, you'd get the parent package(s). eg you'd get Ada.Strings and Ada.Strings.Unbounded but not Ada.Strings.Fixed
This particular proposal is from Apple.
When Herb Sutter/Microsoft make proposals for C++, it's phrased as "Microsoft's proposal for C++0x" or whatever.
http://channel9.msdn.com/Events/GoingNative/GoingNative-2012...
It's also my understanding that there are no "submodules" in Go. If you want what's in the import path "foo/bar", importing "foo" is unrelated. In fact, there might not be anything in "foo", making it an invalid import path, despite "foo/bar" being a valid one.
And then there's the stuff with being able to directly import code from github, code.google.com, etc, but that dovetails in to the same mechanism after downloading.
I have to think that this is a minor repartee between Apple and Google language groups. Google did something clever that breaks with tradition, and now Apple is doing something similarly clever yet backwards-compatible.
I like this dynamic. I certainly appreciate this move as a OS X/iOS programmer, since getting a functional version of Go on iOS is a pipe dream for strategy tax reasons.
http://clang.llvm.org/features.html#performance
Note that "Fast compiles and Low Memory Use" is listed more prominently than any other feature (in primary position, with more text and pretty graphics).
Java has very fast compilation as well, due in no small part to tossing the preprocessor and having fast binary imports.
Just as input, the ISO Pascal eventually got the changes that were available in Turbo Pascal and Mac Pascal (known as ISO Extended Pascal), but by then most people considered Turbo Pascal the _de facto_ standard.
- Java has a much simpler grammar to parse. - Everything is compiled at most once. - Encapsulation is stronger: private class members are not exposed in public class ABI and external code doesn't depend on them (this also extremely improves recompilation times - no need to compile half of the project because you changed a private method).
Except that this is pure marketing. Languages with modules were already compiling as fast as Go does, back in the 80's.
Modula-2, Turbo Pascal, just to name two of many.
I was just calling the attention that some Go devotees seem to think their compile times are something new.
And there are lots of people interested in other sets of major features.
So who knows if C++17 will include it, and even it does, how long one can use it in production.
Until 2017 many native developers might just had moved into Rust, D, Go or whatever comes along.
Just look how computing used to be 5 years ago.
Bjarne believes that for C++17, we only have resource for 1 major feature, 2 medium features, and a dozen minor features.
</quote>
https://www.ibm.com/developerworks/mydeveloperworks/blogs/58...
How do you get away from creating a header file for a closed source module? Without a header, how would users of your module know what they can call? Can you perform reflection on a module to inspect it? Is there some kind of tool proposed, like javadoc or pydoc, to generate documentation for a module?
How does this work with C++ templates? If you don't know in advance what types the template will be instantiated with, how can you pre-compile the code?
I'm sure the authors have thought through all these issues and more; I'd love to read about their solutions.
Creating header: the module needs only the public information for a closed source module. You could generate it from a .h file if you are a consumer of a closed source module, or if you are the producer of the module, you could potentially ship either pre-generated modules or all the information needed to generate a module.
javadoc, etc. It sounds like you could generate the information directly from the C/C++ source since there is now information about what is public and private in the source
Compiling this: template <typename Type> Type max(Type a, Type b) { return a > b ? a : b; } Generates no code, but it does some work in the compiler that can be cached. This is what would be stored in the module.
Re: templates, I thought the expensive part of template processing was code generation rather than parsing. Would caching the AST really be that much of a saving? I admit I've never tried profiling it to find out...
You've just parsed ~6MB of code. You might instantiate 2 or 3 templates from it. The parsing truly is non-trivial.
Do answer some of your questions: Instead of compiling directly to object files/libraries and distributing headers with them, you will be able to distribute modules files. Those will be preprocessed and already be transformed to some vendor specific format.
Templates require some (actually a lot) processing before they can be instantiated. This can be done without knowing any of the types used for instantiation and this can be used to speed-up compilation.
[1]: www.open-std.org/jtc1/sc22/wg21/docs/papers/2012/n3347.pdf [2]: https://github.com/boostcon/cppnow_presentations_2012/blob/m...
For example, class A's code makes use of class B. Class B's code makes use of class A.
a.h looks like:
class B;
class A {
public: void foo (B *);
}
a.cc looks like: #include "a.h"
#include "b.h"
void A::foo (B *) { b->narf (); }
and b.h and b.cc use A in the same way.I think you may be mixing this up with C's #include, which does have recursion issues, hence the need for its absurd #ifdef guards. But Apple did away with that years ago with #import, which is basically #include with built-in guards.
That they chose not to mention the issue makes me nervous that they think they can just ignore the problem. Many previous languages have ignored the problem and said "don't do that." I hope this proposal does not go in that direction.
How will this avoid the LD_LIBRARY_PATH hell of trying to depend on a local module (ie. conflicts)?
How will it work at all with local submodules inside a single project? (ie. I have 200 local c files, each with a header. Now what? A module each? How do we handle simple dependencies between classes and functions?)
How will we import actual macros?
My guess is that the answers are:
1) include path style --module-path=blah
Seems fair, but this is going to be as messy as include paths already are.
2, 3) dont use modules except at a system level.
Thats a shame as far as im concerned, but perhaps I'm wrong. Can anyone else see how these might work?
I understand some of the objections, and the import mechanism doesn't sound like a bad thing, though some of the objections are weird -
Import only imports the public API and everything else can be hidden - who wasn't doing this for libraries or large code modules anyway? Have static functions at the code module level, 'private' headers for sharing functions within a larger logical module and public headers at the logical module or libary level. Is this too cumbersome?
Most of these issues are things you learn really fast how to avoid in production systems, using things like #pragma once, include guards, and proper symbol and header exposure when writing C libraries.
This doesn't really seem like much other than feature bloat for something that works, works well enough, and which probably won't be implemented in a timely fashion by at least 1 major compiler vendor (Hey folks! There's life outside clang and gcc and icc!).
When first parsing a header, the parser can take track about all macros it depends on in the preprocessor state. E.g. a very common macro is the one it reads very first, like `#ifndef __MYFILE_H`.
Then, including a header becomes a function (<macro1, macro2, ...>) -> (parser state update, i.e. list of added stuff). This can be cached.
Thinking about my trips to /usr/include, those headers weren't that useful for coding with but you could get constants and function names at least.
I can't believe nobody else is using the same system, using "go to definition" on a library function is just as natural as using it on one of your own functions. Only difference is in the second case you see the actual code of the function instead of just the function header.
(Replace .NET DLL-files with these new "modules" to make this applicable to C)
Ps. even if you do have the source, in most editors you can collapse the code to hide everything so you don't have to "slog through all the source".
Almost every widespread modern language (all except JavaScript? [1]) has a built-in module system. Adding modules to C is not just some sort of hack to work around preprocessor limitations. Instead I'd say that the C preprocessor is a hack to work around lack of modules.
[1] and modules will probably be added to JS in the next edition of the ECMAScript standard: http://wiki.ecmascript.org/doku.php?id=harmony:modules
Moving from headers to modules is a fundamental change in how the language operates and is guaranteed to further break compatibility between compilers further. I worry that jumping to add these features to the language standard is premature and we should instead look further to optimizations within the preprocessor and linker to see if we can improve performance first.
You can't cache the AST of a #include'd file without breaking the standardized semantics of #include. ccache (which is what you're basically proposing) gets away with it by just ignoring the standard and shrugging if some legal programs break horribly when it's used. The Standards Committee doesn't have that luxury.
You can't change existing functionality while "keeping backwards compatibility". The draft Modules proposal does a much better job of backwards compatibility than what you're proposing, because it allows #include to continue to have the same semantics it has always had, and even provides a clean way forward for using both #include and modules in the same translation unit.
> Moving from headers to modules is a fundamental change in how the language operates and is guaranteed to further break compatibility between compilers further.
How would a standardized module system "break compatibility" between standards compliant compilers? The title of this story is completely misleading: this isn't "Apple's" proposal, this is about LLVM's implementation of the Standard Committee's Module Working Group's draft proposal.
> I worry that jumping to add these features to the language standard is premature
The committee was worried about that too; that's why the draft proposal was held out of C++11 so that vendors could try out implementing it and see how it fared in the real world.
> we should instead look further to optimizations within the preprocessor and linker to see if we can improve performance first.
It's not as though nobody has ever tried to improve pre-processor or linker performance; people have have been looking at those issues since the beginning of C. There is, fundamentally, no way to improve the performance of the current system without breaking semantic compatibility with existing programs.
If you're spamming #includes like that, you need to fix your #includes, not redefine the language.
> "‘import’ ignores preprocessor state within the source file"
I wonder if that would remove specific use-cases where you wouldn't want the import to ignore the state of the pre-processor within the source file?
Overall, I like it!
But as a stand-alone feature, stateless preprocessor includes could be a nice feature to have.
I actually like the preprocessor. I like that I can write code like this
#ifdef DEBUG
#define DEBUG_BLOCK(code) code
#else
#define DEBUG_BLOCK(code)
#endif
void SomeFunction(int a, float b) {
DEBUG_BLOCK({
LOG_IF_ENABLED("Called SomeFunction(%d, %f)\n", a, b);
});
... do whatever it was SomeFunction does ..
}
In other languages that I'm used to there's no way to selectively compile stuff in/out.I like that I can change the behavior of a file for a single include unit
-- foo.cc --
#include "mylib.h"
-- bar.cc --
#include "mylib.h"
-- baz.cc --
#define MYLIB_ENABLED_EXPENSIVE_DEBUGGING_STUFF 1
#include "mylib.h"
because enabling it globally would run too slowI like that I can code generate
// --command.h--
#define COMMAND_LIST \
COMMAND_OP(Stand) \
COMMAND_OP(Walk) \
COMMAND_OP(Run) \
COMMAND_OP(Hide) \
COMMAND_OP(Jump) \
// make enum for commands
#define COMMAND_OP(id) k##id,
enum CommandId {
COMMAND_LIST
kLastCommandId,
};
#undef COMMAND_OP
// --command.cc--
// Make command strings
#define COMMAND_OP(id) #id
const char* GetCommandString(CommandId id) {
static const char* command_names[] = {
#define COMMAND_OP(id) #id,
COMMAND_LIST
#undef COMMAND_OP
};
return command_names[id];
}
// make a jump table for the commands
typedef bool (*CommandFunc)(Context*);
bool FunctionDispatch(CommandID id, Context* ctx) {
static CommandFunc s_command_table[] = {
#define COMMAND_OP(id) id##Proc,
COMMAND_LIST,
#undef COMMAND_OP
}
return s_command_table[id](ctx);
};
Or this class Thing {
public:
void DoSomething();
private:
#ifdef USE_SLOW_LEGACY_FEATURE
// needs access to Thing's internals.
void EmulateOldSlowLegacyFeature();
#endif
};
Yes, I can try to hide the implementation but again, the reason I'm using C++ is because I want the optimal code. Not a double indirected pimpl. If I wanted the indirection I'd be using another language.I love C/C++ and it's quirks. I use it's quirks to make my life easier in ways other some other languages don't. Modules seems like is ignoring some of what makes C/C++ unique and trying to turn it into Java/C#
People saying the preprocessor has issues are ignoring the benefits. I miss the preprocessor in languages that don't have one because I miss those benefits.
You could say, "well, just don't use this feature then" but I can easily see once a project goes down this path, all those benefits of the preprocessor will be lost. You can't easily switch your code between module and include, especially if it's a large project like WebKit, Chrome, Linux, etc.
Leave my C++ alone! Get off my lawn!
A module is a package describing a library
• Interface of the library (API)
• Implementation of the library
That is, this will never replace the way you write applications.All it means is that there is an alternative way to depend on 3rd party libraries that improves compile time if that library supports it.
If you're writing a library, you may choose to support it.
I agree, if this was a proposal to remove #include, #define and #if, it would be doomed as a non-starter; that's simply not practical. I don't believe that it is though.
Edit: Yes, I am saying that I believe this proposal is only practical for getting rid of public header files, the sort you'd find in /include/, and will have no impact on header files inside a single code base.
extern int puts(__const char *__s);
The reason for the double underscore is so the definitions won't be affected if 's' is a macro.It seems to me that that's the genius of this proposal: You can continue to use #include and the preprocessor where you feel it best serves you, because you can freely mix the two in the same compilation unit. It's fairly obvious when you're building things like your hypothetical command.h that you're going to base them on the preprocessor from the start, so the cost of moving from modules to pre-processor based solutions is irrelevant. And note that this proposal is not about removing macros, or removing the preprocessor, so your other two examples will continue to work in either scenario.
Liking macros and the preprocessor is fine and dandy, but there's no need to force template-heavy code to continue to pay insanely high re-re-re-re-re-parsing costs ad infinitum simply to retain those features.
The latest version of D is compelling. It's mostly a reasonable cleanup of C++. It retains all of the power, but the design learns from the decades of experience. The problem is that it's mostly a reasonable cleanup of C++ - it's not different enough to stand out, I think.
i feel that the preprocessor ultimately ends up with the same amount of work, just an extra pass for each included header to build a version to be cached… not to mention the complexity required to handle the multiplicity of pre-processor states required for this. maybe i am being dim and missing the obvious.
tbh, i'd rather they made their compiler work properly, like respecting alignment on copies with optimisation turned on, or implementing the full C++ 11, before adding language features to fix problems that nobody really has.
For future reference: In large C++ projects, it's not at all unusual for greater than 90% of the compile time to be spent parsing (and re-parsing, and re-parsing, and re-parsing, ad nauseam) header files.
> what about macros in include files?
The AST of the header is persisted. If the parser can parse macro definitions, macros will continue to work normally. The only case where you would need to fall back to #include is those rare times in which you want defining something in the source file to alter the parsing of the header (and even in most of those, you should just define the constant in the call to the compiler eg) "clang foo.c -DWITH_FEATURE_X")
> i feel that the preprocessor ultimately ends up with the same amount of work, just an extra pass for each included header to build a version to be cached… not to mention the complexity required to handle the multiplicity of pre-processor states required for this. maybe i am being dim and missing the obvious.
The pre-processor merely slurps the text of the included file into the including file; the compiler then parses the entire gigantic soup of <text of source file> + <text of all files included by source file>. The semantics of this require that every included header be re-parsed once per compilation unit. Imagine you have 5 .cpp files, each containing 200 characters, and each #including iostreams (which weighs in at roughly 1 million characters). Each .cpp file, post-preprocessor phase, will be 1 million, 200 characters long. A full compilation of the project will require the parsing of 5 million, 1 thousand characters. Any subsequent full build will require parsing the full 5 million, 1 thousand characters. Changing one .cpp file will result in the need to parse 1 million, 200 characters.
In this proposal, by contrast, an included file need only ever be parsed once; its AST can then be persisted and referenced eternally. In our above example, the iostreams header will be parsed once, and each .cpp file will be parsed once. This means a full build, the very first time iostreams is ever referenced in any compilation on the system, will require parsing 1 million, 1 thousand characters. Any subsequent full build will require parsing merely 1 thousand characters. Changing one .cpp file will require parsing merely 200 characters.
not to mention that this is not a problem if you encapsulate your use of standard libraries properly... maybe 10-20 compilation units have to use it if you like to split your stuff into files a lot.
standard headers are poorly written/designed by including so much crap everywhere. why can't i have specific - per function headers which include minimal stuff?
fix the headers, not the preprocessor.
Well, then you're pretty much 100% wrong in most C++ projects.
I honestly don't know what to tell you here. That persisting header ASTs between translation units is faster than re-parsing should be trivially obvious, and if it isn't trivially obvious, then the mere fact that precompiled headers and ccache dramatically speeds up builds ought to make it empirically obvious.
The facts just aren't on your side.
> namely that whatever preprocessed import module thing is created, it still has to be included into the compilation unit somehow
Well, yes, obviously. In the current model the compiler slurps the header into the source file, and parse the entire combination, resulting in the parse tree of the header + the parse tree of the rest of the file. In the proposed model the compiler pulls the parse tree of the header out of cache and just builds the parse tree of the file. Since in C++ header parse trees are often quite expensive to build (since template declarations have to live in the headers and their parse trees are incredibly expensive to build), this ought to be a blindingly obvious win.