Warp, a fast preprocessor for C and C++
code.facebook.com
code.facebook.com
First off, a text preprocessor is a classic filter program, and ranges-and-algorithms is a classic filter program solution. Hence, if that didn't work out well for Warp's design, that would have been a massive failure.
The classic preprocessor design, however, is to split the source text up into preprocessing tokens, process the tokens, and then reconstitute output text from the token stream.
Warp doesn't work like that. It deals with everything as text, and desperately tries to avoid tokenizing anything. The ranges used are all ranges of text in various stages of being preprocessed. A major attempt is done to minimize any state kept around, and to avoid doing memory allocation as much as possible.
Warp doesn't use much of any classic algorithms. They're all custom ones built by carefully examining the Standard description of how it should work.
What are the main stages/algorithms that constitute the processing pipeline?
I'm mainly trying to understand the high-level roadmap/design enough that I can use the source code itself to answer my more detailed questions. :)
[My first real project used Zortech C++ for OS/2, thanks for the fond memories...]
I'm credited with a few innovations (like NRVO), but for the most part I just put together novel combinations of other peoples' ideas :-)
On a related topic could you tell us about coroutines/fibres in D ... How are they implemented ? Can they be called from C ? are they are used in D's standard library, examples of a few notable use cases (I guess one would be Vibe.d), examples of asynchronous idioms in D.
Since that's a lot of questions, pointing to documents would be fine too.
..and thanks for answering questions here.
The D runtime library does have fibers:
Other than the speed processing, is there anything about how Warp operates that would make it easier to implement a distributed build system? It seems like the preprocessor can sometimes play a role in whether builds are deterministic enough to be successfully cacheable (__FILE__ macros and such).
The same should work for any other compiler, provided they don't use preprocessors with custom behavior.
I don't see any reason why Warp cannot be used in a distributed build system.
http://gcc.gnu.org/onlinedocs/cpp/Traditional-Mode.html
The other major area of difference seems to be that pre-ANSI people used foo//bar to paste tokens (whereas now we use ##). If we're talking about C, that's easy to update; apparently some Haskell folks can't do that for their own reasons. Again it's a use case which is not preprocessing of C or C++, so it seems OK to ignore it if you're implementing a C preprocessor (as opposed to a generic macro expander usable with Makefiles and Haskell).
Warp does support the obsolete gcc-style varargs, but other obsolete practices it discards.
In any case, I haven't seen any Makefiles big enough to benefit from faster preprocessing.
I'd like to learn more about this. I spend a fair amount of time building on HPC systems. Frustratingly, compiling on a $100M computer is typically 50x slower than compiling on a laptop due to the atrocious metadata performance of the shared file system. Configuration is even worse because there is typically little or no parallelism. Moving my source tree to a fast local disk barely helps so long as system and library headers continue to reside on the slow filesystem. A compiler system that transparently caches file accesses across an entire project build would save computational scientists an enormous amount of time building on these systems.
- Carmack
- Abrash
- Alexandrescu
- Kent Beck
Cannot say I'm not a little jelly of those who get to spend time with these fine gents.
Maybe Facebook might end up being the Xerox PARC of our modern times.
I'm not trying to discredit what Facebook is doing, or the people there, but it'll take time to build up to that level of talent. Having a few remarkable individuals is a great start. Having an entire department filled with them is going to be hard work.
Bell Labs made significant contributions to Physics (solid-state and optics), Chemistry, Electronics, Computer Science (operating systems, graphics, speech recognition, networking), Materials Science, Communications (here they pretty much invented an entire field of study) and the design of so many everyday things (they had an entire team that focused on just the design of your telephone wire) that is would be impossible to imagine what the world would be like if they didn't exist.
If Google didn't exist, search would still have been solved, maybe a decade or so later. If Claude Shannon didn't work at Bell Labs when he did, it could well be the case that we wouldn't have come up with such an elegant theory of Information even today, and therefore the world would look nothing as it does.
Facebook has been around very little time by those standards, and Google only slightly longer.
Time will tell if Facebook, Google and Microsoft can live up to their contemporaries.
> WB: I can guarantee that you are wrong about where your code is spending most of its time if you haven't run a profiler on it.
Definitely a good thing for all programmers to think about!
But Warp is designed to be a drop-in replacement. It doesn't produce char-by-char exactly the same output as cpp, as the whitespace differs, and the decisions about when/where to produce linemarker records are different.
But the output is functionally identical (any differences are bugs in either Warp or cpp).
Nope, it's not, for all X.
I know of a project that takes around 3 hours to build with Clang when you do a "make all".
Top-of-trunk clang from http://llvm.org/apt/ takes ~2.1 seconds on my machine to preprocess the file to /dev/null. warp (compiled with -O4 -frelease -fno-bounds-check) takes ~2.8 seconds.
So this test case is faster with clang even without precompiled headers. It is hard to make a benchmark for clang's precompiled headers because the AST is lazy-loaded from the PCH. You would have to actually have code use values from the header.
EDIT: I forgot another advantage of clang: the preprocessing and compilation and assembly are all done within the same process, eliminating process creation overhead.
http://clang.llvm.org/docs/Modules.html
When they're usable for C++, there should be little reason to use any special preprocessor, since every header file (not just a static common set) need only be compiled once into a binary format rather than being included into N source files. It can't happen any sooner for me... but as of recently, they're still very broken, so I'm still waiting.
While it seems that everybody has almost the same idea how modules should look as part of the language almost no one can agree on how they should be specified or what part of a module should actually be specified at all.
As long as there is only one compiler and no guarantee that you will not have to change everything once again as standardization is complete no one is going to touch a large cross-platform codebase and add module support.
I also think that a modern language that uses includes instead of modules is just outright insane.
Using the C preprocessor in a non-superset-of-C language would be a pretty odd choice.
example: http://hackage.haskell.org/package/system-filepath-0.4.10/do...
The same overall design could be used to implement many textual macro preprocessors, such as runoff, makefiles, etc.
I've written a couple of macro processing programs, including a make, so I know this kind of program can work with that.
Or another way of putting this is—you speak compellingly about both the algorithm side and the constant side (no tokenizing), but you don't speak about "real world" performance when interacting with many levels of caches, processes, and filesystems. Could you speak to this at all?
BTW, of course if a preprocessor is integrated with another program, Warp can't replace it. It can replace standalone use of a preprocessor.
I would ask if there are any improvements between using wrap or just using clang.
I had the same question about how it compares to clang though. I suppose now that it's open source, someone can do some testing.
Yes, I think you are correct and this is what I was getting at. It seems like a propaganda exercise to promote D, you only have to look at the end of the article to see this "And join the D language community for the D Conference 2014 on May 21-23 in Menlo Park, CA."
For the record I am very much aware of who Andrei is and the influences he has on certain languages.
clang has a pretty optimized preprocessor (It even has things like using SSE2 to do some neat things with character processing).
GCC has a fairly modern one, but it's still beatable (and clang beats it handily).
I have serious doubts. Overall compilation time is kind of irrelevant to this discussion, because Warp is just a preprocessor.
While you can get some X speedup to gcc by replacing the preprocessor, X, as a factor of overall compilation time, is usually 0.2-0.5 in most cases, depending on size of file.
I expect the gains warp gets over gcc overall from preprocessing to be similar to those clang gets over gcc overall from preprocessing.
(Though it depends on the size of files being compiled, etc).
Most companies that want actual fast overall compilation and have the resources, build caching distributed compilation infrastructure (Google, Facebook).
As mentioned, if warp is really that much faster than clang's preprocessor that it mattered, clang would be fixed :)
I would expect you guys to post the numbers comparing it to the two most used things out there :)
I had some trouble using the ubuntu 13 packages for gdc, so i downloaded it from the gdc project binaries as of the latest available there, as recommended by the readme.
Using that to compile warp with gdc with the flags it suggests (-release is not recognized by gdc, -O3 is), i get a warp that works.
For including every file in /usr/include/boost/*.hpp in one .cc file (which produces roughly 16 megabytes of C++ code), we get:
[dannyb@mainserver 12:40:56] ~ :) $ time gcc -E e.cc >f
In file included from e.cc:101:0:
/usr/include/boost/spirit.hpp:18:4: warning: #warning "This header is deprecated. Please use: boost/spirit/include/classic.hpp" [-Wcpp]
# warning "This header is deprecated. Please use: boost/spirit/include/classic.hpp"
^
gcc -E e.cc > f 3.18s user 0.25s system 97% cpu 3.528 total[dannyb@mainserver 12:40:51] ~ :) $ time clang -E e.cc >f
In file included from e.cc:101:
/usr/include/boost/spirit.hpp:18:4: warning: "This header is deprecated. Please use: boost/spirit/include/classic.hpp" [-W#warnings]
# warning "This header is deprecated. Please use: boost/spirit/include/classic.hpp"
^
1 warning generated.
clang -E e.cc > f 1.42s user 0.14s system 93% cpu 1.657 total[dannyb@mainserver 12:40:33] ~ :( $ time ./warp/fwarpdrive_gcc4_8_1 -I/usr/include -I/usr/include/c++/4.8 -I/usr/include/x86_64-linux-gnu/c++/4.8 -I/usr/include/x86_64-linux-gnu -I/usr/lib/gcc/x86_64-linux-gnu/4.8/include/ -I/usr/lib/gcc/x86_64-linux-gnu/4.8/include-fixed/ e.cc >f
cla/usr/include/boost/spirit.hpp(18) : warning: "This header is deprecated. Please use: boost/spirit/include/classic.hpp"
./warp/fwarpdrive_gcc4_8_1 -I/usr/include -I/usr/include/c++/4.8 e.cc 2.88s user 0.06s system 95% cpu 3.080 totalI've repeated these timings 10 times, and they are within 0.5% of these numbers each time.
I've also tried this on a large C++ project i have, that generates about 200 meg of preprocessed source (that i can't share, sadly) and got similar relative timings. I also tried it on some smaller projects. Based on data i have so far, clang blows warp out of the water by a factor of 2 in most cases i've tried it.
The above tests include stdout IO, but the relative numbers are the same without it:
[dannyb@mainserver 12:48:24] ~ :( $ time gcc -E e.cc -o f
In file included from e.cc:101:0:
/usr/include/boost/spirit.hpp:18:4: warning: #warning "This header is deprecated. Please use: boost/spirit/include/classic.hpp" [-Wcpp]
# warning "This header is deprecated. Please use: boost/spirit/include/classic.hpp"
^
gcc -E e.cc -o f 3.14s user 0.27s system 99% cpu 3.418 total[dannyb@mainserver 12:48:33] ~ :) $ time clang -E e.cc -o f
In file included from e.cc:101:
/usr/include/boost/spirit.hpp:18:4: warning: "This header is deprecated. Please use: boost/spirit/include/classic.hpp" [-W#warnings]
# warning "This header is deprecated. Please use: boost/spirit/include/classic.hpp"
^
1 warning generated.
clang -E e.cc -o f 1.41s user 0.13s system 94% cpu 1.631 total[dannyb@mainserver 12:48:40] ~ :) $
(I reordered this one to make the timings in the same order as they were before)
[dannyb@mainserver 12:47:38] ~ :( $ time ./warp/fwarpdrive_gcc4_8_1 -o f -I/usr/include -I/usr/include/c++/4.8 -I/usr/include/x86_64-linux-gnu/c++/4.8 -I/usr/include/x86_64-linux-gnu -I/usr/lib/gcc/x86_64-linux-gnu/4.8/include/ -I/usr/lib/gcc/x86_64-linux-gnu/4.8/include-fixed/ e.cc
/usr/include/boost/spirit.hpp(18) : warning: "This header is deprecated. Please use: boost/spirit/include/classic.hpp"
./warp/fwarpdrive_gcc4_8_1 -o f -I/usr/include -I/usr/include/c++/4.8 e.c 2.93s user 0.02s system 99% cpu 2.953 totalWarp is definitely faster than GCC, though.
There's a command (I forget at the moment what it is) that will tell cpp to list all its predefined macros. It's quite a few. You'll need to do that for clang to get an equivalent list, then drive Warp with that.
You'll be able to tell if it is taking the same path or not by using a diff on the outputs that ignores whitespace differences.
The reason Warp doesn't predefine all that stuff is because every install of gcc has a different list, and it's completely impractical to try and keep up with all that.
I'm also familiar with the innerworkings on llvm and gcc (having hacked a lot on both), and generated the list of include paths i used with warpdrive (emulating gcc 4.8.1) to be exactly the same as GCC on my system uses for 4.8.1.
I also verified the preprocessed output is "sane" in each case, as per diff.
Andrei wrote on Reddit:
"We build warp at Facebook using gdc with -fno-bound-checks -frelease -O4."
http://www.reddit.com/r/programming/comments/21m0bz/warp_a_f...
In any case, building warp with this brings the timings down to 2.31 seconds, so clang is still 40% faster (1.41 vs 2.31)
In any case, at least on my side, i don't have time to further explore, i'd love to see cases where warp is faster, but i haven't found them.
(There is also a certain irony of saying you don't post numbers because you get accused, then saying i used the wrong flags, but ...)
And, as your numbers show, suggesting the change in compiler flags was entirely justified.
In the lexer, it is used for block comment skipping. It will find the end of block comments 16 characters at a time (on both PPC and x86).
During line number computation, it will also find newlines 16 characters at a time.
This could actually (nowadays) be done 32 characters at a time on newer processors, but isn't.
You do start to hit two issues though as oyu increase the size of the skipping:
1. Alignment 2. If the average block comment/line is < 64 characters, you may lose more time performing the instruction and then counting the trailing zeros in the result to find the place it ended.
I have no numbers to back up whether this matters, of course :)
It adds AVX2 and SSE4.2 instruction support. It makes no discernible difference performance wise that i can find :)
gcc: 3.14s user 0.27s system 99% cpu 3.418 total
clang: 1.41s user 0.13s system 94% cpu 1.631 total
warp: 2.93s user 0.02s system 99% cpu 2.953 total
2.31s (with recommended build settings)1. wchar_t is ushort on Windows, uint on Linux. 2. Warp uses a slightly modified file reader from the one in the D standard library, customized to the platform.
Grep the source for "version(Windows)".
But it's actually not that tricky, since you can just change the defines it makes (after all, if you are maintaining your own toolchain, you are maintaining your own toolchain).
In fact, you are already doing it with warp to emulate GCC's defines.
Suffice to say, we've done it before to provide clang diagnostics but build with GCC.
This is really not meant as a dig (really!), but more of "why i figured facebook was not trying to make a transition". The companies trying to do so, are contributing heavily to LLVM to make that transition :)
Well, it is written in D. Walter (not Andrei) is talking about how he leveraged D's features to his advantage when writing Warp.