Include-what-you-use: A tool to analyze includes in C and C++ source files
include-what-you-use.org
include-what-you-use.org
That said, I do highly recommend its use.
I've also recently been making libraries I write compatible with users that run IWYU by annotating all public headers with IWYU pragma comments that export symbols/transitive includes correctly, etc.
[1]: https://github.com/include-what-you-use/include-what-you-use...
[2]: https://github.com/include-what-you-use/include-what-you-use...
Sometimes it gets some things wrong, so you have these escape hatches to control it: https://github.com/include-what-you-use/include-what-you-use...
Don't blindly apply its suggestions - test them, skim to see what it got wrong, sprinkle some "// IWYU pragma: keep" to help it out in corner cases. The tool is more like a linter, you don't follow everything that your linter tells you to, no?
FWIW: when I run this tool my experience tends to be that it adds more includes than it removes, because I guess I rely too much on transitive includes.
But the specific problem of include-what-you-use will still be encountered if you include directly from C libraries like system headers or library dependencies.
Do you have evidence that the situation has changed? Last I checked it still remains the case that modules inhibit parallelism and hence result in slower builds in most practical work loads. But of course if you have evidence the contrary I'd be happy to see it.
yeah, but the same graph never shows modules being faster... it only ever shows them being the same or slower. If I'm going to put in all that work, the result should be *faster*
What significantly improves compile-times is Pre-Compiled Headers (PCH), which most compilers have supported for decades.
The study you mention, does not show data for them.
Having ported one >1 million LOC C++ app to use modules in two compilers, the compile time improvement of modules over PCH was not distinguishable from noise.
Modules have many advantages, like better encapsulation, etc.
The main thing people want from them seems to be better compile times, which is the one thing they don't deliver, at least over the PCH solutions that have existed for decades, are already supported by all build systems, etc.
Compared to modules, PCHs are "zero-effort" and deliver performance instantaneously.
CMake supports these with all major compilers...
I'll just google "<your build system> pre-compiled headers" and see if there is a flag or option that you can enabled.
You will definetly need quite a bit of fine tuning for apps over 500k LOC or so, but if your app is under that, and you are splitting code between .h and .cpp files appropriately, just flipping a flag might get you 80% there.
The speed ups you see people get from PCHs is like 20-30% faster compile-times. So they are more a "nice to have" feature than something that will solve your compile-time problems.
If your app is structured in such a way that it takes 20 min to compile, this can cut it to 15 min at most, but that would probably still suck. If you want more, then you'd need to consider other solutions like distributed build caches (sccache, etc.).
Modules inhibit parallelism because modules are ordered along a DAG and must be compiled from the root of the DAG down to the leafs in order. So consider a traditional setup as follows:
A.cpp <- A.h <- B.h <- C.h <- D.h
B.cpp <- B.h <- C.h <- D.h
C.cpp <- C.h <- D.h
D.cpp <- D.h
All four of those cpp files can be built in parallel, even though you're right that all of the header files are being reparsed multiple times. My claim is that parsing header files is incredibly cheap, it's translating the .cpp files that's expensive because cpp files are where the bulk of the semantic analysis and type checking is performed.
With modules, the same compilation model looks like this:
A.mxx <- B.mxx <- C.mxx <- D.mxx
There's no longer header/source and there's no longer redundancy, but I can't build this in parallel anymore. I have to first build D.mxx, then C.mxx, then B.mxx then A.mxx in serial.
$ clang -x c++ -E - <<<"#include <vector>" | wc -l
27378
Other headers are similar: algorithm 23103
array 23450
memory 15909
random 52107
thread 31424
tuple 9240
(Of course, a bunch of this code is shared, e.g. including both thread and vector is “only” 35713 loc total, not 60kloc.)I believe C++ compilers have SIMD-accelerated lexers/parsers because of the sheer explosion of code due to headers and templates.
Modules do not have any effect on the linker one way or another. They are independent of it.
Some C++ developers can keep using their pre-historic UNIX like tooling, whereas others will embrace the fusion of compiler, linker and build system.
Other languages with modules have a similar issue. Go is the only language I know of that makes it a hard compiler error to import an unused module.
It was part of a huge push where we spent months focusing primarily on code health and performance. The guy who did this part of the work said it needed a bit of manual intervention to really work, back then, but in the end it helped us eliminate a lot of includes and really speed up compilation.
The very measurable gains also led to new guidelines for how to write our code to try and maintain the speed we'd gained.
Of course, that approach was only straightforward by imposing other constraints. It would not work for a random project.
One of the downsides of high language complexity is that “simple” concepts like this require a whole compiler to be able to handle every case.
Another downside of high language complexity is that we probably only care about unnecessary includes because there is such a cost to just referring to things. If module references were cheap, easy to cache, etc. then it wouldn’t matter if we have a few extra ones.
If this is in third-party libraries, you can use IWYU Mappings [3] to map the "private" headers (usually the transitive include) to the public interface. An example that I use for the PEGTL library [4].
[1]: https://github.com/include-what-you-use/include-what-you-use...
[2]: https://github.com/anand-bala/signal-temporal-logic/blob/800...
[3]: https://github.com/include-what-you-use/include-what-you-use...
[4]: https://github.com/anand-bala/signal-temporal-logic/blob/800...
I wonder if there is any way to achieve that sort of thing with C/C++ and other languages.
Ruby though... I really hate it. Overloading on numbers for example. Which module lets you do 5.some_verb? Where did you import it?
In python and C you have imports local to the file or included headers. In ruby it doesn't matter as long as it's been imported somewhere in the same runtime??
Absolutely bonkers.
PS: I am quite new to ruby so please enlighten me if I have it all wrong :)
It's been a while since I really used Ruby in anger but to discover where these things came from you need to just look up the documentation of your dependencies and their dependencies, and see where something got added to. For me most of it came from ActiveSupport, and I ended up using that library by itself in other projects than just Rails.
Exactly, which seems directly opposed to the Ruby ethos of happy developers.
I do appreciate the magic of having stuff like 5.bytes or 8.days, but I don't see the reason for having runtime imports or at least warnings that you're using modules required elsewhere.
How close can this tool get to that goal?
Anyway, I'm a beginner in C/C++ world and the most convincing solution I've found to use in my personal project is the Single Compilation Unit approach (https://en.wikipedia.org/wiki/Single_Compilation_Unit). It is exemplified in the Handmade Hero github repository (which I'm afraid is available for paying users only). Essentially, the whole program is divided into modules, each within its own single cpp file. The modules are then all included in the SCU, which is the only file passed to the compiler. There can be no circular dependencies between modules (as then, there would be no order of including them in SCU which would work). In HH's case, there seems to be an absolutely minimal number and volume of headers and they only define data structures, never declare functions.
Include-what-you-use: Clang tool to analyze includes in C and C++ source files - https://news.ycombinator.com/item?id=10958186 - Jan 2016 (40 comments)
Include what you use, remove superfluous #includes - https://news.ycombinator.com/item?id=2582115 - May 2011 (1 comment)
It is possible to explicitly open this module, so that one can use `map` instead, but that's generally not wise.
I find having to write a long list of include directives at the top of a file quite annoying, and this also does not betray in what module exactly bindings are defined that one might encounter in the code below them. If I encounter, say, `Net.Tcp.open` in Ocaml code, I know that this function is defined in `./net/tcp.ml`.
> CAVEAT
> This is alpha quality software -- at best (as of July 2018). It was originally written to work specifically in the Google source tree, and may make assumptions, or have gaps, that are immediately and embarrassingly evident in other types of code.
> While we work to get IWYU quality up, we will be stinting new features, and will prioritize reported bugs along with the many existing, known bugs. The best chance of getting a problem fixed is to submit a patch that fixes it (along with a test case that verifies the fix)!
https://github.com/include-what-you-use/include-what-you-use...
Further useful docs:
Why Include What You Use? https://github.com/include-what-you-use/include-what-you-use...
What Is A Use? https://github.com/include-what-you-use/include-what-you-use...
Why Include What You Use Is Difficult https://github.com/include-what-you-use/include-what-you-use...
Fingers crossed.
https://cmake.org/cmake/help/latest/variable/CMAKE_LANG_INCL...
It also has link what you use.
https://cmake.org/cmake/help/latest/prop_tgt/LINK_WHAT_YOU_U...
The code base I'm working in is very large and I have a recurring problem where I see a term (class/variable/etc) being used in a cpp file, and want to know which header file contains the definition.
What's the quickest, easiest way to do this?
I've been using grep, but the size of the code base, combined with the large number of #includes in each cpp file, makes this inefficient.
I believe I can use ctags/vim, but I last used that circa 2000 and I'm curious to know what other static analysis solutions have cropped up since then.
Does IWYU address this scenario? I'm using clang as a compiler if that's at all relevant.
You can find a Language Server Protocol implementation for your editor at [2] (I don't think it lists __all__ clients, but it should include the most popular ones).
EDIT: I realized that this is a vague answer, so let me clarify.
An LSP implementation (especially clangd) provides actions like `go-to definition` or `find references` that you would find in full-featured IDEs like CLion (which is also amazing BTW). Since you mentioned vim, I am guessing you use it and don't necessarily want to let go of the hand-crafted vimrc you have created. Adding an LSP plugin to Vim is incredibly easy and gives you these "IDE" features with customizable mappings.
Other responses, thanks for your input. Just want to clarify that I have tried VS and VSCode with limited success (sometimes search works, sometimes it doesn't, and my biggest gripe is an occasional lack of transparency into what's going on under the cover). I think any solution is going to require some investment on my part and LSP sounds like a good investment.
A good IDE will have a feature to let you locate the declaration and/or definition of any variable or type.
I've found that a lot of IDEs have that feature completely broken. Qt Creator, for example, is easily confused and comes with all kinds of Qt garbage^H^H^H^H^H^H^H^H baggage. CLion is a resource hog and often just hangs. Visual Studio is usually pretty good -- assuming you're using Windows. VS _Code_ is "okay" but I've found it's more of a headache to set up. I don't have experience with XCode since I've never used OSX for development.
I've found the most reliable way is to learn how to use `grep` and pair that with understanding where to search; the project source directory of course but also system headers and any libraries installed to non-system locations. That knowledge translates to usefulness in other workflows too.
Is that correct? Do you use if for functionality beyond the 'jump to definition/jump back to previous context'?
I really should get with the times and try LSP like others suggested.
- Yes, ctags/vim would work
- You could use something like vscode
- Consider checking out cscope. With cscope you can also build a reverse index which lets you find where things are called. It can be used with something like vim but also has a pretty nice TUI.
Please check out the Guidelines and the FAQ, [0][1] regarding how best to participate. HN is friendly to curiosity, but not to off-topic comments, which is why you've been downvoted. If you'd like to discuss how to become a programmer, either find a thread where that's being discussed, or submit an Ask HN thread, following the style of this thread [2].
[0] https://news.ycombinator.com/newsguidelines.html