Everything You Never Wanted to Know About CMake
izzys.casa
izzys.casa
Given how CMake centric CLion is, and how abusively dumb the CMake DSL is, your project sounds like a great thing to teach CLion about.
Issues with CMake:
- The DSL is not very good. It needs proper functions, for one.
- Non-hermetic builds (also mentioned in this thread)
- No ability to easily query the build DAG
- Headers are not modelled properly (they should be a dictionary of paths -> paths, not a list of include directories)
- Build folders cannot reliably be used between configurations, leading to confusion and cache misses
- Everything is convention driven. It does not model the build graph properly.
- Globs do not work properly (maybe that has changed recently?)
- The cache is not portable across a network, or even between folders on the same machine
The C++ community deserves better!
Others in this thread have mentioned some modern alternatives:
- Buck Build (Facebook, Uber, AirBnB)
- Bazel (Google)
- Pants (Twitter)
- Please (Thought Machine)
It doesn't actually matter which of these succeeds. They all model the builds in such a way that you can easily transpile between them.
It certainly seems like specifying a build process and performing a build process are two distinct things.
To your point, integration will allow the two to be iterated upon faster, since changes can be more easily released together.
Ninja view their build idea as "assembly like". While assembly surely allows you producing very fast results, it's not easy to use at all, if you are using it directly.
That's not correct. Complexity arises from the inability of expressing what you want and lack of abstractions over platform details.
Ninja for instance is to primitive to be productive and does not have any understanding of C++.
Buildsystems like buck allow you to express things in a declarative manner eg.:
cxx_library(
name = 'foo',
srcs = glob(['src/**/*.c']),
exported_headers = glob(['*.h'])
)
cxx_binary(
name = 'app',
srcs = ['main.c'],
deps = [':foo', ':bar']
)
buck implements sophisticated (disk & network) caching, optimization and scheduling strategies to make things fast.Furthermore it will use services like watchman (if available) to precompute what things need building if you change some files for fast incremental builds.
Lastly it will strip and sort symbols in your binary to make sure the hash of your binary is always the same for the same set of inputs.
All this complexity is handled for you and all you need to do to run your executable is
buck run :appSo not sure what you argument is about. It's like saying that assembly has no high end abstractions. Sure, it doesn't. It in itself is not an argument against splitting the build into several passes.
That said, Buck looks like an interesting build system.
In general, to have proper handling of librarires and etc. the language itself should support the notion of modules, like Rust does. Then you can implement sane tools (cargo). C++ is still crippled in this regard. Though there are some ideas how to improve it:
https://medium.com/@dmitrygz/brief-article-on-c-modules-f582...
Safe to say, if you are building any C++ specifically, that is entirely and utterly negligible.
Admittedly with rust not having an ABI you’d still have to rebuild all the packages whenever the compiler was updated, but I’d like to see rustup support toolchains installed to the system.
What do you mean by this exactly? Perhaps you are thinking of the old include_directories() function rather than the now-recommended target_include_directories()? This attaches one or more include directories to each target rather than having a single global list of all include directories. You can indeed have a target for each directory of your source files (this is also recommended practice).
This even works for imported targets. These are targets that represent existing prebuilt libraries on your system, and they are increasing returned by find_package in place of the old pair of ${FOO_LIBRARIES} and ${FOO_INCLUDE_DIRS}. Internally they call target_include_directories() with PUBLIC (or INTERFACE) modifier that means that those directories will be inherited by other targets that link against this one.
So your code might look like this:
#include <foo/bar.hpp>
The project folder structure might look like this: .
└── src
└── include
└── foo
└── bar.hpp
And the mapping might look like this: {
"foo/bar.hpp": "./src/include/foo/bar.hpp"
}
This enables the build system to know exactly which headers are meant to be exposed by a library to its dependees. The build system can then tell you:(1) Exactly where a header comes from
(2) Exactly which headers a library exports
(3) If two libraries will collide in terms of headers
(4) If someone is trying to use a header that is not explicitly exported by a library.
(5) If some is accessing a header in the wrong way (e.g. abusing the layout of source-files and not using the correct include-path)
This approach scales very well. Constructing the mapping can be done via globs (globs work properly in Buck), so most projects just do:
exported_headers = subdir_glob([
('include', '**/*.hpp'),
])
Additionally, the RHS of the mapping might be another build rule, thus supporting generated headers in hermetic builds.I've used generated header files (protobuf) with CMake and they work fine. Although I prefer to create them at configure time rather than build time, even though that's philosophically incorrect, just because that way I can see them in the generated IDE projects and they don't break autocomplete.
You can also wire up Protobuf headers in the map.
As ugly as things like Maven and Gradle are, if you follow all of their conventions they get out of your way.
In addition:
- no cross-platform way to require a compiler version or way to set compiler flags
- linkers are not abstracted from you (good luck trying to get cross-platform support for relocatable binaries; cmake really bungles up rpath etc)
- also good luck trying to link both static and dynamic libraries and targets
- third-party libraries aren't really a first-class thing
- no simple way to specify "build all these things into the target directory with a conventional file layout"
80% of projects need to deal with these and not a lot else. Why is there not a list of "follow these conventions and you don't really need to write or look at cmake file"?
(To be fair: compilers, linkers, and OSes share some blame in not being more consistent but the goal of a competent build-system is to handle these things for the 80% case.)
It doesn't really matter how terrible the syntax and internal implementation are if most reasonable use-cases don't need to look at it.
The compiler version thing is definitely an issue. That said, if you're worried about having to litter `if/endif` calls all over the place, I recommend you look into CMake's generator expressions.
relative rpath support was added very recently and will be in the CMake 3.14 release. It's a shame it took this long to make it in.
third party libraries can be imported via add_library, and then setting the imported location. This allows, I should note, the ability to link against both static and dynamic libraries.
IXM is being developed to make it easier to do the layout system as well as make the 80% you've mentioned here basically a non-issue, and comes with a fairly decent (in my opinion) default layout. Because it's not even in an alpha state I've not had time to document it.
(Also, I am working on a build system replacement for CMake, which includes a compiler frontend translator so you can solve the cross platform compiler flag issue. CMake is simply being used to "brute"strap the project until it is self hosting)
Now, because they are switching to what every sane build system was 5 years ago, but over the course of 10s of 3.x versions, your CMake buildfiles are about to explode in complexity and frankly, stupidity. Lots of "does this target exist already because I need to support some older version alongside bleeding edge" and usually, if not, just copy that stuff over from their HEAD.
Everything they touch goes terrible. We just need to stop with like version 3.10, pretend CMake is dead and slap a DEPRECATED warning at the beginning. It's the only way out.
From what I've seen cargo gets most of the parts right. You can have a programmatic pre-build script bit can't take over the whole building process.
Here’s a few examples:
Dynamically using git describe to determine version and build number (which is actually just the # of commits that ever occured, to allow reproducable builds)
versionCode = cmd("git", "rev-list", "--count", "HEAD")?.toIntOrNull() ?: 1
versionName = cmd("git", "describe", "--always", "--tags", "HEAD") ?: "1.0.0"
Putting git infos into certain classes buildConfigField("String", "GIT_HEAD", "\"${cmd("git", "rev-parse", "HEAD") ?: ""}\"")
buildConfigField("String", "FANCY_VERSION_NAME", "\"${fancyVersionName() ?: ""}\"")
buildConfigField("long", "GIT_COMMIT_DATE", "${cmd("git", "show", "-s", "--format=%ct") ?: 0}L")
Using string interpolation to change the output file to include the version setProperty("archivesBaseName", "Quasseldroid-$versionName")
But the worst part is that often on StackOverflow you only find ugly hacks adding custom tasks, removing tasks, modifying them, etc with ugly hacks.And even worse, sometimes there’s just no good alternative at all.
Which is why gradle took so long to clean up some parts of their API, or introduce parallel builds.
Most of the build issues that a user will face is there system not matching the developers.
We can have pretty complicated logic (arbitrary logic in most build systems) for probing the build system, and much of this logic will inevitably get lost when compiling a linear script.
Once you put all of the complexity in the compiled script, I don't think you've gained anything over "make clean ; make" (or equivelent)
Probably one of the biggest features in a configure/build/packaging software package is ubiquity, IMO. Ideally C/C++ would have had a prescribed-but-not-required thing like cargo/setuptools/go get/etc. IMO CMake is probably the least bad offering these days.
That cmake is in the place it is nowadays is a reflection of the state of cross-platform build systems 10 years ago, and possibly some marketing.
Interesting alternatives worth checking out: Meson and gn (the latter can be used for building llvm)
Additionally I want to mention my company released also a package manager that uses buck as a packaging format: https://github.com/LoopPerfect/buckaroo
And so far over 320 libraries have been ported to buck and are maintained by our bots: https://github.com/buckaroo-pm
* It's meant to be cross-platform, so a well structured CMake file will work in Windows.
* CMake modules allow you to include source-based libraries without a lot of drama. This is especially useful for cross-compiling or embedded use-cases.
* Out-of-source builds are supported with no additional work on my part. This is great for CI generating debug, release, and other build variants from a single cloned repository.
A well structured more conventional makefile can also run on Windows. The tree I'm working on builds for Unix and Windows with the same makefile. (No cygwin or WSL either, just GNU make running on Win32. A few ifdefs and strategically placed variables.)
CMake has simply sucked less, overall, than the competitors of its time, for most common build tasks. That era is over, now that the new generation of build systems is here (Bazel, Buck, Pants, Please).
For example, with GCC you can use the -M options, and then you can -include the result in your makefile. This was not always possible. You had to manually specify the .h files for every .c, and if you omitted one, you could get a successful build but a broken program!
Then consider the process of building a shared library, which is different on different platforms, but if you restrict yourself to GCC on Linux with GNU Binutils it’s damn easy.
There are a few other, minor failings of Make. But you’re right, it’s very comprehensible and straightforward. I still say that it’s a low bar, though, by modern standards.
* "Everyone" knows GNU make.
The lead at my old job was obsessed with CMake. It was an improvement over Boost.Build(barf..), and to the degree it replaces something like autotools I'm all for it, but it was yet another DSL in a company filled with DSLs. It took us 3 months to get new devs up to speed because of all the shit they had to learn. I can take a sw dev(embedded sw, because that was the role) off the street and there's a 90% chance they're conversant in GNU make. About 90% have no idea how to write a CMake file. When I looked at the CMake files, I noticed source and target rules were spread out over the file, which seemed like a really bad idea to me.
Also, for my purposes (using both Windows and GNU/Linux at work), portability is a tremendous benefit.
With non-hermetic builds, it's much more difficult to verify the correctness of your build rules, which means that you might get non-reproducible builds or you might get incorrect incremental builds. With hermetic builds, it's always safe to do an incremental build, no matter what state the repository is in.
You then need to build that code using a consistent set of tools and libraries, which is where the chroot comes in.
One big chroot around your whole build system isn't enough, you can run into nondeterminism problems due to relative ordering of different build steps during execution. This is why it's nice that the new build systems make each individual step hermetic, because you can make a separate chroot for each individual step (Bazel executes each build step in a separate sandbox).
This allows you to cache all intermediate results, share the cache, and get the same results both for clean and incremental builds, repeatably. It's also easier to remove nondeterminism from individual build steps rather than looking at the whole build.
So you're saying this takes care of race conditions in the build where things may happen out of order even if you have a fixed environment?
- Bazel https://bazel.build/
- Buck https://buckbuild.com/
- Pants https://www.pantsbuild.org/index.html
- Please https://please.build/
Cmake seemed to work well, meaning, I could just open the GitHub project and generate an Xcode project file and bam, it just worked. However, agreeing with everyone here, holy hell, trying to start a new project with it, or migrate another project over to it is quite a futile experience.
It has several escape hatches in the weirdest places. Need to pass a conditional compiler flag? Good luck with that exercise. A compiler flag was added from a previous macro and there doesn't seem to be a way to keep it from doing that. I don't know, the whole system seems nice when you don't have to touch the make file, it just works, but I can only imagine the hair splitting effort it took to get that make file in a workable state.
Maybe just use python/etc to write a better wrapper for Makefiles? There are so many build system to choose these days.
- message(STATUS ...) or message(FATAL_ERROR ...) for printing stuff out at configure time, e.g.
message(STATUS "libcurl found: ${HAVE_LIBCURL}")
- cmake -E echo for printing stuff out at build time, e.g. add_custom_command(TARGET lorenz POST_BUILD
COMMAND ${CMAKE_COMMAND} -E echo "lorenz command line:"
COMMAND ${CMAKE_COMMAND} -E echo '$<TARGET_FILE:lorenz>' ${LORENZ_ARGS}
VERBATIM)
- diagnostic-level MSBuild output for dependency issues - you can configure this in Visual Studio, Tools, Options, Projects and Solutions, I think (it's around there somewhere). You might not expect much from MSBuild debug output, considering how annoying the rest of it is, but it's actually extremely comprehensive, and I've found it useful for figuring out even rather weird stuff. There's quite a lot of it, though, so get a cup of teaA passing familiarity with the MSBuild syntax might be helpful, but I've managed to do mostly without.