Moving SciPy to the Meson Build System
labs.quansight.org
labs.quansight.org
As a user (not a developer) of SciPy, this is the big win. My "exotic platform" is embedded Linux distributions via Buildroot[1]. This opens the door to many downstream libraries becoming available as well, such as pandas and scikit-learn.
I haven't followed Meson closely in about 3 years, but I also got the sense that Windows support sometimes lagged. If that's true, it's going to be a tough sell for the many large C++ projects who adopted CMake almost exclusively due to its support for Windows and Visual Studio.
On the other side I'm using Java with Maven and it is - a big burden. It's build in dependency retrieval system also isn't helping. Maybe I just don't like XML - because it is not human-readable ;)
The other big one is that you need to have a strong familiarity with Make, and use it often. Familiar because there's no way a person who hasn't actually been through some form of Make documentation in detail can even guess at what things like $@, $?, and $^ do, or accurately decipher its macro replacement syntax, or any of that. And use it often because, even if you were deeply familiar with Make in the past, if you haven't touched it in a few years, you're unlikely to reliably remember it without help.
:) it devolves into a custom shell script
If you use it Maven-style, with relatively simple build scripts and all the complexity pushed out to plugins, it's very readable. A fair bit more readable than Maven, in my opinion, but I wouldn't care to argue that point in particular.
If you use it the way Stack Overflow tells you to do it, so that your build scripts are basically glorified ad-hoc Groovy programs, then, yeah, it's just impossible.
Unfortunately, since the only way to really understand any of Gradle is to understand all of Gradle, and understanding all of Gradle is a huge effort, it's kind of a catch-22. For most people, the only really sensible way to use it is to Google for advice and then do what Stack Overflow tells you to do. But that invariably leads to a result that is not even remotely sensible.
Long story short, my hot take is this. Maven is write-only. Gradle (ideal usage) is read-only. Gradle (normal usage) can be neither read nor written.
But I have never been satisfied with the solutions for out-of-tree builds with python - being able to debug-in-place extensions _and_ eventually install in a single package structure seems to be somewhat in conflict with the way that Python allows you to structure packaging, without lots of environment variable hacks.
The CI for boson seems like it runs on platforms where Python definitely is available, but also I notice the CI uses samurai, a reimplementation of ninja with a similar motivation: https://github.com/michaelforney/samurai
Ninja is in C++ so I am even more confused at Sanurai.
Is this just an implementation-diversity thing? (which is great!)
Would be great to hear why those were the only two candidates that they considered (and not e.g. Bazel or Nix).
I suspect Bazel was ruled out because it requires the JVM and it has limited open source uptake relative to CMake (huge open source userbase), and Meson (limited presence in open source scientific software, but adopted by GNOME and systemd).
Based on that, I just sort of assumed that, in addition to all the fairly well-documented up-front challenges that these folks had identified and were talking about, there must also be some impassable barrier lurking around in there that nobody finds until they've already sunk a lot of time and money into trying to get it working. I don't know what that is, and I'm not curious enough to spend the better part of a year trying to find it for myself. I ended up choosing Gradle.
It went from “maybe run stdlib Python in an activated venv“ to actually working.
Depending on how minimal you want to get, I think native build systems where you're not just building for yourself, but distributing source bundles that need to be built by others, have to start the discussion with Autotools, CMake, and Meson. The reason is just that they're already required by everything else anyway, so you're not adding any burden to the downstream packagers by using them.
If you're just building for yourself, have at it, get as bespoke and obscure as you want, but SciPy isn't being built and deployed primarily on NumFOCUS's own servers. It's being distributed as a tarball to its users who mostly build it themselves. You should make every effort to use a build system they're likely to already have and already be familiar with.
In particular, much of the focus here seems aimed at users who want to cross-compile for embedded Linuxes running on ARM. I doubt those people want to pull in a JVM.
Kind of reminds me of when I was working for a geointelligence agency charged with building and integrating third-party ground processing algorithms into a user-facing web tasking framework. We have to retrieve and build the dependencies, too. xerces-c, Armadillo, MKL, no problem. Just keeping downloading, either cmake .. or configure, make, then make install, over and over.
Then one of them requires TensorFlow and it becomes a two-month research project trying to figure out how to build it (sorry Google, but the military does not trust your pre-built binaries).
Scipy is very, very hard to cross-compile. Like, almost impossible. I eventually had to compile it on-target and store the outputs for inclusion in later images.
Distutils in general is wretched at cross-compilation tasks, should be retired.
If you want to make that comparison, meson is more restrictive than cmake, but tends to have more functionality built-in to make up for it.
bazel is more all-encompasing than either, as it's more about reproducible and distributed builds than a "better makefile" type system.
It's "cmake in role only" - both are build systems, that's where the similarity falls apart.
The syntax is very pythonic, however, I think that's true of many ergonomic centric formats.
There's the 3.0+ CMakeLists.txt syntax, which is a nice declarative target-based system.
There's the *.cmake language syntax for finding dependencies and other scripting, which is a C'thonian horror mix of imperative and declarative bits that the trench coat tries to hide.
Then there's the creepy pre-3.0 stalker, the old syntax for both. Non-target based, imperative bits sneaking in everywhere, makes Autotools look sane.