Rust in illumos
wegmueller.it
wegmueller.it
When distros are charging themselves with things like security, shared libraries do become a bit load bearing. And for users, the idea that you update some lib once instead of "number of software vendor" times is tempting! But as a software developer, I really really enjoy how I can ship a thing with exactly a certain set of versions, so it's almost an anti-feature to have a distro swap out functionality out from under me.
Of course there can be balances and degrees to this. But a part of me feels like the general trend of software packaging is leaning towards "you bundle one thing at a time", away from "you have a system running N things", and in that model I don't know where distro packagers are (at least for server software).
I think what happens in practice is package maintainers do the work to get things up to date, and so only include a subset.
As I understand it, the process goes something like this:
- Debian wants to package a Rust tool
- They examine the tool's Cargo.toml file to determine its immediate dependencies
- Those Rust dependencies are converted into Debian package names
- Generate a Debian package definition that lists those dependencies under "Build-Depends"
- When building the Debian package in an isolated environment, Build-Depends packages are installed, providing things like `rustc` and `cargo` as well as putting the dependencies' source into the "local dependencies" directory
- Cargo runs, configured to only look in the "local dependencies" directory for dependencies
- the resulting binary is scooped up and packaged
If it were up to me, I’d consider carving off part of the Debian package namespace - like cargo—* for cargo packages, and just auto-create them (recursively) based on the package name in cargo. There are a few things to figure out - like security, and manually patched packages, and integrating with cargo itself. But it seems like it should be possible to make something like this work with minimal human intervention.
The only downside I can see is disk space usage on the builder machines, but it's source code... is it really that big?
That sounds close to what I had in mind, but I assume you run into problems with the finer points of Rust's dependency model. Like if I depend on libfoo v1, and you depend on libfoo v2, there are a lot of situations (not all) where Cargo will happily link in both versions of libfoo, and applications that depend on both of us will Just Work. But if libfoo needs to be its own Debian package, does that mean every major version is a different package?
That's the first question that comes to mind, but I'm sure there are others like it. (Optional features? Build vs test dependencies?) In short, dependency models (like Cargo's and Debian/APT's) are actually pretty complicated, and ideally you want to find a way to avoid forcing one to represent all the features of the other. Not that I have any great ideas though :p
Just curious, how do you handle security monitoring and updates when doing this?
I've been exposed to two major approaches:
At $FAANG we're on a treadmill where company CI force builds your stuff every time anything in your requirements.txt releases a new version (semver respected, at least) because in the absence of targeted CVE monitoring (which itself won't cover you 100%), that's the only practical way they see to keep up with security updates. If a dependency change breaks your project and you leave it broken for too long, peeps above you start getting autocut tickets. This leads to a lot of ongoing maintenance work for the life of the project because eventually you get forced into newer minor and eventually major versions whether you like it or not.
Outside of $FAANG I've gotten a whole lot of mileage building customer software on top of Debian- or Ubuntu-packaged dependencies. This allows me to rely on the distro for security updates without the churn, and I'm only forced to take newer dependencies when the distro version goes EOL. Obviously this constrains my library selection a lot, but if you can do it there's very little maintenance work compared to the other approach.
I'd like to hear how others handle this because both approaches obviously have considerable downsides.
The joys of not having to deal wit the extremes of any problem...
Personally, the update story is why I find developing on Windows much easier than developing on Linux. Instead of hundreds of small, independent packages, you can just target the _massive_ Windows API. You can take the "living off the land" approach on Linux too, but it's harder, often relying on magical IOCTLs and magical file paths rather than library function calls.
Not that I think the Win32/WinRT API is particularly well-designed, but at least the API is there, guaranteed to be available, and only breaks in extremely rare circumstances.
With the European Cyber Resilience Act coming into full effect late 2027/early 2028 (depends on when it's finally signed) this will not always be an option anymore for a lot of people.
a) You'll have to provide SBOMs so you can't (legally) "hide" stuff in your binary
b) Security updates are mandatory for known exploitable vulnerabilities and other things, so you can't wait until a customer asks.
This will take a few years before it bites (see GDPR) but the fines can be just as bad.
While legislatively requiring SBOMs is obviously a good idea, it might also unintentionally incentive companies to hide dependencies by rolling their own. Afterall, it's not a "dependency" if it's in the main code base you developed yourself.
Not sure how likely this is, but especially in case of network stacks or cryptographic functions that could potentially be disastrous.
But the CRA also has provisions here as you have to do a risk assessment for your product as well and publish the result of that on your documentation and hand rolling crypto should definitely go on there as well.
Again... People can just ignore this and probably won't get caught anytime soon....
In some ways, the pressure of requirements like this may be positive: “Do we really need a dependency on ‘leftpad’?”
But it will also add pressure to make unsound NIH choices as you suggest.
I just hope there will be actual fines, even for "small" companies. GDPR enforcement against most companies seem rather lackluster if it happens at all. If the ECRA ends up only applying to companies like Google, it'll be rather useless.
Good way to have malpractice lawsuits filed against you and either become bankrupt or go to jail.
https://dsc.duq.edu/cgi/viewcontent.cgi?article=3915&context...
Why Are Unpatched Vulnerabilities a Serious Business Risk? https://prowritersins.com/cyber-insurance-blog/unpatched-vul...
Remember, the desire here is simply to keep up with security updates from dependencies. There is no customer requirement to be using the latest dependencies, but this approach requires you to eventually adopt them and that creates a bunch of work throughout the entire lifecycle of the project.
I'm seeing an uptick in what they call "atomic" systems, where this is exactly the case, but last I tried to install one, it didnt even register correctly on boot up, so until I find one that boots up at all, I'll be on POP OS.
Soft forks are also good, but doesn't illumos have a stable driver ABI, such that you could just make your drivers completely independently? I thought that was how it had ex. Nvidia drivers ( https://docs.openindiana.org/dev/graphics-stack/#nvidia )
The rough areas where I think things would be useful, and probably also roughly the order in which I would seek to do them:
* Rust in the tools we use to build the software. We have a fair amount of Perl and Python and shell script that goes into doing things like making sure the build products are correctly formed, preparing things for packaging, diffing the binary artefacts from two different builds, etc. It would be pretty easy to use regular Cargo-driven Rust development for these tools, to the extent that they don't represent shipped artefacts.
* Rust to make C-compatible shared libraries. We ship a lot of libraries in the base system, and I could totally see redoing some of the parts that are harder to get right (e.g., things that poke at cryptographic keys, or do a lot of string parsing, or complex maths!) in Rust. I suspect we would _not_ want to use Cargo for these bits, and probably try to minimise the dependencies and keep tight control on the size of the output binaries etc.
* Rust to make kernel modules. Kernel modules are pretty similar to C shared libraries, in that we expect them to have few dependencies and probably not be built through Cargo using crates.io and so on.
* Rust to make executable programs like daemons and command-line commands. I think here the temptation to use more dependencies will increase, and so this will be the hardest thing to figure out.
The other thing beyond those challenges is that we want the build to work completely offline, _and_ we don't want to needlessly unpack and "vendor" (to use the vernacular) a lot of external files. So we'd probably be looking at using some of the facilities Cargo is growing to use an offline local clone of parts of a repository, and some way to reliably populate that while not inhibiting the ease of development, etc.
Nothing insurmountable, just a lot of effort, like most things that are worth doing!
Just pointing this out because I had a fair share of issues with Cargo and ultimately moved to Bazel and a bit later to BuildBuddy as CI. Since then my builds are reliable, run a lot faster and even stuff like cross compilation in a cluster works flawlessly.
Obviously there is some complexity implied when moving to Bazel, but the bigger question is whether the capabilities of your current build solution keep up with the complexity of your requirements?
For at least one project at Oxide we use buck2, which is conceptually in a similar place. I’d like to use it more but am still too new to wield it effectively. In general I would love to see more “here’s how to move to buck/bazel when you outgrow cargo) content.
https://github.com/bazelbuild/examples/tree/main/rust-exampl...
I wrote all of those examples and contributed them back to Bazel because I've been there...
Personally, I prefer the Bazel ecosystem by a wide margin over buck2. By technology alone, buck2 is better, but as my requirements were growing, I needed a lot more mature rule sets such as rules OCI to build and publish container images without Docker and buck2 simply doesn't have the ecosystem available to support complex builds beyond a certain level. It may get there one day.
I fully agree with the ecosystem comments; I’m just a Buck fan because it’s in Rust and I like the “no built in rules” concept, but it’s true that it’s much younger and seemingly less widely used. Regardless I should spend some time with Bazel.
https://github.com/diesel-rs/diesel/blob/master/examples/pos...
The parallel testing with dangling transactions isn't Diesel or Postgres specific. You can do the same with pure SQL and any relational DB that supports transactions.
For CI, BuildBuddy can spin up Docker in a remote execution host, then you write a custom util that tests if the DB container is already running and if not starts one, out that in a test and then let Bazel execute all integration tests in parallel. For some weird reasons, all tests have to be in one file per insolated remote execution host, so I created one per table.
Incremental builds that compile, tests, build and publish images usually complete in about one minute. That's thanks to the 80 Core BuildBuddy cluster with remote cache.
GitHub took about an hour back in April when the repo was half in size.
There is real gain in terms of developer velocity.
Right now we have decades of accumulated make stuff, with some orchestrating shell. It's good in parts, and not in others. What I would ideally like is some bespoke software that produces a big Ninja file to build literally everything. A colleague at Oxide spent at least a few hours poking at a generator that might eventually suit us, and while it's not anywhere near finished I thought it was a promising beginning I'd like us to pursue eventually!
In an open source world where multiple shipping operating systems are based on (different snapshots of) the same core code and interfaces, it seems much more valuable to try and push straight to Committed interfaces (public, stable, documented) where we can.
There are plenty of cases where people use Uncommitted interfaces today in things that they layer on top in appliances and so on, but mostly it's expected that those people are paying attention to what's being worked on and integrated. When folks pull new changes from illumos-gate into their release engineering branches or soft forks, they're generally doing that explicitly and testing that things still work for them.
Also, unlike the distributions (who ship artefacts to users) it doesn't really make much sense for us to version illumos itself. This makes it a bit harder to reason about when it would be acceptable to renege on an interface contract, etc. In general we try to move things to Committed over time, and then generally try very hard never to break or remove those things after that point.
I find this to be overbroad and incorrect in the larger case.
I want one version that has been reasonably tested for everything but the parts I'm developing or working with. Having just one version for everyone generally means it's probably going to work, and maybe even has been tested.
For my own stuff, I am happy to manage the dependencies myself, but moving away from the debian approach harms the system in general.
Well, rust does support shared libraries. Call things by their name: what distros want is for all libraries to have a stable ABI. And I guess some sort of segmentation that the source of one library doesn't bleed into the binary output of another, and I think the latter part is even more tricky than ABI stability. Generics are monomorphized in downstream crates (aren't C++ templates the same?), so I think you'd either have to ban generic functions and inlining from library APIs or do dependent-tree rebuilds anyway.
Those repos really need some basic high-level information about what they are and how they work, the Forge doesn't even have a README.
This is silly. Virtually all illumos development is done by Oxide. There is no other upstream to speak of.
As far as I am concerned, illumos is Oxide.
Also, it's illumos, in lower case.
user@machine:~/illumos-gate$ git log --since 2023-09-10 | grep ^Author | awk -F @ '{print $2}' | sed 's/>$//' | sort | uniq -c | sort -n | tail
11 mnx.io
16 grumpf.hope-2000.org
18 gmail.com
21 oxide.computer
29 hamachi.org
36 richlowe.net
37 racktopsystems.com
71 fiddaman.net
91 fingolfin.org
125 me.com
I'm not going to post names because that feels a touch icky, but a little bit of cross-comparison suggests that one of those domains should be combined with oxide (read: I search for names and they work there) in a way that does probably make Oxide more than 50% of commits in the last year. Though even if that's true it doesn't make them the only player or anything.Our (Oxide’s) distro isn’t even listed on the main page for illumos.
Also the author is using the lower case i as far as I can see?
They are but they're no longer contributing to illumos...
So what did Samsung do with Joyent? Just a failed acquisition or were there plans?
Just in name. It's mostly Samsung infrastructure.