Erase your darlings: immutable infrastructure for mutable systems
grahamc.com
grahamc.com
It's a pretty great system for such an application. All the details of all the specialized, rarely-touched, hard-to-remember moving parts (nftables syntax, dual-stack DNS resolution, RS-232 serial connection parameters, etc. etc.) are all neatly collected under /etc/nixos for future me to puzzle out... and under version control, backed up offsite. It would be pretty easy for me to swap out failed hardware or upgrade it.
I wouldn't mind getting more infrastructure set up on these lines, and then maybe figure out a good setup for NixOps.
It wasn't too bad learning a little Nix to keep the configs DRY, modularized and parametrized. I find the results clean & readable, even though I'm hardly a grade-A FP propellerhead.
My main complaint was that nftables rules need to be expressed as dead strings instead of proper objects in the Nix language, which limits their composability. This would be a nice thing for the wish list.
I mean, I get that NixOS can pin versions of everything, and we've all been bitten by a new server build on identical CM code failing because of an upstream version change. But it's an eternal problem: pinning all the versions means that you a now micromanaging a zillion different versions of shit that you don't want to care about.
In my previous job where I used it heavily, every module had an enable parameter. If it was true, the module would follow one logic branch (install package, write config, perform action) but if it was false then it would follow another (remove package, delete config, perform action). This was a basic standard, and a new module would be rejected if it wasn't possible to toggle it on/off and configure it from Hiera (hierarchical yaml data). It wasn't perfect, but once the complex stuff was dockerized the puppet codebase shrank a lot. Most of the differences between our VMs were defined in Hiera (edit a yaml, toggle something on, override some default).
This is one area where NixOS shines: The ability to remove things from your configuration with no mental overhead. Even if you take great care to do all the cleanup work in your runbook, something is bound to get missed and you still haven't accounted for the state of the system before the runbook was run the first time.
For NixOS it's easy: "Not in config" = "It's not there". All system configurations are themselves normal Nix expressions that get evaluated into derivations which are then built into immutable paths under /nix/store. Most files in /etc are simply symlinks to files in the Nix store, and the activation script takes care of creating and deleting those symlinks when a new configuration is activated.
It's the difference between "you can use this tool safely" and "the tool makes sure that you will use it safely", basically.
This is one common question people ask when I introduce them to Nix and its ability to pin packages. They are often surprised to hear that you only need to pin one thing, which is the version of the Nixpkgs package collection.
Nixpkgs [1] is the central repo that contains all the Nix expressions for the "official" packages. It contains everything Nix needs to know to build a package. When you rebuild your system configuration, Nix will resolve all the packages using your local checkout of Nixpkgs into store derivations (.drv).
A .drv file uniquely identifies a specific version of a package. When any input to the expression change (source URLs, dependencies, etc.), it will result in a different derivation that will be installed under different paths in /nix/store. In other words, different inputs -> different outputs.
This is where binary distribution comes into play. Nix builds packages in sandboxes with tight isolation (no access to network or outside fs). Because the resulting store paths depend on the inputs being the same, for the same store path you will pretty much [2] get the exact same thing no matter where it's built. So Nix will lookup the binary cache for a matching object with the same store path. If it fails to find a match (e.g., because you modified the expression), Nix will simply build it. Binary cache can thus be seen as a transparent optimization that will get you binaries that are the same and will behave the same, as if you built it locally.
To summarize, when you use NixOS, you already have the source (Nixpkgs) to be able pin everything. In fact, you cannot rebuild the system without it. If the expressions don't change, the results won't. The only thing you need to to is to actually pin it to a specific commit and keep track of it in your VCS.
NixOS predates puppet by 2 and chef by 7 years.
- Actually reproducible builds, as opposed to "only if all the external factors involved, like time and third party packages, are identical between builds". One thing this enables is building your system packages in your automated pipeline and just downloading the resulting cache.
- Much more compact than either. For an example, the last state of my own Puppet configuration was [1] - five directories with hundreds of files. The current state is 13 files total, only one of which contains the Nix code actually needed to set up dozens of packages and services for a full desktop OS. That file clocks in at 346 lines, and comes in at a fraction of the size and complexity of the less featureful Puppet.
- Nix shell[2], where you can build a set of packages for use within a single project without having to worry about it borking other projects (like adding things to $PATH).
- Very few shell commands are ever needed.
- Built-in features I haven't seen in either of the other systems: allowing unfree packages either globally or individually, setting up UEFI, LUKS and LVM with a handful of lines, and stopping the compilation immediately if there's a problem rather than continuing by default.
The only caveat is the Nix language itself. The whole thing might be more complicated than the Puppet language, but you can see for yourself that the current configuration[3] isn't exactly rocket science.
[1] https://gitlab.com/victor-engmark/root/-/tree/daa391e1d957f7...
[2] https://gitlab.com/victor-engmark/root/-/blob/master/shell.n...
[3] https://gitlab.com/victor-engmark/root/-/blob/7f52eccc8b4c4c...
If so then it could genuinely replace any existing CM system. But displacing the existing tech, rewriting the whole CM codebase and getting the team over the learning curve is a monumental.
I get that Puppet isn't 100% reproducible, but my point (in another part of this thread) is that time marches on and there are only so many things you need to (or want to) 'pin'. In the past, our practice was to leave software versions all 'latest' except for known issues or requirements. Otherwise you're in a constant fight with the security folks over old software versions. If we hit an issue on a puppet run in our test environment, we put the brakes on the puppet rollout until it's. Once it gets through a few days in test, we promote it to the next environment, and eventually do a staged release through the production environments.
You have to get past runbooks to actual scripts to do this, but that's not a bad goal to have. Change history in wikis leaves a lot to be desired, and when you're trying to go from beginner to mastery, the change history can help a lot.
You're gonna find out real fast if #2 was also required to fix the problem. Or you can find out in four months when the servers need a fix for an advisory.
Etckeeper does that for /etc using git. It can even push/pull changes.
And something else can update or rollback software deployments. Apt, yum, zypper do that just fine.
"Hey, what's the store hash of the thing that's failing, so I can grab it and reproduce it locally" is a thing you can just do with Nix.
SystemImager and rsync modules did what Kickstart and CFEngine could not: take me from COTS to Custom in this lab or one in another country.
Kiwi did a pretty good job for a bit, but locked you to SUSE.
Nix seems like a runtime environment rather than an OS, but I haven’t spent the sweaty hours with it that I have the others.
Ansible and Debian work equally well on a laptop VM, Cloud Provider or colo rack.