Nix is a better Docker image builder than Docker's image builder
xeiaso.net
xeiaso.net
I have 2 systems running Nix, and I'm afraid to touch them. I've already broken both of them enough that I had to reinstall from scratch in the past (yes yes - it's supposed to be impossible I know), and now I've forgotten most of it. In theory, Nix is idempotent and deterministic, but the problem is "deterministic in what way?" Unless you intimately understand what every dependent part is doing, you're going to get strange results and absolutely bizarre and unhelpful errors (or far more likely: nothing at all, with no feedback). Nix feels more like alchemy than science. Like trying to get random Lisp packages to play nice together.
Documentation is just plain AWFUL (as in: complete and technically accurate, but maddeningly obtuse), and tutorials only get you part of the way. The moment you step off the 80% path, you're in for a world of hurt, because the underlying components are just not built to support anything else. Sure, you can always "build your own", but this requires years of experiential knowledge and layers upon layers of frustration that I just don't want to deal with anymore (which is also why I left Gentoo all those years ago). And woe unto you if you want to use a more modern version than the distribution supports!
The strength of Docker is the chaos itself. You can easily build pretty much anything, without needing much more than a cursory understanding of the shell and your distro's package manager. Or you can mix and match whatever the hell you want! When things break, it's MUCH easier to diagnose and fix the problems because all of the tooling has been around for decades, which makes it mature enough to handle edge cases (and breakage is almost ALWAYS about edge cases).
Nix is more like Emacs: It can do absolutely anything if you have the patience for it and the deep, arcane knowledge to keep it from exploding in a brilliant flash of octarine. You either go full-in and drink the kool aid, or you keep it at arm's length - smiling and nodding as you back slowly towards the door whenever an enthusiast speaks.
So in theory Nix would be perfect. But it's not, because it's so different. Get a tool from a vendor => won't work on Nix. Get an error => impossible to quickly find a solution on the web.
Anyway, out of that frustration I've funded https://www.stablebuild.com. Deterministic builds w/ Docker, but with containers built on Ubuntu, Debian or Alpine. Currently consists of an immutable Docker Hub pull-through cache, full daily copies of the Ubuntu/Debian/Alpine package registries, full daily copies of most popular PPAs, daily copies of the PyPI index (we do a lot of ML), and arbitrary immutable file/URL cache.
So far it's been the best of both worlds in my day job: easy to write, easy to debug, wide software compatibility, and we have seen 0 issues due to non-determinism in containers that we moved over to StableBuild in my day job.
Huh, I have never had this issue with apt (Debian/Ubuntu) but frequently with apk/Alpine: The package's latest version this week gets deleted next week.
not necessarily, eg snapshot.debian.org
> pinning on hashes is architecture dependent
can't you pin the multi-arch manifest instead?
I still like StableBuild for protection against package deletion, and mirroring non-pinnable deps
But I also agree with all the flaws of Nix people are pointing out here.
Free …
Number of Users 1
Number of Users 15GB
Is that a mistake or if not can you explain please?
Free => Access to all functionality, 1 user, 15GB traffic/month, 1GB of storage for files/URLs. $0
Pro => Unlimited users, 500GB traffic included (overage fees apply), 1TB of storage included. $199/mo
Enterprise => Unlimited users, 2,000GB traffic included (overage fees apply), 3TB of storage included, SAML/SSO. $499/mo
Here's a related discussion about reproducible builds by the Docker people, where they provide some more details: https://github.com/docker-library/official-images/issues/160...
Any source of that claim?
> or when you depend on some packages installed through apt - as you'll get the latest version (impossible to pin those).
Well... please re-read my previous comment - we do Java thing so we use any JDK base image and then we slap our distribution on top of it (which are mostly fixed-version jars).
Of course if you are after perfection and require additional packages then you can install it via dpgk or somesuch but... do you really need that? What about security implications?
Any tag like ubuntu:20.04 -> this tag gets overwritten every time there's a new release (which is very often)
https://hub.docker.com/r/nvidia/cuda -> these get removed (see e.g. https://stackoverflow.com/questions/73513439/on-what-conditi...)
Besides, if you really need utmost stability you can use image digest instead of tag and you will always get exactly the same image...
If you want reproducibility, the first step is to copy everything to a storage you control. Luckily, this is pretty cheap nowdays
Very nice project!
I've work many years on bare metal. We did (by requirement) acceptance tests, so we did need deterministic builds, before such thing had even a name, or at least before it was mentioned as much as nowadays.
Redhat has a lot of tooling around versioning of mirrors, channels, releases, updates, etc. But I'm so old that even foreman and spacewalk didn't exist, redhat satellite was out of the budget, and the project was migrating from the first versions of CentOS to Debian.
What I did was simply use DNS + Vhosts (dev, stage, prod + versions) for our own package mirrors, and bash+rsync (and of course, raid+backups), with both, CentOS and Debian (and our project packages).
So we had repos like prod/v1.1.0, stage/v1.1.0, dev/v1.1.0, dev/v2.0.0, dev/2.0.1, etc Allowing us to rebuild things without praying, backport bug fixings with confidence, etc
Feels old and simple, however I think it was the same problem/issue that people gets now (re)building containers.
If you need to be able to produce the same output from the same input, you need the same input.
BTW about stablebuild: nice project!
Documentation is often just plain erroneous, especially for the new CLI and flakes, not even edge cases. I remember spending some time trying to understand why nix develop doesn't work like described and how to make it work like it should. I feel like nobody ever actually used it for its intended purpose. Turns out that by default it doesn't just drop you into the build-time environment like the docs claim (hermetically sealed with stdenv scripts available), it's not sealed by default and the commandline options have confusing naming, you need to fish out the knowledge from the sources to make it work. Plenty of little things like this.
>In theory, Nix is idempotent and deterministic
I surely wish they talked more about edge cases that break reproducibility. Things like floating point code being sensitive to the order of operations with state potentially leaking from OS preemption, and all that. Which might be obvious, but not saying obvious things explicitly is how you get people shoot themselves in the foot.
That’s profoundly cursed and also something that doesn’t happen, to my knowledge. Unless the kernel programmer screwed up, an x86-64 FPU is perfectly virtualizable (and I expect an AArch64 FPU too, I just haven’t tried). So it doesn’t matter where preemtion happens.
(What did happen with x87 is that it likes to compute things in more precision than you requested, depending on how it’s configured—normally determined by the OS ABI. Yet variable spills usually happened in the declared precision, so you got different results depending on the particulars of the compiler’s register allocator. But that’s still a far cry from depending on preemption of all things, and anyway don’t use x87.
Floating-point computation does depend on associativity, in that nearestfp(nearestfp(a+b)+c) is not the same as nearestfp(a+nearestfp(b+c)), but the sane default state is that the compiler will reproduce the source code as written, without reassociating things behind your back.)
We don't need to be looking for such contrived examples actually, nixpkgs track the packages that fail to reproduce for much more trivial reasons. There aren't many of them, but they exist:
https://github.com/NixOS/nixpkgs/issues?q=is%3Aopen+is%3Aiss...
Less than a couple of thousand packages are reproduced. Nobody has even attempted to rebuild the entirety of the nixpkgs repository and I'd make a decent wager on it being close to impossible.
That depends whether you are okay with chaos.
It appears that you are, so it is suitable tool for you. Choose the right tool for the right job.
---
Docker is a poor choice for people who are interested in deterministic/reproducible builds.
I’ll use docker, I like docker, but I can see the point of how it’s not necessarily advantageous if stability is your main goal.
Sure, your compiler, your hardware, or your distro might be compromised, but if you follow the chain all the way through you does indeed validate version X does result in SHA y, there's now less things were blindly trusting.
It also helps with things like rolling back to earlier versions when you don't still have the binary kicking around without having to revalidate the binary.
If you're not getting the same SHA on different hardware, weeks apart, even if it's good enough for you, it's not reproducible
https://grahamc.com/blog/erase-your-darlings/
Inside my user account, I don’t bother. I don’t like Home Manager.
https://elis.nu/blog/2020/05/nixos-tmpfs-as-root/
That way it doesn't even need to bother with resetting to ZFS snapshots, instead it just wipes root on shutdown and reconstructs it in RAM on reboot.
Then, optionally, with some extra work you can put /home in tmpfs too:
https://elis.nu/blog/2020/06/nixos-tmpfs-as-home/
That setup uses Home Manager, so maybe it's not for you, but worth mentioning if we're talking about making all state declarative and reproducible. You have to use the Impermanence module and set up some soft links to permanent home folders on different drive or partition. But for making all state on the system reproducible and declarative, this is the best way afaik.
That said I'll just mention that ZFS support on NixOS is like nothing else I've seen in Linux. ZFS is like a first-class citizen on NixOS, painless to configure and usually just works like any other filesystem.
https://old.reddit.com/r/NixOS/comments/ops0n0/big_shoutout_...
Nix Doc are horrible but I've found that ChatGPT4 is awesome at troubleshooting Nix issues.
I feel like 90% of the time I run into Nix issues, it's because I decided to do something "Not the Nix way."
That has been the case for as long as I can remember. I gave up on Nix about 5 years ago because of it, and apparently not much has changed on that front since then..
Could you mention a bit about how they broke? I'm curious to see how that state looks, as from my perspective switching to a previous configuration seems to cover everything.
I've been a nixos user for years and I generally had the opposite problem: the latest of the package you want is not available but hey here's a version from months ago - or just build it yourself (which is not hard, oftentimes updates work fine with no build change, you just point at a different version).
Also rebuilding everything at every update take forever (I had a few nix-shells with ai dependencies that would take hours to upgrade).
I love the concept of nix but I'm back to Arch, binary bleeding edge packages and AUR for less supported stuff.
I think that might have changed with the release of flox: https://flox.dev, it’s basically seamless (and that’s not surprising coming from DE Shaw).
Nix doesn’t really make sense without flakes and nix-command, those things are documented as experimental and defaulted off. The documentation story is getting better, but it’s not there. nixlang is pretty cool once you learn it well, but it’s never going to be an acceptable barrier to entry at the mainstream. It’s not really the package manager it’s advertised as, nix-env -iA foo is basically never what you want. It’s completely unsurprising that it’s still a secret weapon of companies with an appetite for bleeding-edge shit that requires in-house expertise.
flox addresses all of this for the “try it out and immediately have a better time” barrier.
Nix/NixOS or something like it is going to send Docker to the same dustbin Subversion is in now, but that’s not going to happen until it has the GitHub moment, and then it’ll happen all at once.
Most of the complaints with Nix in this thread are technically false, but eminently understandable and more importantly (repeat after me Nix folks): it’s never the users fault.
I’m aware that I’m part of a shrinking cohort who ever knew a world without git/GitHub, so I know this probably sounds crazy to a large part of the readership, but listen to Linus explaining to a room full of people who passed the early Google hiring bar why they should care about a tool they feel is too complicated for them:
This is true for git, but not so true yet for Nix, so I'm not sure a GitHub-like moment helps.
But anywhere else you just brew/apt/rpm install whatever and nix develop —impure, which is easier than most intermediate git stuff and plenty of beginner stuff. git and Nix are almost the same data structure if you start poking around in .git or /nix/store, I might not understand what you mean without examples.
But all my guesses about what you might mean are addressed well by flox.
If you drop me an email at b7r6@b7r6.net (I also just joined your slack) I’d love to give my feedback on this or that nitpick.
But overall, well done friends, very very nice stuff.
Why is docker bad at this? In order to enjoy the caching benefit, each time you build a docker image you want it to output as much existing layers as possible. So running apt-get install python3 today should result in the exact same layer as yesterday, if there are no new updates. But this requires the all the files to be exactly the same, including the metadata like creation time. As docker layers are cached by hashing the files.
Now, Nix already does storing dependencies by hash. So the layers will always be the same with the same version and same configuration.
When said caching, I meant on the nodes that run the containers. With Nix you can also update only one layer, while keeping the other layers the same.
The Dockerfile format imposes a hierarchical relationship between layers. This quickly becomes very annoying, since dependencies usually form dependency graphs, not dependency trees.
Alternative tools, like nix (probably bazel too), are not bound in the same way. They can achieve fine grained caching by mapping their dependency graph to docker layers, which is something that can not be expressed with a Dockerfile.
The final result need not be. You can build a bunch of things then merge the results in a final stage without any hierarchy (this is "COPY --link" in a Dockerfile).
With nix, re-usability is very high. It's a function that is very baked in at very low levels of it's design. This comes with up front complexity but getting to these reusable layers is basically forced.
Docker is very simple and often touts reusable layers, but in practice is not. Unless you tackle that complexity.
Making reproducible and reusable content takes effort. Other tools are not designed for that. As a result the getting to the same state requires a similar amount of complexity. Worse, with docker you can never be sure that you actually succeeded in your goal of reproducibility.
An analogy could be rust. Rust has up front complexity, but tackling that complexity gives confidence that memory safety and concurrency primitives are done correctly. It's not that C _can't_ achieve the same runtime safety, it's just requires a lot more skill to do correctly; and even then memory exploits are reported on a near daily basses for very popular and widely used libraries.
Complex problems are complex. And sooner or later you'll need to face that complexity.
What you are describing is chaining a bunch of commands together. Yes, this forms a dependency chain stored in separate layers and is part of the cache chain.
Nix suffers the exact same problems with reproducibility. The thing it provides is the toolchain of dependencies that are reproducible. Docker does not provide your dependencies.
If the inputs change then so does the output. If the output itself is not reproducible (like, say an artifact with a build-time embedded in it) then you have something that is inherently not reproducible and two people trying to build the same exact nix package will have different results.
EDIT: Fixed a sentence I apparently got distracted while writing and didn't complete (about layer caching).
> The thing it provides is the toolchain of dependencies that are reproducible. [...] If the inputs change then so does the output. If the output itself is not reproducible (like, say an artifact with a build-time embedded in it) then you have something that is inherently not reproducible and two people trying to build the same exact nix package will have different results.
There are no guarantees they are reproducible. The only guarantees Nix gives you is that the build environment is the same which allows you to make some claims about the system behaving the same way. But they are certainly no guarantees about artifacts being bit-by-bit identical.
The content address of a docker image is a json blob (referencing other objects).
The content address of a Dockerfile "RUN" command is the content address of what came before it and the command being run.
In the case of Nix it's addressed by the input. Not the content of the build. It's an important distinction and one Nix also makes.
https://nixos.org/manual/nix/stable/command-ref/new-cli/nix3...
But doing this is going to give you a slight headace as most of the package repository in Nix is not checked for reproducible builds and there are no way to guarantee the hashes are actually static.
We are saying the same thing here, I'm just trying to point out this is exactly how docker build works, but rather it is more about what you are willing to put into your docker build.
Docker layers are completely independent from each other. A docker layer is a sha256 sum of that layer. Separately there is an image manifest, which is also fetched with a sha256 of that manifest, states the order of the layers and at runtime those layers are stacked up on each other.
With docker, there is no explicit dependency chain. A layer is just a tarball or some JSON. Some tooling can take advantage of this fact. How nix builds docker images takes advantage of this.
Nix on the other hand, the output hash is not tied to the output hash, because the output hash is irrelevant to how it was produced. You also cannot know in advance the output hash of something. IE. If I do say "echo foo > bar.txt", I cannot know the sha256 sum of bar.txt until the code runs. But before the code runs I can know the hash all the inputs that will create bar.txt.
This fundamental difference means two builds, executing the same code can share the outputs. Provided that the build environment is trust worthy.
While docker build can/does output OCI images, that is only an output. How that output comes to be is not the output itself, same as the nix side the article is talking about.
I see the confusion now. The nix image builders OCI layer's contents are _only_ nix store paths which _do_ include the input.
Nix store paths are are guaranteed to never overlap each other, and such the order of layering them in the docker manifest does not matter. But the docker layers are just tar balls of nix store paths. Each layer in the image has no dependence on previous or future layers at all. It is just one or more nix store paths.
What I'm saying here through multiple different threads is, buildkit and nix build things the same way. `docker build` is not just a Dockerfile builder, its actually a grpc service (with services running on both the docker CLI and in the daemon). This service is actually very generic. It includes builtin support for Dockerfiles, which just converts the Dockerfile format into what buildkit calls "LLB", which is analogous to LLVM IR.
What I'm also saying is, people are comparing "docker build" with a Dockerfile that's using a package manager that's not even provided by docker to nix. This is not an apples to apples comparison, and in fact you can implement nix packaging using buildkit (https://github.com/reproducible-containers/buildkit-nix).
I'm also saying that `Dockerfile` does actually support merging dependencies without being order dependent (this is `COPY --link`). But also, you can drive buildkit operations without going through Dockerfile. You can also plug in your own format with the `syntax=<some/image>` at the top of your file. This isn't "convert to dockerfile", its "convert to LLB", which is all the Dockerfile frontend does.
Finally, I'm saying nix isn't in and of itself some magic tool to have a reproducible build. You still have to account for all the same things. What it does do, at a package management level, is make it easier to not have dependencies that change automatically over time (which has its own plusses and minuses).
Isn't a Dockerfile just a sequence of dependencies, rather than a tree?
Nix is absolutely the best tool for what I want to achieve but it has those dark forsaken corners that just suck your soul out dry.
I love it but sometimes it feels like being a Morty on Rick’s adventure to the compilerland.
Before someone asks. I've been using macOS for a long time. I've never seen remnants like this from a program. Sure, there are often directories left in ~/Library/Application Support/, but this was more than that. Unfortunately, I didn't write down the details, but I ran across the bits in at least 3-4 places.
- ~/.orbstack
- Docker context that points to OrbStack (for CLI)
- "source ~/.orbstack/shell/init.zsh" in .zprofile/bash_profile (to add CLI tools to PATH)
- ~/.ssh/config (for convenient SSH to OrbStack's Linux machines)
- Symlinks to CLI tools in ~/.local/bin, ~/bin, or /usr/local/bin depending on what's available (to add CLI tools to existing shells on first install — only one of these is used, not all)
- Standard macOS paths (~/Library/{Application Support, Preferences, Caches, HTTPStorages, Saved Application State, WebKit})
- Keychain items (for secure storage)
- ~/OrbStack (empty dir for mounting shared files)
- /Library/PrivilegedHelperTools (to create symlinks for compatibility)
Not sure what the best solution is for people who don't use Homebrew to uninstall it. I've never liked separate uninstaller apps, and it's not possible to detect removal from /Applications when the app isn't running.
And cough since we’re at with - did you consider Nixpkgs distribution?
I’m slowly moving deeper and deeper into ecosystem and use Home Manager for utilities that I use often (and use nix shell/nix run for one offs). Some packages are strictly GUI and while they aren’t handled flawlessly (self-updaters) it’s nice to have them on a single list.
Yet based on your list it’s a definitely a nixventure…
And yet, you are the only(?) one with that knowledge, so the alternative seems to be replying to HN threads with a curated list of things that a user must now open iTerm2 and handle by themselves. Something, unless I'm mistaken, that computers are really good at doing (Gatekeeper and privilege elevation nonsense aside)
Even just linking to the zap portion of your brew cask could go a long way since it would be the most succinct manifest if I correctly understand what it does
I agree this should be documented, but I still appreciate uninstallers.
Also, I'm a little confused about your statement:
> Not sure what the best solution is for people who don't use Homebrew to uninstall it.
You said at the start you've "been meaning to update the Homebrew cask to be more complete on zap" ... does that mean Homebrew uninstall will not do a complete job?
But unfortunately cross compiling quickly broke when I started doing mild customization (and one reasons I’m doing this is a complex setup that’s very sensitive to version changes).
In the end solution was to “simply” get darwin.linux-builder up but that pulled a lot of weight behind it.
It works, but it’s not the first time I spent my time on nix-ventures.
If something does break, rollbacks are free and an integral part of Nix.
This sort of twenty-minute adventure?
The rougher part is the speed of it; it's a one-two punch between the hardware & the fact that Docker has to emulate a Linux VM.
Things are straight forward on Linux. You build your binary, place in a docker container and you are done. The nix code will also be straight forward. If you can build your code, then creating a container is just one more operation away.
Unfortunately docker requires Linux binary and you are on Mac. So the docker desktop actually runs a Linux VM and performs all operations on it, abstracting this away from you.
Nix doesn't do that and you have two options:
1. Do cross compilation, the problem is that for this to work you need to be able to cross compile down to glibc, the problem is that while this will work for most community used dependencies you might get some package where the author didn't put effort making sure it cross compile. To make things worse the Hydra that populates standard caches that nix uses, doesn't do cross compile builds, so you will run into lengthy processes that might potentially end with a failure.
2. You can have a Linux builder, that you add to your Mac and configure to send build jobs for x86_64-linux to that builder. Now you could have a physical box, create a VM, or even have a NixOS docker container (after all docker ion Mac runs inside of the VM).
The #1 seems like the proper way, while #2 is more of a practical way.
I think you are running into issues, because you're likely trying #1, and that requires a lot of experience not only with Nix, but also with cross compiling. I wish Nix's Hydra would also build Darwin to Linux cross compilation as that would not only provide caches, but also help making sure the cross compilation doesn't break, but that would also increase costs for them.
I think you should try the #2 solution.
Edit: looks like there might have been an official solution to this problem: https://ryantm.github.io/nixpkgs/builders/special/darwin-bui... I haven't used it yet.
I'm using `clang` from `pkgs.pkgsCross.musl64.llvmPackages_latest.stdenv` to cross-compile Rust binaries from ARM macos to `x86_64-unknown-linux-musl`. It _works_, but every time I update my `flake.nix` it rebuilds *the entire LLVM toolchain*. On an M2 air, that takes something like 4 hours. It's incredibly frustrating and makes me wary of updating my dependencies or my flake file.
The alternative is to switch to dockerized builds but:
1) That adds a fairly heavyweight requirement to the build process
2) All the headache of writing dockerfiles with careful cache layering
3) Most importantly, feels like admitting defeat.
I pass the output of each stage of a toolchain as a dependency to the next stage. By chaining the stages, changes made to a single stage only require a rebuild of each succeeding stage. The final stage is the default of the flake, so you can easy get the complete package.
In addition, I can debug along the toolchain by entering a single stage env with nix develop <stage>
Not sure if this is the most optimal way, but it appears to work in modularizing the rebuild.(using 23.11)
"I love it but sometimes it feels like being a Morty on Rick’s adventure to the compilerland."
I would be very interested to know where the difference is; is nix including things it doesn't need to? Is the non-nix build not including things it should?
As stated above, I havent done any trimming on the resulting image, so There's too many stuff in the image.
And sorry for gradle2nix, I'm working on an improvement that's less of a hack.
Don't be. Thanks for your work. Excited to learn about the improvement. Can you tell more about what you have in mind?
It looked something like this: https://gist.github.com/takeda/17b6b645ad4758d5aaf472b84447b...
So what I did was:
- link everything with musl
- compile python and disable all packages that I didn't use in my application
- trim boto3/botocore, to remove all stuff I did not use, that sucker on it's own is over 100MB
The thing is what you need to understand is that the packages are primarily targeting the NixOS operating system, where in normal situation you have plenty of disk space, and you rather want all features to be available (because why not?). So you end up with bunch of dependencies, that you don't need. Alpine image for example was designed to be for docker, so the goal with all packages is to disable extra bells and whistles.
This is why your result is bigger.
To build a small image you will need to use override and disable all that unnecessary shit. Look at zulu for example:
https://github.com/NixOS/nixpkgs/blob/master/pkgs/developmen...
you add alsa, fontconfig (probably comes with entire X11), freetype, xorg (oh, nvm fontconfig, it's added explicitly), cups, gtk, cairo and ffmpeg)
Notice how your friend carefully extracts and places only needed files in the container, while you just bundle the entire zulu package with all of its dependencies in your project.
Edit: tadfisher seems to be more familiar with it than me, so I would start with that advice and modify code so it only includes a single jdk. Then things that I mentioned could cut the size of jdk further.
Edit2: noticed another comment from tadfisher about openjdk_headless, so things might be even simpler than I thought.
https://discourse.nixos.org/t/how-to-create-a-docker-image-w...
I agree with your assertion regarding the language though. I think nix-lang makes it harder to get into Nix.
I wouldn't expect a layman to be able to grok that file. That's fine though - it's not for laymen.
> I wouldn't expect a layman to be able to grok that file. That's fine though - it's not for laymen.
This is the kind of comment that makes me want to stay far, far away from Nix and the Nix "community".
I actually pointed out that mkDerivation is something helpful to learn - that's one thing I wish someone made me sit and learn when I first got exposed to Nix. It unlocks a lot.
https://github.com/NixOS/nixpkgs/blob/master/pkgs/developmen...
What that file is doing, is building a package, and it essentially is a combination of what Makefile and what RPM spec file does.
I don't know if you're familiar with those tools, but if you aren't it takes some time to know them enough to understand what is happening. So why would be different here?
- it contains information how to actually build the application
- how to set up a dev environment
- how to build application with musl
- how to build application with glibc
- how to build python with only with expat, libffi, openssl, zlib packages
- how to take botocore and patch it up to only have cloudformation, dynamodb, ec2, elbv2, ssm, sso, sts clients
Try to get all of that into a single Dockerfile and see how complicated mess you end up with.
The actual docker configuration is here:
https://gist.github.com/takeda/17b6b645ad4758d5aaf472b84447b...
It might be still confusing to you at first, as you're used to list of incremental steps how to get to the final result, while this description instead is declarative (you're describing not the steps to do, but what the final image should be).
It's basically comparing bash script with bunch of "aws" CLI invocations to a terraform or cloudformation file.
Wonder if there is a good term for this. I have been jokingly referring to this as 'headed' and headless as 'beheaded'.
I've switched to using a headless OpenJDK build from Nixpkgs as a baseline instead of Zulu, to remove all the unnecessary dependencies on GUI libraries. Then I've used pkgs.jre_minimal to produce a custom minimal JRE with jlink.
The image size now comes out to 161MB, which is slightly larger than the demo_jlink image. This is because it actually includes all the modules required to run the application, resulting in a ~90MB JRE. The jdeps invocation in Dockerfile_jlink fails to detect all the modules, so that JRE is only built with java.base. Building my minimal JRE with only java.base brings the JRE size down to about 50MB, the resulting (broken) container image is 117MB according to Podman.
I've also removed the erroneous copyToRoot from your call to dockerTools.buildImage, which resulted in copying the app into the image a second time while the use of string context in config.Cmd would have already sufficed.
I've also switched to dockerTools.buildLayeredImage, which puts each individual store path into its own image layer, which is great for space scalability due to dependency sharing between multiple container images, but won't have an impact for this single-image experiment.
This is mostly a JRE size optimization challenge. The full list of dependencies and their respective size is as follows:
/nix/store/v27dxnsw0cb7f4l1i3s44knc7y9sw688-zlib-1.3 125.6K
/nix/store/j6n6ky7pidajcc3aaisd5qpni1w1rmya-xgcc-12.3.0-libgcc 139.1K
/nix/store/l0ydz31lwa97zickpsxj2vmprcigh1m4-gcc-12.3.0-libgcc 139.1K
/nix/store/a3n1vq6fxkpk5jv4wmqa1kpd3jzqhml9-libidn2-2.3.4 350.4K
/nix/store/s5ka5vdlp4izan3nfny194yzqw3y4d1z-lcms2-2.15 445.3K
/nix/store/a5l3w6hiprvsz7c46jv938iij41v57k6-libjpeg-turbo-2.1.5.1 1.6M
/nix/store/r9h133c9m8f6jnlsqzwf89zg9w0w78s8-bash-5.2-p15 1.6M
/nix/store/3dfyf6lyg6rvlslvik5116pnjbv57sn0-libunistring-1.1 1.8M
/nix/store/a3zlvnswi1p8cg7i9w4lpnvaankc7dxx-gcc-12.3.0-lib 7.5M
/nix/store/657b81mfpbdz09m4sk4r9i1c86pm0i8f-app-1.0.0 19.0M
/nix/store/1zy01hjzwvvia6h9dq5xar88v77fgh9x-glibc-2.38-44 28.8M
/nix/store/b1fhkmscb0vff63xl8ypp4nsc7sd96np-openjdk-headless-minimal-jre-21+35 91.4M
There's not much else that can be done here. glibc is the next largest dependency at ~30MB. This large size seems to be because Nixpkgs configures glibc to be built with support for many locales and character encodings. I don't know if it would be possible or practical to split these files out into separate derivations or outputs and make them optional that way. If you're using multiple images built by dockerTools.buildLayeredImage, glibc (and everything else) will be shared across all of them anyway (given you're using roughly the same Nixpkgs commit).https://github.com/max-privatevoid/hackernews-docker-challen...
If you're already using Docker but want to gradually adopt Nix, there is an alternative approach outlined by this talk: https://youtu.be/l17oRkhgqHE. Instead of migrating both the configuration AND container building to Nix straight away, you can keep the Dockerfile to build the nix configuration.
The biggest downside is that you don't take advantage of layers at all, but the upside is that you can gradually adapt your Dockerfiles, and reuse any Docker infrastructure or automation you already use.
But... is a lot of that irreproducibiity (apologies for that word) because there's no guarantee one of the docker layers will be available?
And... does Nix have some guarantee to the end of the universe that package versions will stay in the repository?
But to briefly answer your specific questions: Docker files are commonly not reproducible because they contain arbitary stateful commands like `apt-get update`, `curl`, etc. For a layer with these kinds of commands to be reproducible you would need a mechanism to version and verify the result.
Nix provides such a mechanism, and a community package repository with versioned dependancies between packages. These are defined in a domain specific language called Nix (text files) and kept into a git repository. This should be familiar if you've used a package manager with lock files before.
You can guarentee the package version will stay in the repository by pinning your build to an exact commit hash in the repository.
One issue that really bugged me was to build multi-arch images, it actually wants to execute stuff as the other architecture, and only supports using qemu with hardware virtualization for that. My build machine (and workstation) is a VM, so I don't have that. I do have binfmt-misc, though, so if you just happened to fork and exec the arm64 "mkdir" to "mkdir /tmp", it would have worked. Of course, this implementation is a travesty when docker layers are just tar files, and you can make the directory like this:
echo "tmp uid=0 gid=0 time=0 mode=0755 type=dir" | bsdtar -cf - @-
(As an aside, I'm sure this exact layer already exists somewhere. So users probably don't even have to download it.)Every time I try nix, I feel like it's just a few months away from being something I'd use regularly. nixpkgs has a lot of packages, everything you could ever want. They all install OK onto my workstation. But "I need bash, python, build-essential, and Bazel" doesn't seem like something they're targeting the docker image builder at. I guess people just want to put their go binary in a docker image and ... you don't need nix for that. Pull distroless, stick your application in a tar file, and there's your container. (I personally use `rules_oci` with Bazel... but that's all it does behind the scenes. It just has some smarts about knowing how to build different binaries for different architectures and assembling and image index yaml file to push to your registry.)
You should be able to cross compile binaries for other architectures without actually running them. As long as the package's build files support it of course.
> and only supports using qemu with hardware virtualization for that
That doesn't sound right. You can use qemu for architectures that be only software emulated too.
The minimal example is discussed here:
https://discourse.nixos.org/t/how-do-i-get-a-shell-nix-with-...
I don't want to say it should be as simple as using pkgCross (https://nix.dev/tutorials/cross-compilation.html), but... are some specific issues with the usual process that you're running into?
I never saw went being used so im a bit confused where that came into play for you
> I spent a half a day or so relatively recently trying to build our CI base image with Nix (at the recommendation of our infra team), but it was huge, and some stuff didn't work because of linking issues.
You must be talking about the official Nix Docker image[1], which indeed is huge. I've been using it for years for a handful of projects, but if the size is an issue you can use the method mentioned in the article and build a very minimal image with only the stuff you specify.I think Nix really needs some articles outlining how to play well and smoothly transition from an existing system piece by piece.
and the dx is still pretty bad IMO.
for example, i prefer devbox DX just because i can add pkg like this `devbox add python@3.11`.
also, looking at 120 lines flake.nix. it's not exactly "easier"
https://github.com/Xe/douglas-adams-quotes/blob/main/flake.n...
1. Defines a Go binary (i.e., how to build it)
2. Defines a Docker image that uses said Go binary as an entry point
3. Defines a NixOS module that creates a systemd service that runs the Go binary (only relevant on NixOS)
4. Defines a NixOS test for the module that ensures that the NixOS module actually creates a systemd service that runs the Go binary as expected. The NixOS test framework is actually quite impressive - tests run in a QEMU VM that also runs NixOS :)
Note that only (1) and (2) are relevant to the linked article (+ some of the surrounding boilerplate).
I don't think most people's flakes are more complex than the alternative (which would be, I don't know, a bunch of different script, maybe some ansible playbooks?) but it is a bit daunting when all that complexity is wrangled into a single abstraction.
It doesn't add kitschy terms like "Cell" and "growOn", it's just a way to standardize attribute names according to where things live in the filesystem.
So in their repo, /foo/bar/baz.nix has the attribute path depot.foo.bar.baz
I will say that to understand how it works you need to have a solid grasp of the nix language. But once established I think it's a pattern that noobs could learn in a very quick and superficial way.
I've even used Nix a little bit (though early days, before flakes) and it makes absolutely no sense to me.
Then this std background of "We bring clarity, clear mental picture, manageable Nix projects throughout the lifecycle." appealed a lot.
If the target audience is limited to those already highly experienced (and yet frustrated) with Nix then I guess it's fine.
(PS. This cell breakdown reminded me a bit of Atomic Design for front-end UI to make that easier.)
Because it certainly looks like the things you listed are separate/orthogonal and should be in separate modules/files.
Having many years of Java experience this is the reason why I stick to Maven (not moving to Gradle) - it is opinionated and strongly encourages fine grained modularisation.
Nix absolutely allows that, to the point where I'm surprised that the linked example doesn't separate them. Most of my flakes have a flake.nix that's just a thin wrapper around a block that looks like
devShell = import ./shell.nix { inherit pkgs; };
defaultPackage = import ./default.nix { inherit pkgs; };
...I am aware of flake-parts that are supposed to offer that but the ecosystem of flake-parts is on the smaller side of things...
You can pull in arbitrary code from pretty much anywhere as an input
https://ryantm.github.io/nixpkgs/languages-frameworks/maven/
nix is a build tool a similae way that docker is a build tool. You define build scripts that call the tools you already use. The major difference is that a docker file gives you no way tl be sure you can reproduce that build in the future. A flake on the other hand gives you a high degree of trust.
What we have today is that there is a multitude of dependency managers and - what's worse - all of them are _also_ build tools targeted at specific language.
Nix has a unique position because it is not language specific. Where it is lacking is missing standards and reusable libraries that would simplify common tasks.
I am comparing to Maven because I am looking for multi-platform Maven alternative. There is Bazel but its dependency management is non-existent. There is Buck2 which is great in theory but the lack of ecosystem makes it a non-starter.
Nix is the only contender in the space that offers almost everything and has a chance to become a de-facto standard of software delivery thanks to this.
What's missing though is... easy to use canned solutions similar to maven plugins.
EDIT: grammar
Nix has composable, reusable, and shareable functions. Nix is a full, be awkward, programming language. You'll find functions and flakes for nearly everything you might want to do. An example of one that is more plugin like is sops-nix.
Though have never used maven or maven plugins, I may be missing your overall point.
Want a packaged spring boot application? There is a plugin for that. Want to package it in a container image? Just add a plugin dependency. Want to build an RPM or deb package? Add another two plugin dependencies. All various artifacts are going to be uploaded to a repository and made available as dependencies.
Missing a specific plugin? You can easily implement it in Java as a module and use it inside your project (and expose it as a standalone artifact as well).
I can’t find anything similar in Nix ecosystem.
Having a language allowing for this is not the same as having solutions already available.
For example, `dockerTools` provides different functions like creating an image and other things. There is a module of fetchers, functions that retrieve source files from GitHub and other sources.
But I don't think there are many language-specific functions, like the ones you are describing. I can't think of any, except for the ones that build a Nix package from a language-specific package.
or
https://github.com/nix-community/flakelight
Their aim is to create an ecosystem of reusable Nix libraries. But it is tiny.
Writing a Maven plugin (ie. a reusable piece of configuration management logic) is easy because you can use any library from the vast Java ecosystem (just add a library as a dependency of your plugin).
Doing the same in Nix requires recreating these libraries in Nix.
Looks like Maven might simply be the right choice...
And don't get me wrong - I know it _can_ be done in Nix (as it is a programming language) - the question is _how easy_ it is to create/maintain it.
Let's say my solution requires a Java Spring Boot application, React client, custom Postgres extension and Python ML code.
- Java stuff
- node and npm
- I don't know what you would need for Postgres extensions
- Python, pip or poetry (or whatever you use for Python package management)
Here is an example flake: https://bpa.st/ZUQQ
Under `devShells`, you have `packages`. You can put there any package from nixpkgs, not just Python packages like I have there.
To make it cleaner, you could have a variable for each platform and concatenate them under packages like
`packages = javaPkgs ++ nodePkgs ++ pgPkgs ++ pyPkgs`.Having said that - Java is in general IDE targeted language and once you have an IDE - many small files is not an issue anymore.
This example is a "let's do a trivial thing the way you'd do a big serious thing". If you wanted to treat it as a one-off, I'm sure you could cut it down to 40 lines or so.
Hrmmmm...this is such a TOUGH decision.
cramming everything into one RUN line to save space in a layer.
I really wish instead of:
RUN foo && \
bar && \
bletch
You could do: LAYER
RUN foo
RUN bar
RUN bletch
LAYER
or something similar.maybe even during development you could do:
docker build --ignore-layer .
then at the end: docker build .Docker should address it one day, but until then, we just have to do it in real world scenarios.
RUN foo
RUN bar
RUN baz
and then just build it with something like... I think kaniko did this last I looked?... that smashes the whole thing into a single layer. Obviously that has other tradeoffs (no reuse of layers) but depending on your usecase it can be a good trade to make (why yes, I did cut my teeth in an environment where very few images shared a base layer, why do you ask?).It's not always best or most efficient to squash everything into one giant layer.
[0]: https://docs.docker.com/reference/dockerfile/#shell-form
[1]: https://www.docker.com/blog/introduction-to-heredocs-in-dock...
[2]: https://docs.podman.io/en/latest/markdown/podman-build.1.htm...
That binary + config file are effectively as close as we're going to get to a "flat pack" on linux.
Im not sure what is the forest and where are the trees but this example shows exactly what we have lost sight of.
[1] https://guix.gnu.org/manual/en/html_node/Invoking-guix-pack....
[1] https://nix.dev/tutorials/nixos/building-and-running-docker-...
What you linked is the equivalent of:
FROM scratch
COPY hello-world /hello-world
Of course that's small (hello-world is statically linked). Try to add coreutils (or any other small package) and you'll see what I mean. In my experience the size of a Docker image built with some nix packages is greater than the Debian counterpart. I don't know why though.
However I'll give you that it could be smaller in more complex examples. For example glibcLocales is for all locales which is quite chunky but your application only needs one locale.
Default is none (i.e. like "FROM scratch" in a Dockerfile); you can specify a baseImage if needed, but I haven't had to yet. It works by copying parts of the nix store into the image as needed, but see also below.
> wouldn’t the image size be large if the glibc is copied over again
The original Nix docker-tools buildImage did suffer from poor reuse of common dependencies. Docker already has a way to reuse parts of images (e.g. if you build 7 images where the first N lines of a Dockerfile are the same, the 7 images will use a shared store for the results of running the first N lines). There are several backends for Docker storage that accomplish this in various ways (e.g. FS overlays, tricks with ZFS/btrfs snapshots).
Nix docker-tools now has a "buildLayeredImage" that uses this ability of Docker to share much of the storage for the dependencies, so if you build several images that all rely on glibc, you only pay the cost of storing glibc in docker once.
Instead, it's just a simple root directory placed somewhere and you run the container on that. Much more straightforward.
Write your pipelines and more in languages you already use. It unlocks the advanced features of BuildKit, like the layer graph and advanced caching.
nit: this post makes a point about deterministic builds and then uses "latest" for source image tags, which is not deterministic. I've always appreciated Kelsey's comment that "latest" is not a version
They also recently released their "github actions" replacement <https://news.ycombinator.com/item?id=39550431> but holy hell their documentation is just aggressively bad
Anything in particular you'd like to see improved in the docs?
https://docs.dagger.io/quickstart/562821/hello just emits "Hello, world!" which is fantastic if you're writing a programming language but less helpful if you're trying to replace a CI/CD pipeline. Then, https://docs.dagger.io/quickstart/292472/arguments doubles down on that fallacy by going whole hog into "if you need printf in your pipline, dagger's got your back". The subsequent pages have a lot of english with little concrete examples of what's being shown.
I summarized my complaint in the linked thread as "less cowsay in the examples" but to be honest there are upteen bazillion GitHub Actions out in the world, not the very least of which your GHA pipelines use some https://github.com/dagger/dagger/blob/v0.10.2/.github/workfl... https://github.com/dagger/dagger/blob/v0.10.2/.github/workfl... so demonstrate to a potential user how they'd run any such pipeline in dagger, locally, or in Jenkins, or whatever by leveraging reusable CI functions that setup go or run trivy
Related to that, I was going to say "try incorporating some of the dagger that builds dagger" but while digging up an example, it seems that dagger doesn't make use of the functions yet <https://github.com/dagger/dagger/tree/v0.10.2/ci#readme> which is made worse by the perpetual reference to them as their internal codename of Zenith. So, even if it's not invoked by CI yet, pointing to a WIP PR or branch or something to give folks who have CI/CD problems in their head something concrete to map into how GHA or GitLabCI or Jenkins or something would go a long way
Here's what the very first page of our documentation says (https://docs.dagger.io). I'd love suggestions for making it more clear.
Welcome to Dagger, a programmable tool that lets you replace your software project's artisanal scripts with a modern API and cross-language scripting engine.
[...]
Dagger may be a good fit if you are...
- Your team's "designated devops person", hoping to replace a pile of artisanal scripts with something more powerful.
- A platform engineer writing custom tooling, with the goal of unifying application delivery across organizational silos.
- A cloud-native developer advocate or solutions engineer, looking to demonstrate a complex integration on short notice.
Benefits to development teams:
- Reduce complexity: Even complex builds can be expressed as a few simple functions.
- No more "push and pray": Everything CI can do, your local dev environment can do too.
- Native language benefits: Use the same programming language to develop your application and its delivery tooling.
- Easy onboarding of new developers: If you can build, test and deploy, they can too.
- Caching by default: Dagger caches everything. Expect 2x to 10x speed-ups.
- Cross-team collaboration: Reuse another team's workflows without learning their stack.
Benefits to platform teams:
- Reduce CI lock-in: Dagger functions run on all major CI platforms - no proprietary DSL needed.
- Eliminate bottlenecks: Let application teams write their own functions. Enable standardization by providing them a library of reusable components.
- Save time and money with faster CI runs: CI pipelines that are "Daggerized" typically run 2x to 10x faster, thanks to caching and concurrency. This means developers waste less time waiting for CI, and you spend less money on CI compute.
- Benefit from a viable platform strategy: Development teams need flexibility, and you need control. Dagger gives you a way to reconcile the two, in an incremental way that leverages the stack you already have.
> https://docs.dagger.io/quickstart/562821/hello just emits "Hello, world!" which is fantastic if you're writing a programming language but less helpful if you're trying to replace a CI/CD pipelineWe went back and forth on this. On the one hand, starting with "hello world" makes it plain that Dagger Functions are "just functions", and then gradually introduces more concepts, such as a native `Container` and `Directory` type. On the other hand, you're right that "hello world" is not useful on its own, so you need to read through more pages before a realistic example. Note that this criticism applies equally to all "hello world" examples everywhere.
In any case, this is a valid criticism and I'm tempted to try going straight to a build function as you requested.
> Related to that, I was going to say "try incorporating some of the dagger that builds dagger" but while digging up an example, it seems that dagger doesn't make use of the functions yet <https://github.com/dagger/dagger/tree/v0.10.2/ci#readme> which is made worse by the perpetual reference to them as their internal codename of Zenith. So, even if it's not invoked by CI yet, pointing to a WIP PR or branch or something to give folks who have CI/CD problems in their head something concrete to map into how GHA or GitLabCI or Jenkins or something would go a long way
You are absolutely right, we have started porting over our CI to functions, but have not finished yet.
It's the way building containers ""should"" be.
Here's my base image, here's my files, here's my command, write the into an image.
Of course you can customize the container further if needed.
(Also, you don't need a base image for rules_oci either; most people choose to start with one.)
All this bazel-is-our-savior complex burns down when we want to build a tiny bit complicated thing with it. And we unfortunately do try to do that (our devbox image), which with full caching takes long minutes to an hour, and it's a freaking thousands pine mess instead of using a lined dockerfioe with pinned versions.
I absolutely hate bazel and its broken unmaintained Google-abandoned rulesets and I wish we either used either Docker or Buck2 for everything.
rules_docker was fundamentally flawed -- and overreaching in scope -- in ways that rules_oci is not.
> broken unmaintained Google-abandoned rulesets
rules_oci last commit was 12 hours ago, and it's been actively maintained for years.
(This criticism is even weirder from someone pushing Buck2. Like, it's a great tool. But, apparently it warranted a complete rewrite too, eh?)
> instead of using a lined dockerfioe with pinned versions
You'll pin every transitive version?
> All this bazel-is-our-savior complex burns down when we want to build a tiny bit complicated thing with it.
Let me rephrase. I would not use Bazel to build images from rpm or deb packages.
But....would you use Nix for installing rpm or deb packages????? I believe you've lost the thread.
Unfortunately booting the entire system with /init is largely broken, especially without --privileged. This would be an amazing feature if it didn't require so much extra tinkering.
Docker build are deterministic and easily reproductible, you use a tagged image, that's it, it set in stone.
The 0.01% of Dockerfile that don't work, what does it even means, what does not work?
The other thing is about that buildGoModule module so now you need somehow a third party tool to use or build Go in a Docker image, when using a Dockerfile you just use regular Go commands such as go build and you know exactly what is going on and what args you use to build the binary.
As for the thing about using Ubuntu 18 which is out of date and not finding it, most orgs have docker image cache especially since docker hub closed access to large downloads, but more importantly there is a reason that's its not there anymore, it's not secure to use it, it's like wanting to use the JVM 6, you should not use something that is out of date security wise.
Building an image using nix solves many problems regarding not only reproducible environments that can be tested outside a container but also fully horizontal dependency management where each dependency gets a layer that's not stacked on one another like a typical apt/npm/cargo/pip command. And I don't have to reverse engineer the world just to see what files changed in the filesystem since everything has its place and has a systematic BOM.
And that all relies on discipline. Just like using a dynamically typed programming language can in theory have no type errors at run time, if you are careful enough.
FROM base-image@e70197813aa3b7c86586e6ecbbf0e18d2643dfc8a788aac79e8c906b9e2b0785
RUN pkg install foo=1.2.3 bar=2.3.4
RUN git clone https://some/source.git && cd source && git checkout f8b02f5809843d97553a1df02997a5896ba3c1c6
RUN gcc --reproducible-flags source/foo.c -o foo
but that's (IME) really rare; you're more likely to find `FROM debian:10` (which isn't too likely to change but is not pinned) and `RUN git clone -b v1.2.3 repo.git` (which is probably fixed but could change)...And then there's the Dockerfiles that just `RUN git clone repo.git` and run with whatever happened to be in the latest commit at the moment...
The only reason nix isn't dog slow is that it has really strong caching so it doesn't have to build everything from source.
You can have as many "FROM"'s as you want. "FROM scratch" along with "ADD" is also valid (for non-image dependencies).
From there you do not need to copy things, you can mount the reference into another stage directly.
Also, this is the Dockerfile format, the underlying build API's are _far_ more powerful than what Dockerfile exposes.
A Docker FROM is essentially the equivalent of a dependency in nix... but each RUN only has access to the stuff that comes from the FROM directly above it plus content that has been COPY-ed across (and COPY-ing destroys the ability to share data with the source of the COPY). For Docker to have a similar power to nix at building Docker images, you would need to be be able to union together an arbitrary number of FROM sources to create a composed filesystem.
People use the Dockerfile format because it is accessible. You can still use "docker build" with whatever format you want, or drive it completely via API where you have the full power of the system.
I'm not sure what you mean by 'You can still use "docker build" with whatever format you want'. As far as I'm aware, "docker build" can only build Dockerfiles.
I'm also not sure what you mean when you mention gaining extra abilities to make layered images via the API. As far as I can tell, the only way to make images from the API is to either run Dockerfiles or to freeze a running container's filesystem into an image.
Buildkit operates on "LLB", which would be equivalent to llvm IR. Dockerfile is a frontend. Buildkit has the Dockerfile frontend built in, but you can use your own frontend as well.
If you ever see "syntax=docker/dockerfile:1.6", as an example, this triggers buildkit to fire up a container with that image and uses that as the front end instead of the builtin Dockerfile frontend. Docker doesn't actually care what the format is.
Alternatively, you can access the same frontend api's from a client (which, technically, a frontend is just a client).
Frontends generate LLB which gets sent to the solver to execute.
Applying this information to the topic at hand:
Given what Buildkit actually does, I bet someone could create a compiler that does a decent job transforming nix "derivations", the underlying declarative format that the nix daemon uses to run builds, into these declarative Buildkit protobuf objects and run nix builds on Buildkit instead of the nix daemon. To make this concrete, we would be converting from something that looked like this: https://gist.github.com/clhodapp/5d378e452d1c4993a5e35cd043d.... So basically, run "bash" with those args and environment variables, with those derivations show below already built and their outputs made visible.
Once that exists, it should also be possible to create a frontend that consumes a list of nix "installables" (how you refer to specific concrete packages) and produces an oci image out of the nix package repository, without relying on the nix builder to actually run any of it.
This would subsume the purpose of e.g. https://nixery.dev/
Dockerfiles being at their core a set of instructions for producing a container image could of course be used to make a reproducible image. Although you'd have to be painfully verbose to ensure that you got the exact same output. You would actually likely need 2 files, the first being the build environment that the second actually get built in.
Or you could use Nix that is actually intended to do this and provides the necessary framework for reproducibility.
[1] https://en.wikipedia.org/wiki/Impact_driver#Manual_impact_dr...
This is achieved by building inside of a chroot, with blocked network access etc. Only the dependencies that are explicitly listed in the derivation are available.
Nix is the toolchain, which of course has its advantages.
Both are mostly running well-controlled shell commands and hashing their outputs, while tightly controlling what's visible to what processes in terms of the filesystem. The difference is that nix is just enough better at it that it's practical to rebase the whole ecosystem on top of it (what you refer to as a "toolchain") whereas Docker is slightly too limited to do this.
Using this approach you get something reproducible. Using apt-get in a docker file is an antipattern
We have 2-3 service updates a day from a dozen engineers working asynchronously — and we allow non-critical packages to float their versions. I’d say that successfully applies a fix/security patch/etc far, far more often than it breaks things.
Presumably we’re trying to minimize developer effort or maximize system reliability — and in my experience, having fresh packages does both.
So what’s the harm, precisely?
Regardless.... I can give a few reasons that it matters, off the top of my head:
1) Debugging: It can make debugging more difficult because you can't trace your dependencies back to the source files they came from. To make it concrete, imagine debugging a stack trace but once you trace it into the code for your dependency, the line numbers don't seem to make any sense.
2) Compliance: It's extremely difficult to audit what version of what dependency was running in what environment at what time
3) Update Reliability: If you are depending on mutable Docker tags or floating dependency installation within your Dockerfile, you may be surprised to discover that it's extremely inconsistent when dependency updates actually get picked up, as it is on the whim of Docker caching. Using a system that always does proper pinning makes it more deterministic as to when updates will roll out.
4) Large Version Drift: If you only work on a given project infrequently, you may be surprised to find that the difference between the cached versions of your mutably-referenced dependencies and the actual latest has gotten MUCH bigger than you expected. And there may be no way to make any fixes (even critical bugfixes) while staying on known-working dependencies.
nix provides the toolchain of reproducible artifacts... and then uses that toolchain to build a graph of content addressable inputs in order to produce a content addressable output.
So yes they are very different, but not in the way you are describing. Using nix, just like using docker, cannot guarantee a reproducible output. Reproducible outputs are dependent on inputs. If your inputs change (and inputs can even be a build timestamp you inject into a binary) then so does your output.
But you can achieve the same result if you use similar techniques with a bash script.
Just create a snapshot of the OS repo, so apt/dnf/opkg/ etc will all reproduce the same results.
Make sure _any_ scripts you call don't make web requests. If they do you have the validate the checksums of everything downloaded.
Have no way to be sure that npm/pip/cargo's package build scripts are not actually pulling down arbitrary content at build time.
You seem to be comparing 2 different things.
There are still ways that package will not be fully reproducible, for example if it uses rand() during build, Nix doesn't patch that, but stuff like that is fortunately not common.
It doesn't break as you scale. If you don't need that, then keep using Docker. (Personally, for me "scale" starts at "3 PC's in the home", so I eventually switched all of them to NixOS. I don't have time to babysit these computers.)
> Docker build are deterministic and easily reproductible
No, they definitely aren't. You don't really want to go down this rabbit hole, because at the end you realize Nix is still the simplest and most mature solution.
I build a lot of Docker images using Nix, and while yes it’s generally more pleasant than using Dockerfiles, the 128 layer limit is really annoying and easy to hit when you start building images with Nix. The workaround of grouping store paths makes poor use of storage and bandwidth.
Yes, one of the main downsides of Docker images using Nix is the 128 layer limit. It means we have to use a heuristic to combine packages into the same layer and losing Nix’s package granularity. When building containers with Nix packages already on a Nix binary cache you also have to transform the Nix packages into layer tarballs effectively doubling the storage requirements.
Nix-snapshotter brings native understanding of Nix packages to the container ecosystem so the runtime prepares the container root filesystem directly from the Nix store. This means docker pull == Nix substitution and also at Nix package granularity. It goes a bit further with Kubernetes integration that you can read about in the repo.
Let me know if you have any other questions!
I assume it's mostly in the field of ... "you're making a semi-large investment on this enough that you're doing semi-custom kubernetes deployments with custom containerd?"
Or maybe the thought is nix-snapshotter users are running k8s/kubelet with nixos anyway so its not a big deal to swapout/add containerd config?
For other distributions, nix-snapshotter works with official containerd releases so it’s just a matter of toml configuration and a systemd unit for nix-snapshotter.
We run Kubernetes outside of NixOS, but yes the NixOS modules provided by the nix-snapshotter certainly make it simple.
Can one even get away with abusing DaemonSets/hostDir/privileged on hosted clusters to modify their own installations? Or is `nix-snapshotter` just sort of out of the question on those provided solutions?
I’m not sure what you mean by modifying your own installation. Like running k8s on NixOS and then using a nix-snapshotter based DaemonSet to modify k8s on the host? At first glance it seems like vanilla k8s can do this already, nix-snapshotter just makes it more efficient / binary matching.
Some folks used to use DaemonSets + hostDir to do that node configuration instead of SSH. Which is weird, but less weird than "you can't autoscale nodes anymore because you have a manual bootstrap step".
Or am I just absolutely missing something?/
https://github.com/nlewo/nix2container
It is truly magical for handling large, multi-layered containers. Instead of building the container archives themselves and storing them in the nix store, it builds a JSON manifest that is consumed by a lightly patched version of skopeo that streams the layers directly to either your local container engine or the registry.
This means you never rebuild or reupload a container layer that is unchanged.
Disclosure: I contributed a change in nix2container that allows cheaply pulling non-Nix layers into the build, just using the content hashes from their registry manifests.
nix2container wraps all of that and it automatically runs it in behind the scenes when you call nix run.
The whole image generation is much more efficient as well. In the standard streamDockerImage you get a script that generates a docker layers and the image. In nix2container all layers are stored in nix cache so subsequent run doesn't regenerate them. I believe that was the main goal behind this solution.
Another benefit (and I think it is even better than the caching) is that it also allows you manually specify what dependencies go to each layer.
The automatic way that the code offers, is nice in theory, but because docker has a limit of 128 layers, in practice it starts nicely then the last layer is a clump of remaining dependencies that didn't fit in previous layers.
With nix2container I managed for example to create the first layer which contains just Python and all of its dependencies, then the next layer are my application dependencies (Python packages) the last layer is my application.
With this approach a simple bugfix in the application only replaces the last layer that's the size of a few kb. Only a change of Python version will rebuild the whole thing.
Frequently it's not packaged or some version was but the upstream build system changed an you have to do it yourself.
If you don't either benefit significantly from Nix's upsides (which most projects really don't), or you don't enjoy this and do it for fun, it's not a productive use for your time. If you can convince your employer to pay you to shave that yak, it's a net waste.
[0] https://guix.gnu.org/manual/devel/en/html_node/Contributing....
[1] https://guix.gnu.org/manual/devel/en/html_node/Defining-Pack...
[2] https://guix.gnu.org/manual/devel/en/html_node/Defining-Pack...
[3] https://guix.gnu.org/manual/devel/en/html_node/Package-Trans...
The problem isn't contributions, it's reviews
I've noticed that I get fast merges if the package are popular
It's kind of like seeing something is written in Lisp or Scheme... somebody decided that functional systems and running code aren't worth it. Instead they'll build a system which is perfect in all ways in theory but nobody can or will ever use.
If what you need is available, however, then it can be so much better.
Where Debian/PIP/NPM/Cargo already do work.
Nix as an instruction set, no, you still have to declare what you want, same as a dockerfile.
it's just... Docker will do what a shell script will.
Nix is like hoping someone's done the upstream finagling to get your particular dependencies happy.
That's just possible due to nix being more granular? Is that right?
Can I really build nix from 5 years ago? No src gone? No cache server gone? Nothing?
I mean yeah Ubuntu as a base is shitty.
Would you answer my assumption? Is this easier with nix because it's more granular?
Like if I have a cacerts alone I can select it?
And can I really build stuff from 5 years ago?
Is there a better alternative like Alpine but with glibc that isn't Debian?
Just override the font if you dislike it that much. That's the great thing about the interwebs, it's all just text, you can format it however you want.
Assuming you're using something based off chrome: https://chromewebstore.google.com/detail/font-changer/obgkji...
Or Firefox: https://addons.mozilla.org/en-US/firefox/addon/refont/
What I'm getting at is that they might (unintentionally) be complaining about a rendering bug and not the actual font...
"Just use an extension that changes the website bro!"
As another commenter points out, you may have some rendering issue. Alternatively, you may just not like the font. Can't please everyone.
Using nix, or Dockerfile, or similar systems are, today, fundamentally additional complications to support building containerized systems that are not pure Go or pure Java etc. So we should stop recommending them as the default.