The Magic Nix Cache, a GitHub Action for speeding up your Nix workflows
determinate.systems
determinate.systems
You can get fast startups if you're willing to define your containers upfront (dockerTools, nix2container) and/or adopt a dynamic container-server (nixery, flakehub). And you can get reasonably fast substitutions if your binary cache is on MinIO in the same cluster as the workers.
But I feel like there's still room for a "magic" /nix/store that skips the copying and decompression stage altogether— something that works using standard nix invocations (like Magic Nix Cache), but presents itself as a Kubernetes Volume, so that in cases where a path already exists on-node, the existing files (in the cache pod) are simply mounted/served directly into whatever container ran a nix command.
I don't feel like I really know enough about either k8s or nix to assess the practicality of such a thing, but the thought of lightning-fast substitutions for arbitrary Nix workflows is massively appealing.
Several years ago I looked at implementing a custom component for k8s which would exchange nix store paths instead of containers, substitute, and bind mount them in at run time. It was an interesting experiment, but was Quite Difficult to pull off for someone who wasn't already familiar with k8s.
I've seen some projects similar to what you're describing though: the lightning-fast substitutions. It was incredible! They had the benefit of a fabulously fat network connection, though, and I'm not sure the experience translates very well. We will see!
Definitely there's an obviousness to the concept of "magic" Nix stores in various spaces, and I know the tvix project seeks to realize some of this as well— reusing OCI tools to supply the build sandbox, leveraging existing container orchestrators for job management and queuing, all those goodies. So I'm excited to keep watching the space.
Mostly a digression since it's not Nix, but I've wondered a bit about this sort of thing when building AUR packages on Arch Linux. The very last step of the builder is to compress the package, which in the vast majority of my uses is followed immediately by installing it, which of course decompresses the package. I've wondered why there isn't some (non-default) option to say "I don't need to keep the package itself around; just install it as soon as its built". I'm sure for my specific use case there's a simpler solution, but I've always wondered if there's a hacky way to get around things more generally by making a tool that can mimic the expected compression API but then creates a "fake" compressed artifact that no-ops (or maybe puts in a valid header followed by non-compressed data) and then injects that implementation into the PATH. You'd be able to invoke it with something like `fakecompress --zstd makepkg -si`, and it would invoke `makepkg -si` with the no-op zstd implementation.
https://wiki.archlinux.org/title/makepkg
Tar would also be the answer to how to do fake compression (keeping files together and in order in one file) IF you can choose the "compression" library/tool.
Note that when PRs merge to the default branch, their cache doesn't come with them. This is how GitHub Action's cache works, as a security measure. However: subsequent rebuilds will, and PRs off the default branch will too.
I basically expose the daemon socket into the docker container so that it requests builds from the host. It means that everything is cached right on local disk. If you need more oomph than one machine will provide the cache won't be shared between different machines without extra effort but you can do a lot of building on a single machine (especially if a lot of stuff is using Nix so cached).
It caches on two levels (instance's /nix/store on EBS and then also binary cache on S3).
Of course the attack surface is quite large, so I wouldn't expose this to the public. But using this for my repos and trusted developers is fine with me. It is basically impossible to accidentally do harm. Also note that GitLab forks and Merge Requests from forks run in the author's repo, so they won't use your runners. So it is only people with push access that will use them.
More concretely, let's say you have a python backend that uses poetry. Do you just use `poetry install` in your derivation for python-deps? Do you use something like poetry2nix or node2nix and do all of your package management in nix?
Do you have some examples for this? The ones that come to mind are things like bare metal/vms with nix, or perhaps disnix, but those are a pretty hard sell over more popular orchestration systems, and I’d like to have more alternatives.
- getting dev environment on CI to be identical to user dev - with minor changes the project is not rebuilt or or rebuilt minimally - the caching works across branches, so for example merging a feature branch to master, if nothing changes the build on master will be very quick
I created something similar to nix-cache for gitlab, but I had to create a dedicated runner running NixOS.
If I could use NixOS for deployment, at that point I would just point the same binary cache to the machine and use the same derivation to build the app. Because the app was already build by CI, it would just download the compiled version. No need for artifactory or similar. In that scenario (you using poetry) you probably would just use poetry2nix to generate the application.
If the OS is not NixOS, but you still want to deploy via nix, then IMO this[2] looks interesting, basically it packages everything in self extracting archive. That you can extract and then run the app.
Other alternatives are these bundlers[3], which includes building toArx (works in a way similar to the previous one but pretends everything is in a single file), RPM, DEB, docker (you would have more control over it if you would use the code directly instead of a bundler though)
And the last option (probably the most obvious one) is that you can simply just use the tool to build the package. Since you're using poetry, then you can generate a wheel from it.
[1] https://github.com/takeda/nix-cde/blob/master/contrib/gitlab...
Another option might be to use pnpm instead of Yarn and cache your pnpm dependencies. pnpm actually works a bit like Nix in that it creates a pnpm-lock.yaml file with content-based hashes for the full package.json dependency tree. This enables it to quickly determine which parts of the dependency tree it needs to build and which are already available.
[1]: https://github.com/NixOS/nixpkgs/blob/master/doc/languages-f...
Node packages sometimes pull additional files from the internet in a postinstall script, or do other funky stuff that's incompatible with Nix. So the idea that you can construct a pure derivation from a package-lock.json or yarn.lock file is a pipe dream.
That means that node_modules (as created by a package-lock.json) can't really be cached or be built from caches since it depends on the particular version solution found by npm for a particular project.
So there's only so much that Nix can do. It can cache it about as well as using a naive caching scheme with actions/upload-artifact or similar (create a tarball of your node_modules and just cache it across runs, update when you need to).
Basically node_modules is inherently large. If you want better performance for caching dependencies use a better programming language environment.
- uses: actions/cache@v3
with:
# npm cache files are stored in `~/.npm` on Linux/macOS
path: ~/.npm
key: ${{ runner.os }}-${{ hashFiles('**/package-lock.json') }}https://github-runners.www-cachix-org.pages.dev/github-runne...
prs from forks run only main workflows, and these run nix which is kind of isolated enough.
i guess one could attack with some infinite nix store bomb.
But you can use it for things like:
- declare your application with all dependencies explicitly, so when someone else wants to build it they can (I would argue this is the primary purpose and rest is just built on top of that)
- common dev environment (so other developers can get the same dev environment as you with all exact same build tools)
- build toolchain (for example if you do embedded environment)
- you could use it as a replacement for homebrew/mac ports
- a configuration file holding your .dot files (home manager)
- if you use NixOS (OS that was built around Nix) then you have OS with a built-in configuration management (i.e. salt/puppet/ansible/chef) that is truly declarative
I think this[1] also shows some crazy stuff you can do with it.
Regarding question around Docker, the great thing is that if you define your application as a nix derivation, then you can easily generate docker that just contains your application and dependencies (the docker in that case is just a deployment unit). The reproducible environment is what docker promised, but practically failed to deliver. Instead of deliver reproducibility, it actually delivered repeatability.
Yes, actually I got interested in Nix because with requirements.txt I could only define python dependencies, and I had no control over for example installing postgresql C library that psycopg2 depends on.
> Then you pair it with a containerization technology like Docker or Podman
You don't have to, but you can, given that in most places containers are being used a lot of people use nix that way.
> And if you use NixOS, you can skip the last part?
Yes, although keep in mind that for example the requirement is to use Kubernetes then you would have to use Kubernetes. But if you need to create for example an EC2 instance. You can use Ubuntu + ansible or you could use NixOS.
Honestly I don't have much experience with NixOS as at my workplace we are mandated to use specific distro for everything.
Edit: I forgot to add additional benefits with using NixOS compared to Ubuntu + ansible for example. All updates are atomic. You either end up with the new configuration or the old configuration, there's no in-between as it would happen with ansible. Second big benefit is easy way to rollback.
Docker mainly does two things: there's Docker images as a way of sharing some container-image, and the container runtime to allow running those images for tasks or services in isolated/fresh ways.
Docker images make it easy to distribute software which runs the same everywhere.
Nix is a package manager which tackles that problem, but without using container images: Nix is for distributing software so it has the same behaviour everywhere.
Nix users are often very enthusiastic about Nix because it also enables all sorts of neat developer experiences. e.g. Nix is great for setting up development environments.
In terms of Nix-vs-Docker, Nix is also capable of building Docker/OCI images. So, you could use Nix instead of writing a Dockerfile.
Using a dockerfile with a step like "RUN nix ..."? You can, but it strikes me as a cumbersome way of doing things. Mitchell H describes this way of doing things here: https://mitchellh.com/writing/nix-with-dockerfiles
Whereas, I'd reckon the more idiomatic thing to do is to build the Docker image with Nix code. -- You're going to get a precisely defined environment to share across workstation/CI/etc. Some example Nix code for building Docker images is here https://github.com/NixOS/nixpkgs/blob/master/pkgs/build-supp...
so locally i use process-compose to avoiding waiting docker builds.
I suppose maybe it will only work if I split my build up into multiple steps such that Nix will know to skip those first steps. If Nix knows that, I suppose the Magic Nix Cache also knows?
As for splitting build into steps. I am not entirely sure what are you trying to do, but you should not need to do that. Each package in Nix has a hash, which is generated from things like hash of the source code, of the dependencies, compile flags, system architecture etc. This means that you could have multiple versions of the same package in /nix/store with different compile options or different dependencies. When you need a given dependency you know all the information to generate the hash and can easily know if you need to build it or you can use cached version. This is what makes binary cache a pretty much plug and play and don't need to worry what files to cache or whether you should split build into stages.
The long and short of it is a merkle tree of hashed inputs :).
we're at https://less.build if you want to take a peek -- we will look at adding S3 support! :)
that's where we are beginning and then we're working to add spicier features on from there. but first and foremost, we want to serve the ccache/gradle cache crowd with a fantastic protocol-agnostic backend which "just works" and pays for itself in terms of time saved.
and thanks for asking :) we are very new