The unreasonable effectiveness of VMs in hacker pedagogy
hiandrewquinn.github.io
hiandrewquinn.github.io
==> default: Clearing any previously set network interfaces...
==> default: Preparing network interfaces based on configuration...
default: Adapter 1: nat
==> default: Forwarding ports...
default: 22 (guest) => 2222 (host) (adapter 1)
==> default: Booting VM...
==> default: Waiting for machine to boot. This may take a few minutes...
default: SSH address: 127.0.0.1:2222
default: SSH username: vagrant
default: SSH auth method: private key
default: Warning: Remote connection disconnect. Retrying...
...
Timed out while waiting for the machine to boot. This means that Vagrant was unable to communicate with the guest machine within the configured ("config.vm.boot_timeout" value) time period.
I cannot find any config.vm.boot_timeout in VagrantfileSo, just set config.vm.boot_timeout to 900 in the vagrant file, and see what happens.
Bringing machine 'default' up with 'virtualbox' provider...
==> default: Checking if box 'debian/bookworm64' version '12.20240212.1' is up to date...
==> default: Machine 'default' has a post `vagrant up` message. This is a message
==> default: from the creator of the Vagrantfile, and not from Vagrant itself:
==> default:
==> default: Vanilla Debian box. See https://app.vagrantup.com/debian for help and bug reports
almost immediately. Looks like timing out was not the real issue. I can run vagrant ssh now.Realizing there's no cc and no git, I try
$ apt-get update
Reading package lists... Done
E: List directory /var/lib/apt/lists/partial is missing. - Acquire (13: Permission denied)
Hmm, still some challenges ahead...sudo apt update
Also, isn’t Vagrant dependent on VirtualBox? Because the latter doesn’t run on ARM Macs.
Added a footnote suggesting a workaround with Parallels for any confused ARM Mac users out there, searching desparately for this fabled Virtualbox program. Appreciate the catch.
> Ubuntu VMs on demand for any workstation
Which may be suitable for many Linux-based tutorials, but not necessarily all. For instance, we deploy software to RHEL (historical reasons more than anything else) at work and might want a RHEL-based tutorial for things unique to administrating it versus administrating Ubuntu or other distributions.
Added benefit was mainly some sort of workspace isolation since you could simply “vagrant in” you different projects via WebStorm/PHPstorm in parallel.
I was given a MacBook Pro 13” with then massive 16gb exactly for this reason: to handle VMs via Vagrant.
Before that, everything was kind of tricky and fragile, mostly relying on a LAMP stack without isolation.
Vagrant was fantastic at the time for the job.
It sure would be nice if Apple provided a nice interface into their emulation for various projects like VirtualBox / Bochs / PCem / Qemu / DOSBOX / VMWare / Bochs / WINE to plug into.
It also kind of sucks that the entire overall body of PC / DOS / Windows / x86 emulation and virtualization is locked in all these silos despite the open source nature. The problem probably is that there are so many gotchas to document and cross-annotate across the projects that it basically is impossible without some dedicated team of very talented technical documenters.
The thing is I don’t really care. I was worried I’d need to run Windows or Linux X86. Turns out, I don’t. I only had to run Windows once for a few moments in the last 3 years since I had an Apple Silicon. Surprising, actually.
Do you re-use an existing artefact or do you build a new one?
You can put Nix in Docker and in Vagrant and get the same result every time. You an also not use Nix, create a Docker image or a Vagrant box, and whenever you instantiate one of those you get the same result every time. You an also involve Nix and not pin to a specific release and now you get a different result every time. That would be the same as creating a Dockerfile with a mutable tag or Vagrant file with a dynamic setup.
Usually you don't really want "the more deterministic way to get all, and only, the things you absolutely need.". That is something that is probably only really useful in CI when you need reproducible builds. For everything else you probably want to be tracking a stable release for the version you target. Especially since most people aren't working to get a deterministic result, but a 'close enough' result to get the job done.
> Usually you don't really want "the more deterministic way to get all, and only, the things you absolutely need."
IMO, you do, as otherwise you'll need to fix additional issues every time your CI runs, no?
>> Usually you don't really want "the more deterministic way to get all, and only, the things you absolutely need."
>IMO, you do, as otherwise you'll need to fix additional issues every time your CI runs, no?
That is why I wrote about the specific exception for CI, but since we're talking about humans experimenting in a box, we're not talking about CI.
No, in nix you pin the commit hash of nixpkgs, which gives you the exact same package set every time. But with a Docerfile, even if you add it to your project's repo then any `apt-get install something` will still give you a different something depending on when you build the image. In nix, you always get the same something.
> That is why I wrote about the specific exception for CI
I meant that you keep the dev environment the same as CI because if you develop in a different environment than the CI then whenever CI runs then you'll get new issues (because you weren't using the same environment).
As for CI, dev environments, not relevant here. I only mentioned in passing because CI is a common use case for hardcoded immutable reference. We're talking about someone wanting a box to experiment in, not CI. Not even software engineering either.
You can both consume that specific one, but also reproduce it locally if you want to. Depending on your intent, bandwidth and trust you can pick the method you desire.
So building the same Dockerfile twice gets you the same hash. This does of course not work if you change something each build, like pulling in some random dependencies that you didn't also pin to a hash, or if you don't use a fixed point in time for the filesystem. We also sign the images so you get some protection against (future) hash collisions. Works for about 900 different images in our internal registries so far. We also used nix, but only a handful of developers actually enjoyed it, so that project got killed off.
You can use any search engine and search for "Dockerfile reproducible build" if you want to learn more.
Docker itself doesn't really get you reproducible builds. I assume you haven't actually tried to achieve it if you still believe it. Docker is like any arbitrary linux machine you start installing things on: you need some other system to get reproducibility.
It's also not a normal workflow to build the same image many times, except for validation, and you do that with frozen sources, not arbitrary references. You build the image, sign it, distribute it. It doesn't get rebuilt on the destination.
This way we can reproduce at will, but also make use of the very large ecosystem of suppliers, consultants and communities that already exist.
Edit: the other resources (source, source image) are also packaged as OCI images and signed and stored so you have everything available in the registry, even if the source repositories were to cease to exist. Because everything else we run interfaces with OCI registries by default we don't have to re-invent that wheel.
If you want to do that, you'd need to keep your artifact around, and I guess you'd also need to back up the docker images you depend on.
I'm passing X to doubt.
This is just odd analysis. Nix is strictly superior to Docker, except the learning curve. You have to really work at it to make any modern Nix setup non-reproducible.
Either way, you are missing the point: the problem here isn't having perfect reproduction. It's about having an environment you can learn in without having to worry about breaking things. That's why the author references vagrant since it's about as low a barrier to entry as disposable environments can be (and even still a barrier to high for some).
EDIT:
Both vagrant and docker are useful tools, it's just that for creating any dev environments nix will be better due to determinism.
And just in case there is a language barrier: box does not refer to a specific technical implementation of a box, it's just a term to denote a border between the user's system (which they don't want to break) and "something else".
Edit: and just in case the word 'someone' trips over a language barrier, I'm not referring to 'anyone' but to the persona (the 'someone') who might want to try something out because they saw something and thought it was cool. Not someone with package manager experience, not a software engineer, not a sysadmin.
We're also not 'building a learning environment'. Learning environment is a proxy for disposable environment, which as a relatively simple concept is already a step too far for the average "how do I become a cool hacker like on TV" case (which is where the unreasonable effectiveness from the title comes in) where someone might take their first steps and wanting to try something out without breaking their current environment.
For popularity, I wouldn't actually mind if nix weren't all that popular, but it may just have the largest number of packages available for install for any package manager.
Perhaps I can assume that you have only used nix personally or at a small company?
Tell me you don't know what you're talking about, without telling me you don't know what you're talking about.
If you can do this, with flakes, or in pure-eval mode, it's a stop-everything bug in Nix.
Also "Nix isn't the box" again betrays how little you know about Nix. I can take a Nix program and package it into a VM or OCI with literally zero effort. Similarly I can transform any NixOS configuration into a VM for literally any platform I can image, similarly trivially.
Like, I literally can type a single command and boot my current machine configuration in a completely disposable VM. I can trivially build any rev of my machine over the past 3 years and get an *IDENTICAL* nearly bit-for-bit replica of that machine from any point in time.
Edit: come to think of it, this is probably why we stopped using nix for new projects. Having to wrap nix a lot to make it work with reality makes it a bit pointless.
Anyway, Nix isn't the box, and it still will never be the box. It doesn't matter that it could be the box, because the people that want the box are not the people that can utilise nix. This Venn diagram just doesn't intersect.
As for your insistence that it's bit for bit guaranteed: you do really think that anyone in the world who is just looking for a box gives a crap? It doesn't matter how deterministic it is, or how nix does it better than anything else, because the entire factor of determinism doesn't matter in this case.
The only way this is true is if you're sending around a list of pacakges, telling users to install them manually, not pinning nixpkgs, not using a flake. Aka, a situation I've literally never seen or heard of. Ever. Worst case a user comes and asks how to make `nix-env -i` portable and the community collectively gasps and discourages the behavior.
I'm not advocating anything. I acknowledged Nix has a learning curve. I have countless issues opened for UX nags. I'm combating outright FUD and ... I don't want to say... at this point.
>As for your insistence that it's bit for bit guaranteed: you do really think that anyone in the world who is just looking for a box gives a crap? It doesn't matter how deterministic it is, or how nix does it better than anything else, because the entire factor of determinism doesn't matter in this case.
Move those goalposts and ignore the other benefits that spoke to the desire use-case, again, sure!
EDIT: Oh brother, a niv user. Never mind, ignore my post, you probably already know it and don't care. Not worth the time for either of us. Implying that using niv is harder, or somehow not superior to using Docker is really a hot take. I am sad I ever entered this thread.
EDIT2: I'll leave this here and stand by it 110% and I'll insist it proves my point:
>If you can do this, with flakes, or in pure-eval mode, it's a stop-everything bug in Nix.
I.e instead of:
Docker run --rm debian:trixie
You'd do
Docker run --rm debian:trixie@sha256:0ee9224bb31d3622e84d15260cf759bf37f08d652cc0411ca7b1f785f87ac19c
The only disadvantage of the digest approach is that you would need to manually resolve the digest that's correct for your processor arch. Using bare tags like "debian:trixie" can resolve to manifest lists (if so configured) that has Docker automatically find the right digest for your arch.
See e.g. https://github.com/reproducible-containers/repro-sources-lis... for how to use them.
Also, at least for Debian there are official images available that use snapshot.debian.org as package repository from the get-go – unfortunately those images are not published on a daily basis yet.
It seems the overall topic if reproducibility is on the Docker team's to do list, though, see https://github.com/docker-library/official-images/issues/160...
Docker in general can't solve reproducibility - it's the package manager within any container that does that.
Anything you're referring to specifically? I think I made it pretty clear that Docker images based on package repo snapshots are not fully "there" yet.
> Docker in general can't solve reproducibility - it's the package manager within any container that does that.
No doubt, but right now the issue is that the package managers have done their part and the Docker images need to catch up.
If you are not uploading images into registries (why?), you can use "docker save" to turn that image into a tarball and then checksum the tarball.
If the people you are writing a tutorial for can be expected to be familiar with Nix, this is probably a much more future-proofed solution. That's still a small camp of people at the moment, but it does include the creator of Vagrant himself, Mitchell Hashimoto. I maintain that a specific Debian version is a better "lingua franca" for a general audience - that may change in coming years, of course.
I have begun writing a new tutorial called "NixOS: Three Big Ideas" to expose people like myself who haven't given Nix the time of day to explore its concepts in about 20 minutes or so.
The learner already has a pretty good idea of where a computer stops. They can pretty much transfer that knowledge to a VM and not have any problems for a long time. Whereas with Docker, what's a filesystem? What's a process namespace?
And, well, I wouldn't do anything in a container that I couldn't afford to leak out into my host system. In a chroot especially, the files are just sitting there, waiting for a newbie to touch them from the wrong context.
Docker is a lot better on Linux, everywhere else you must also spin a VM of sorts. It makes a difference if you’re running multiple instances since you only pay for the virtualized kernel once.
So much life-force is wasted in trying to make Docker on another OS behave like linux.
It is ironically, easier to run Docker, an linux inside something like a Virtual box an enjoy near native experience.
It does appear to be in some kind of beta
https://web.archive.org/web/20240331100215/https://hiandrewq...
Every single developer on my team has exactly the same environment thanks to a single flake.nix file that I wrote. (And it wasn't that hard. Certainly easier than wasting time troubleshooting individual dev's configuration differences. And certainly less frustrating than Docker.)
vagrant plugin install vagrant-share
Seems like a useful tool also!
If I was new enough to not know what Docker is, I am pretty sure I would be a little confused about that. I don't want a "workflow" do I, I thought I was trying to follow a tutorial. Do I need to "manage virtual machine environments?". Do I need to install both Virtual Box and this Vagrant tool?
I think this is a nice idea and useful, but I wouldn't couch it in terms of being some great technique for beginner tutorials.
VirtualBox nowadays is un-trustable. If you end up installing the wrong package some Oracle lawyer will come and ask you for money.
I'd love for Vagrant to make libvirt/qemu a first-class citizen, to be honest.
The prior step is a lossy conversion, though; there are things arbitrary VMs can know about/control, that OCI container images cannot. Like UEFI, or host core affinity, or the layout of the guest’s physical memory, or host disk storage formatting, or nonstandard behavior for paravirtualized devices like NICs.
Ideally the conversion from VM to OCI image would have some sort of standard + portable schema for encoding this type of VM “pragma” info as latent metadata — which container runtimes would ignore, but which a back-conversion into a VM could pick up and translate into hypervisor-specific VM config.
I think, if that was done, then OCI images would indeed become the only portable format you need for any type of container or VM.
That's not as important anymore, I think, since wacky network/storage topoligies are now mostly wacky k8s topologies in stuff like helm charts. But I imagine that old use case isn't completely gone.
A VM can be a type of re-creatable and disposable environment, if you also include an OS and configuration. But if all you need is a CLI context, the same could be said for chroot, LXC, Docker etc. A VM can be an easier mental model or box for people, but no all people are the same.
Ironic
Also, had a long-forgotten setup with vagrant that I later came back to and couldn't start the VM after a couple of upgrades to vagrant & its dependencies. So take that as you will.
First, there's the whole "VM" part. You are constantly reminded of the border between your machine and the VM: having to explicitly start/stop the VM, SSH into it, set up port forwarding, figure out a core count & RAM size which is large enough to be useful yet small enough to work on everyone's machine, having to set up rsync wrappers because the NFS host mount is unusably slow...
Second, it doesn't really deal well with change. When you're working in a team, you want everyone to have an identical environment. Vagrant is great for setting it up, but it doesn't really have a solution to changing it. Want to add a package? Nuke the VMs, start from scratch, and hope nobody runs into install issues.
Docker, especially with docker-compose, provides pretty much the same but better. It was a no-no when Vagrant first started due to poor Windows support, but I understand that WSL has significantly changed that. You'd have to pay me to go back to Vagrant.
What are you talking about? I used to do all that you're mentioning in the Vagrantfile, via code. Granted, ruby code (which doesn't help), but still in code.
> Second, it doesn't really deal well with change. When you're working in a team, you want everyone to have an identical environment. Vagrant is great for setting it up, but it doesn't really have a solution to changing it. Want to add a package? Nuke the VMs, start from scratch, and hope nobody runs into install issues.
Uh, adding the vagrant file to the code repository, maybe? So that maybe you can decide if a package goes into the vm via a code review? Maybe using the provisioning methods, so that changes can both be reviewed and git-tracked.
Like, have you actually read the Vagrant's documentation ? Or are you one of those people that criticize stuff just because the getting started one-page tutorial doesn't solve 100% of their use-case ?
I've used Vagrant in a team in the past, and it feels like either you have no idea what you're talking about or you have used it only briefly without actually reading any documentation.
> Docker, especially with docker-compose, provides pretty much the same but better.
There are some things docker cannot do, and some things vagrant cannot do.
They're completely different pieces of technology that happen to be able to do similar things (and happen to be used for similar purposes).
All this is still kind of true for Docker on any other platform than Linux. There’s still a VM behind the scenes, and you still have to configure and manage it.
However, I personally do avoid that — when I use Docker “on” macOS, I skip the hassle of actually running the Docker daemon locally, by instead having a little headless Linux mini-PC that exposes the Docker control socket over Tailscale to me, and configuring my local Docker client to connect to it. This not only gets `docker run` streaming to a local PTY, but even streams input files for `docker build` to the Docker daemon, entirely transparently. It’s really nice, as I never have to worry about whether Docker Desktop is running, has enough memory, is taking up my local disk space, etc.
Surely there’s a way to do something like this but for vagrant? A Vagrant “thin client” that runs all the VMs on an IaaS service or in-office hypervisor cluster?
> Vagrant is great for setting it up, but it doesn't really have a solution to changing it. Want to add a package? Nuke the VMs, start from scratch, and hope nobody runs into install issues.
Option 1 (centrally managed Vagrant VMs via MDM): Keep the VM’s rootfs volume separate from its /home + /var volume, and push updates to the rootfs volumes via MDM (ideally using a binary diff system like Courgette.)
Option 2: build a chroot under /opt with versions of everything you care about, using e.g. Nix. Then create an apt/RPM package that in turn uses chef/puppet/etc to update this chroot within the VMs. (This is basically what Gitlab does for their self-hosted installable.)
Still not as good as just pulling a new Docker image… but possible.