We clone a running VM in 2 seconds
codesandbox.io
codesandbox.io
One challenge to clone-and-restore that they don't talk about here is making sure that clones don't behave too similarly (like returning the same cryptographic random numbers). We wrote a paper about that a while back (https://arxiv.org/abs/2102.12892), and the Linux kernel community has been doing some great work in that area recently too.
Once uniqueness has been solved though, VM cloning would become a real solution for serverless hosting (and many of other cases), exciting prospect!
Details here: https://bugs.launchpad.net/bugs/1710341
It was since fixed though I never updated the bug.
src/util/virrandom.c:virRandomOnceInit seeds the random number generator using this formula: unsigned int seed = time(NULL) ^ getpid();
This seems to be a popular method after a quick google but it's easy to see how this can be problematic. The time is only in seconds, and during boot of a relatively identical system these numbers are both likely to be relatively similar across multiple systems which is quite likely in cloud-like environments. Secondly, by using bitwise OR only a small difference is created and if the 1st or 2nd MSB of the pid or time are 0 then it would be easy to have colliding values.
Though problematic from basic logic, I also tested this with a small test program trying 67,921 unique combinations of time() and pid() which produced only 5,693 random seeds using PID range 6799-6810 and time() range 1502484340 to 1502489999.
For example, a simple memset() call across a few gigabytes of RAM inside the VM might slow down a factor of 1000x after a VM clone like this.
Both 'parent' and 'child' VM's see the slowdowns.
Some types of garbage collector also see really substantial slowdowns.
It was a deal-breaker for my project where I did similar sorts of things.
If your RAM is file backed, you end up spending lots of time in the filesystem code too - I used anonymous mappings which really helped there, and called clone() on the VM process to keep them shared.
I suspect if you use huge pages you might see lots of the impact vanish, but obviously that has other downsides.
I've been looking at huge pages recently, I'm going to do some more testing with transparent huge pages today and see if it changes performance. Unfortunately we cannot use reserved huge pages because that doesn't work with shared mmap on say an XFS FS.
Another idea is to make clones use the same memory base layer of their parent, then the pages are already prefaulted and it would deduplicate overall memory usage. Many things to discover still..
I think increasingly we'll see Firecracker used with EC2-like setups of "create a disk image with everything preinstalled and then boot it" rather than using snapshots of running (suspended) VMs.
The main reason why snapshotting became interesting for us, is because we're running development servers defined by our users. A development server could take a long time to start, sometimes minutes.
So even if we can start the VM fast, the most important speedup for us is on the user code that we cannot control.
The opposite case - say the user code binds to an IP:port to run a service. Will the clone try to step over the parent, binding to a port that is already taken?
For IP uniqueness, we give every VM the same IP, but we put every VM in its own network namespace. Then we have iptable rules to rewrite the src/dest IP on every packet that enters the network namespace.
I did some extensive IPv4 and IPv6 ECMP anycast testing a couple years ago where we'd randomly bring up and kill hosts and containers.
The network layer provided the fault tolerance and could be tweaked to react very quickly to missing hosts.
I revisited my proof-of-concept test scripts when I wrote the previous comment. I'll try in the next week to add some additional tests in there to determine stream reliability and packet delay/loss.
UDP of course doesn't have the same benefits.
I'm using ECMP + Anycast in a project I've been developing for the last couple of years (K18S or Keep It Simples Stupids) to effectively replace Kubernetes functionality with standard protocols and tooling that is in almost all distros.
We started out with the challenge of replacing the major parts of CNIs and that is where the ECMP + Anycast work arose from.
Native IPv6 with only VLANs and direct routing (no messing about with IPv4, NAT or overlay networks), ECMP + Anycast gives load-balanced routing to pods with automatic detection of lost hosts. Pods exposed to public get public IPv6 address in addition to a ULA (Unique Local Address, formerly called site-local). ULAs used for private routing.
Systemd-networkd is configured automatically by systemd-nspawn so there doesn't need to be a massive, foreign, orchestration control system.
Systemd-nspawn/systemd-machined to manage container lifecycles with OCI compliant images, or leverage nspawn's support for overlayfs to build machine images from several different file-system images. (rather like Docker's layers but always separate, not combined) but can be used in a pick-and-mix fashion to assemble a container that has several related but separately packaged components.
Configs for /etc/ of each container mapped in from external storage using the same overlayfs method. In most cases everything is read-only but some hosts/pods can be allowed to write into the /etc/ overlay and those changes can be optionally committed to the external storage.
Adopting IPV6 and dropping IPv4 was the best thing we ever did in terms of keeping things simple and straightforward and relying on the existing network protocols and layers, instead of re-inventing it all (badly).
At the time we started Kubernetes didn't even have IPv6 support and even once it did many CNIs couldn't handle it properly.
A very specific use case, I know, but if I could have the CI runners run as needed, we could get instances that are way bigger so our builds run faster, and pay around the same amount since they don't have to sit around when they aren't being used.
Pulling the image and building the container is actually just a matter of a few seconds.
I have no data about it though.
On the other hand, ECS still seems slow compared to k8s where things are nearly instance unless you're measuring so ECS control plane speed might be part of the issue, too
If you need super fast boot times firecracker is definitely worth looking at but should be taken with caveats of what precisely you are going to run there.
I have a question about the copy-on-write example involving VM A and VM B. It says t VM B will directly use all the data from VM A and for any change, it copies the block, writes into it and reads from it after this.
But what if, say, block 2 is changed by VM A and was never written to by VM B? Wouldn't VM B read the changed block 2? Clearly, it doesn't happen cause a fork is a copy, but an explanation of how this is tackled is appreciated!
What if VM A is a new VM? What happens to the block after copy-on-write? Just destroyed?
So the logic is to check if VM A has a new fork. If yes, then start CoW to a new layer of blocks, and leave the current layer to be linked with VM B. If no, just don't use CoW.
I hope I got it right!
Practically, 99% of the forks will be done from the `main`/`master` branch of the repo, which is read-only for everyone on the team. So the mini-pause isn't breaking in those cases.
[1] http://www.cs.toronto.edu/~brudno/public/pdf/lagar2009snowfl...
The major difference seems to be that SnowFlock would start a proprietary server which is responsible for sending memory pages over the network on demand whenever the clone reads them. Some follow up work also added several different prefetching strategies to improve the performance of the cloned VMs while they were still fetching remote memory.
SnowFlock was really targeted at compute-heavy applications. The idea was that you could mostly set up your application in the single VM, clone it, and then after cloning, it could be fairly easy to configure the clones to continue working on the problem in parallel.
My Masters thesis made use of SnowFlock to clone relational databases on demand.
https://www.researchgate.net/publication/221351958_FlurryDB_...
Very interesting blog post nevertheless. Looking forward to read more!
[1] https://github.com/qemu/qemu/blob/7dd9d7e0bd29abf590d1ac235c...
https://github.com/firecracker-microvm/firecracker/blob/main...
In our case we changed Firecracker to use a shared mmap instead of an private mmap, so in our case the dirtied pages were synced back automatically to the backing memory file. The main reason for this was to reduce IO on snapshot time. I'm also looking at other ways we can do this, because using a shared mmap fragments the underlying xfs fs pretty fast. Maybe we can batch writes more instead of writing single pages.
Using XFS with CoW has been the easiest way to enable this, but if there's a way that we can do this purely in-memory, that would be even faster.
That said, for hibernation we would still have to persist to disk, but timing is less important there.
My question is how are you guaranteeing uniqueness, or do you only clone snapshots for a single tenant? [3]
[1] https://github.com/self-actuated/actuated [2] https://github.com/firecracker-microvm/firecracker-go-sdk [3] https://github.com/firecracker-microvm/firecracker/blob/main...
For most of us who are consumers only of these more fundamental infrastructure projects, there's something deeply satisfying about seeing people push these boundaries (very appropriate for HN too). Fly is another similar team/blog
Especially if a machine is snapshotted, restored, snapshotted again, restored and the cycle continues. Even if what’s stored doesn’t get much larger the subsequent snapshot+restore processes take a little longer each time. Each provider has different timelines with vultr saying it can take up to 60 minutes for a snapshot to restore.
My use case is similar but different to code sandbox. I use a beefy remote machine for development and to keep costs low I fire it up and tear it down on demand and pay only for the hours the machine was up. It works fine for me but I just wish snapshotting+restoring was faster on these services. That would make it perfect.
Still. Food for thought for me. Thanks again :)
1) Clone your VM in 1.5s as described in the article
2) Clone your database in a few seconds with Database Lab Engine [1]
3) Something else?
looking forward to the unwritten details / future posts too, particularly:
- How to handle network and IP duplicates on cloned VMs
and
- Turning a Dockerfile into a rootfs for the MicroVM (quickly)
Another (unrelated) test we've done is on overprovisioning memory. We were able to run 200 VMs (all running Vite dev server where a file was changed every second) with 2GB RAM per VM, on a node with 128GB RAM. Because we were mapping the memory files on disk directly to the VM, the VM would automatically "swap" the memory back to the memory file when it had memory pressure. The bottleneck here was CPU.
Unrelated TIL: AWS Fargate has supported Windows since last October. I work at AWS and “specialize” in serverless and I didn’t know that.
On the other hand, CodeBuild has supported Windows containers for years and at least CodeBuild for Linux is based on Fargate, so the service team figured something out. (I had to figure out how to word that. I can’t say “they figured it out” since I work for the same company. But I couldn’t say “we” since I’m so far removed from any service team in the consulting department that it would be disingenuous)
The microVM emulates the minimal possible set of devices needed to run, such as disks and network devices, and in the specific case of firecracker, through the use of the virtio model. So it can theoretically use huge amounts of memory of a large vCPU count and still be a microvm.
I'll make sure we write about the other topics as well. For the network, we run the VM in its own network namespace on the host, and we give every VM the same IP. We then use an iptable rule to rewrite every incoming and outgoing packet to the IP that the host has assigned for the VM.
Another use case I was thinking of was stateful compilers like scala where warming up the compiler is expensive, often a CI task too.
Disclaimer: I’m the author.
- you can dump the image using `docker save <name>`. - you can then get a list of the tarballs in this image by extracting this tarball and reading the file `manifest.json`; `Config` -> `Layers` will give you a list of tarballs (see undocker for how to do this: https://github.com/larsks/undocker) - Untar these in a directory and use linux tools to convert this dir to a rootfs.
Problem is the VM takes twice that time to boot so it's not as impressive ;-)
(yes, it's a different idea to the OP but still pretty neat)
Looking forward to read about networking. That I think is technically also interesting and has been a challenge for us for a bit. Coming to the VMs and lower level topics like kernel, or Linux networking has been really fun for me. Weirdly, things feel much simpler the lower you go for some reason. Probably less abstraction?
A bit of self less promo. We are using Firecracker to create interactive onboarding for devs. We did one for Prisma
We start a Firecracker clone when you visit the website. Everything you do happens in your Firecracker VM. You have access to the terminal and can play around with Prisma.
> Initially, Qemu is booted and its state is saved. On each evaluated command, this state is loaded (giving a usable shell in less than one second), a command is fed on stdin and the output read on stdout.
A quick example to take offline, instantaneous "disk snapshots" (QEMU can do this for live VMs too): Let's assume you already have a disk image of a clean Linux distro, let's call it _base.raw_. Then you can create instantaneous "snapshot"[1] this way:
$> qemu-img create -f qcow2 -b ./base.raw -F raw overlay1.qcow2
[The "-F raw" is specifying the file format of the backing file; this is a good practice to explicitly mention this when creating overlay files;]Once you do this, and boot the VM with overlay1.qcow2, all the new guest writes will go to overlay1.qcow2. And whenever the guest need to refer to some old data it is copied over from the backing file, base.raw into the overlay1.qcow2 file. This lets you take a a backup of the base image, or make more "snapshots" (overlays) based on it.
To take an instantaneous disk snapshot while the guest is running, refer to the docs here[2].
[1] The term "snapshot" here a bit of a misnomer, it is actually called an "overlay" — because the overlay file "refers" to its backing file, which becomes read-only once you create the overlay.
If you also have a versioned filesystem you can efficiently create lots of snapshots that store VM images differentially you can introduce a branchable/versioned environments for the whole backend and tie it to the repo commit hash.
Yes, we did look at live migrations since there's a lot written about it and it's the closest to cloning a running VM. Lots of development in that space!
Online I was only able to find that the way to go seems to be to produce an heartbeat on a rolling schedule, wanted to look into this.
In Lambdas, you can schedule a cloudwatch event similar to the heartbeat you've mentioned.
I don't see anything about graphics in the article - could this approach also be used to clone a VM running a desktop window manager like Gnome or KDE? Or would that rely on GPU memory which is not included in the dumps?
Overall great article!