CRIU, a project to implement checkpoint/restore functionality for Linux
criu.org
criu.org
[1]: https://github.com/qawolf/crik
[2]: The Party Must Go On - Resume Pods After Spot Instance Shutdown, https://kccnceu2024.sched.com/event/1YeP3
Doing this on a networked application is going to be iffy. The restored program sees a time jump. The world in which it lives sees a replay of things the restore program already did once, if restore is from a checkpoint before a later crash.
If you just want to migrate jobs within a cluster, there's Xen.
Does crik guarantee the order of events (saving a checkpoint should be followed by killing the old process/pod, which should be followed by a restoration - the order of these 3 events is strict) and given that criu can checkpoint and restore sockets state correctly - how does that work for kubernetes? The new pod will have a different IP.
Usually clients would connect to a Kubernetes svc to not have the problem with changing IPs. Even for just a single pod I would do that.
The project ended up being a dead end because it turned out running my program in a QEMU whole system vm and then fork()ING QEMU worked faster.
It lets you make good use of CPU time whilst waiting for the network.
It turns out simple heuristics can get 99% accuracy on the question of 'will the server return the same result as last time for this non-cachable response'.
However, since my machine has many CPU cores it made sense to have many 'speculative' copies of the browser going at once.
A regular fork() call would have worked, if not for the fact chromium is multi thread and multi process, and it's next to impossible to fork multiple processes as a group.
If you are guessing at the data anyway, what's the difference?
Why set up an entire speculative execution engine / runtime snapshot rollback framework when it sounds like adding heuristic decision caching would solve this problem?
It'd be cool to predict which resources are speculation safe (ie the cache headers don't permit it, but the content in practice doesn't change) and speculate those resources but not ones which you have repeatedly had a speculation abort (ie actual dynamic resources). If your predictor gets a high enough hit rate, you could probably do okay with just a single instance/no snapshot and use an expensive rollback mechanism (reload the whole page non-speculatively?).
Basically, for the fuzzing purposes speed is paramount so they made some changes to speed up snapshot restoring. Don't know the limitations but since it is used to fuzz full operating systems, there should not be many.
I believe it should be faster than forking because why even patch QEMU otherwise.
A two-digit number of these devices was meant to form a cluster of sorts and using Erlang clustering sounds like a nice solution for that until you realize that the base load of the full mesh that this implies is high enough to use a meaningful chunk of the device's resources.
That embedded constraint is fascinating. With the benefit of 20/20 hindsight, given the same substrate constraint, what direction would you have taken? I was thinking along the lines of maybe a C/C++ state machine library.
Since everyone is treating containers as cattle CRIU doesn't seem to get much attention, and might be why a video and not a blog post was my first introduction.
Yeah, I guess that's probably the reason. If you're engineering your workloads with the idea that the world might "poof" out from under you at any moment you'd never wonder about / reach for something like CRIU.
It's a trick that I'd never much thought about, but now that I've learned it exists (so many years late) I find myself wondering about the path not taken here. It feels like it should be incredibly useful... but I can't figure out exactly what I'd want to do with it myself.
Check out mainframes and Tandem systems for a peek at that path. Lots of support in those systems for the notion your application’s substrate might suddenly go poof, and you need it to recover from where it left off as instantaneously as possible.
It’s expensive.
Nah, it's more like "I don't trust that thing to not cause weird behavior in production".
VM-level snapshots are standard practice[1] because the abstraction there is right-sized for being able to do that reliably. CRIU isn't, because it's trying to solve a much harder problem.
[1]: And even there, beware cloning running memory state, you can get weird interactions from two identical parties trying to talk to the same 3rd service, separated by time. Cloning disk snapshots is much safer, and even there you can screw up because of duplicate machine IDs, crypto keys, nonces, etc.
Im sure there are some niche applications for container checkpointing, but I don’t really see the complexity being worth it. Maybe checkpointing some long running batch jobs could save you some money, but you should just make your jobs checkpoint their state to an external store such a ceph or s3 and make the jobs smart enough to load any state from those stores if they are preempted.
Hopefully though, my trepidation is wrong. What is the most complex piece of software others have run under CRIU in production, and for how long?
small nit: podman is not a docker fork, it's a completely different codebase written from scratch
In fact one of our customer's use cases is exactly what you describe, allowing users to "hibernate" container workspaces.
Just a few months ago I was talking to a startup founder at KubeCon who built a product based on CRIU. Unfortunately I forgot the company's name. (And I can't find that git repo with the prototype anywhere, even in my backups. Sad.)
Unfortunately, I was disappointed to find `lxd stop --stateful` couldn't save any of my LXD containers. There was always some error or other. This is how I learned about CRIU, as it was due to limitations of CRIU when used with the sorts of things running in LXD.
# lxc stop --stateful test
(00.121636) Error (criu/namespaces.c:423): Can't dump nested uts namespace for 2685261
(00.121645) Error (criu/namespaces.c:682): Can't make utsns id
(00.150794) Error (criu/util.c:631): exited, status=1
(00.190680) Error (criu/util.c:631): exited, status=1
(00.191997) Error (criu/cr-dump.c:1768): Dumping FAILED.
Error: snapshot dump failed
LXD is generally used with "distro-like" containers, like running a small Debian or Ubuntu distro, rather than single-application containers as are used with Docker.It turns out CRIU can't save the state of those types of containers, so in practice `lxd stop --stateful` never worked for me.
I'd have to switch to VMs if I want their state saved across host reboots, but those don't have other behaviours regarding host-guest filesystem sharing that I needed.
In practice this meant I had to live with never rebooting the host. Thankfully Linux just keeps on working for years without a reboot :-)
I could be wrong, though. Interesting approach if true
Except I would strongly suggest not doing that as there have been some very nasty security issues fixed as of late.
Found a GitHub issue for this: https://github.com/checkpoint-restore/criu/issues/1430
The issue apparently is newer systemd versions create their own UTS namespace, so suddenly running systemd in a container results in nested UTS namespace. Containers with older versions of systemd, or which don't use systemd, shouldn't have the issue.
One commenter posted in April 2021 that they had a patch to add support for nested UTS namespaces, but they don't appear to have submitted it: https://github.com/checkpoint-restore/criu/issues/1430#issue...
Comment on another issue has suggestion on how to implement nested UTS namespace support: https://github.com/checkpoint-restore/criu/issues/1011#issue...
It doesn't sound like nested UTS namespace support is impossible, just something nobody has got around to implementing.
Comment in CRIU source code says nested namespaces are only supported for mount namespaces (CLONE_NEWNS) and network namespaces (CLONE_NEWNET): https://github.com/checkpoint-restore/criu/blob/b5e2025765b9...
But if you look at the OpenVZ fork of CRIU, you see it also supports PID (CLONE_NEWPID), UTS (CLONE_NEWUTS) and IPC (CLONE_NEWIPC) namespaces: https://bitbucket.org/openvz/criu.ovz/src/d9bf55896015a27df9...
I don't know why these additional features in OpenVZ CRIU don't exist in the upstream.
I think the main blocker to supporting nesting of the other namespace types (user, cgroup, time), is someone getting around to write the code for the support. It is possible some of them pose some kind of architectural issue where some kernel enhancement might be necessary (if that's true of any, I'd say most likely of user), but I suspect for most of them it is simply a matter that nobody has gotten around to it.
The other issue is eventually someone will add another namespace type to the Linux kernel, and then CRIU will need to support that too.
> One of the CRIU features is the ability to save and restore state of a TCP socket without breaking the connection. This functionality is considered to be useful by itself, and we have it available as the libsoccr library.
If you want your original process to continue living after the checkpoint and not lose packets during checkpoint, you can go a pretty long way with the 'plug' tc, IFBs. And if you're aventurous, lots of support for getsockopt/setsockopt and ioctls have been or are being merged within io_uring so checkpointing a big-buffered TCP socket can cost under 100us, even less IIRC.
Back then, CRIU turned out to not be an option for us. E.g. one of the problems was that it was not possible to be used as non-root (https://github.com/checkpoint-restore/criu/pull/1930). I see that this PR was merged now, so maybe this works now? Not sure if there are other issues.
We also considered DMTCP (https://github.com/dmtcp/dmtcp/) as another alternative to CRIU, but that had other issues (I don't remember).
The solution I ended up was to implement a fork server. Some server proc starts initially and only preloads the modules and maybe other things, and then waits. Once I want to execute some script, I can fork from the server and use this forked process right away. I used similar logic as in reptyr (https://github.com/nelhage/reptyr) to redirect the PTY. This worked quite well.
Getting the code cleaned up enough to post it has been on my to-do list for quite some time, and this has inspired me to do it soon!
I wonder if that "trick" can be extended to a full implementation of distributed shared memory, i.e. multiple nodes running separate tasks in a single address space and implementing cache coherence over the network. Probably needs quite a bit of extra compiler/runtime support so it wouldn't really apply to standard binaries, but it might still be useful nonetheless.
I haven't yet figured out how to (neatly) persist this to disk, so I just sort of make a mini server-loop that catches signals and dispatches a fork() for each one, and that's my fast version of the CLI command. Delightfully ugly :)
(The killer app I'm trying to apply this to is LaTeX, so that I can write math notes in Emacs, incrementally, without visible latency. Unfortunately the running LaTeX process is slightly convoluted, and needs a few more tricks to get working in this way. This trick works on the plain TeX command out-of-the-box (it's like a 50x speedup), so I think I'm on the right track...)
See texpresso [1] for one solution that does something like this with the LaTeX process.
Another, more conservative solution is the upcoming changes to Org mode's LaTeX previews [2] which can preview live as you type, with no Emacs input lag (Demos [3,4]).
[1] https://github.com/let-def/texpresso
[2] https://abode.karthinks.com/org-latex-preview/
(Are you by chance the author of org-latex-preview, or is it a coincidence of usernames?)
It's useful because, by design, it's difficult for the process to even notice it's been stopped. And while it's stopped, you can apply arbitrary patches completely atomically.
And Docker is a very convenient way to do this, e.g. workaround the PID limitation.
(Though I really wish it got more attention https://github.com/docker/cli/issues/4245 )
For long running containerised simulations, this saves a lot of time on failures ( as long as you have a safe place to write the snapshots to ) by not restarting from 0 every time.