FreeBSD on Firecracker
usenix.org
usenix.org
The "process sandbox" wars are over. Everybody lost, hypervisors won. That's it. It feels incredibly wasteful after all. Hypervisors don't share mm, scheduler, etc. It's a lot of wasted resources. Google came in with gvisor at the last minute to try to say "no, sandboxes aren't dead. Look at our approach with gvisor". They lost too and are now moving away from it.
I've been wondering about this - are they really?
Now if workloads become less ephemeral and more general purpose, tolerance for startup latency goes up, annd probability of bespoke needs goes up making VM more palatable.
https://cloud.google.com/blog/products/serverless/cloud-run-...
Obviously there might be many reasons for that, but as someone who worked on a similar gvisor tech for another company, it's dead in the water. No security expert or consultant will ever sign off on a process isolation model. Despite of architecture, audits, reviews, etc. There is just too much surface area for anyone to feel comfortable signing off on hostile multi-tenants with process isolation regardless of the sandboxing tech.
Not saying that there are no bugs in hypervisors, but the surface area is so so much smaller.
Disclaimer: I work on this product but wasn't involved in this decision.
Can't help but feel the security concerns are overblown. To support my claim; Well, Google IS using gvisor as part of their GKE sandboxing security..
https://cloud.google.com/blog/products/serverless/cloud-run-...
But at the end of the day, there is a line in the sand around hypervisors vs proc/kernel isolation models. I challenge you to go to a financial or medical institute and tell their CTO "yeah, we have this super bullet proof shared-kernel-inproc isolation model"
The first question you'd get is "Why is this not just part of upstream linux?" Answer that question and realize why you should just use a hypervisor.
Saying gVisor is "ultimately enforced by a normal kernel" is about as misleading & accurate as "KVM is enforced by a normal kernel" -- it is, but it's a very narrow boundary, not the usual syscall ABI.
I won't speak for AWS, but your assumption about what "enterprise grade" cloud vendors do is dead wrong. I know, because I'm working on maintaining one of these systems.
I don't think your beliefs are well founded. AWS's EC2 by default only supoprts shared tenancy, and dedicated instances are a premium service.
We run Triton and SmartOS in production and the linux compatibility works via lx-zones just fine. Only some of the linux-locked software, which usually means docker, needs to go inside a bhyve VM.
But yes, they do virtualize hardware not the kernel. I'm willing to bet you could swap out vanilla containerd with firecracker-containerd for most users and they wouldn't notice a difference given they initialize so fast.
[1]: And for the non-hyperscalers with less tuning, you may be buffering I/O pages both in the guest and the host.
Thanks to KVM, and to the minimal hardware support (no PCI, no ACPI, etc), Firecracker's source is rather simple and even relatively readable for non-experts.
... as long as they're experienced at writing Rust. As a Rust newbie it took me a long time to figure out simple things like "where is foo implemented", due to the twisty maze of crates and uses directives.
I totally get why this code is written in Rust, but it would have made my life much easier if it were written in C. ;-)
Out of curiosity, what development setup do you use?
I imagine that with vanilla EMacs or vanilla Vim you’d have to do quite a bit of spelunking to answer that sort of question.
With a full-blown IDE like for example JetBrains CLion with Rust plug-in installed, it is most of the time a matter of right-click -> go to definition / go to type declaration. (Although heavy use of Rust macros can in some cases confuse the system and make it unable to resolve definitions/declarations.)
And with JetBrains CLion you still have Vim keybindings available as a plug-in.
I switched from Vim to CLion + plug-ins years ago and haven’t looked back since. (Vanilla Vim is still on my servers though so that when I ssh in and want to edit some config files or whatever I can do so in the terminal.)
Whereas in FreeBSD I just grep for ^foo and the one and only line returned is where foo is implemented -- because if there's different versions of foo, they have different names.
Namespaces sound good in principle but they impose a mental load of "developers need to know what namespace they're currently in" -- which is fine for the original developer but much harder for someone jumping into the code for the first time.
Eg, (me picking a random crate on crates.io): https://docs.rs/syn/2.0.29/syn/ or the standard library: https://doc.rust-lang.org/std/option/index.html
It's all generated by the same system from comments in the source. You can generate the same thing for your code.
I do think it's a deliberate tradeoff, having e.g. .push() do something useful for quite a few similar (Vec-like) data structures means you can often refactor Rust code to a similar data structure by changing one line... but it certainly doesn't make things as grep-friendly as C.
The "hardest" part is probably sufficiently emulating Linux userspace accurately: it's a big surface area. That's why I think creating a pseudo-OS target is the best route.
gvisor is not like a seccomp-bpf process sandbox that just ACLs system calls.
At any rate: why is this better than just using KVM and Firecracker? The big problem with gvisor is that the emulation you're talking about has pretty tough overhead.
Regardless, I'm still surprised microkernels aren't more popular in this space, but perhaps the losing the ecosystem of Linux libs/applications is a non-starter.
Even if the idea wasn't fruitful, the conversation was fun. Thanks for engaging and challenging my bad ideas!
Edit: I've also realized I was thinking of Unikernels, not microkernels and I've been calling it the wrong thing all night. *sigh*
Could you link to any specific Linux kernel source that implements support for TCP offload? AFAIK networking subsystem maintainers were always opposed to accommodate TCP offload because it is A Bad Idea.
[0]: https://www.linuxjournal.com/content/userspace-networking-dp...
https://www.kernel.org/doc/Documentation/networking/segmenta...
You wouldn't have to. There's patches for hardware TCP offload using the normal socket syscall ABI. The kernel net stack maintainer is pretty ideologically against them so they're not mainlined, but plenty of people have been running them in production for decades and they're quite mature.
The closest things to what you're describing are unikernels and NetBSD rump kernels.
1. Mechanism to capture syscalls (systrap)
2. Reimplementation of parts Linux kernel
3. Narrow set of calls to outside using Linux syscalls
This thing could be envisioned as
1. Linux syscall ABI as-is, same mechanism[a]
2. Reimplementation of parts Linux kernel
3. virt-io hardware drivers for calls to outside
So the middle part of the sandwich could look the same.
Also, I think it's worth saying that I think the work in maintaining #2 there is exactly why Google Cloud Run migrated away from gVisor. People just kept asking for more and more kernel features.
[a]: Alternatively, #1 could be replaced with unikernel like linking directly with #2.
Me personally, I think HTTP/3's move away from TCP could be really interesting for this sort of stuff. The responsibilities of the kernel could be hugely simplified if all you had were UDP/IP directly hardcoded to virtio (no need for routing, address configuration, ARP, etc), no paging etc, and the only filesystems were EROFS & tmpfs. Of course, Cloud Run's move away from gVisor shows that Enterprise clients would hate it.
Beside Firecracker, there're all sorts of micro-VM being developed right now, such as crosvm, cloud-hypervisor, Kata's Dragonball, all on top of KVM.
https://qemu.readthedocs.io/en/latest/system/i386/microvm.ht...
The "print a message telling the user that we're rebooting, then wait a second to let them read the console before we go ahead and reboot", on the other hand...
If anyone in the OpenBSD world is interested in speeding up your boot process I'd be happy to share tips and code. It's a bit daunting to start with but with some good tools it becomes a lot of fun.
I understand the reasons for no "how-tos". But sometimes they make sense for people like me. I wouldn't mind delving a bit deeper given some direction.
I use FreeBSD for everything from my colocated servers, to my own PC. By no means am I developer; seasoned Unix Admin at best. Bare-metal forever but welcome to the future. Especially anything that contributes to the OS.
However I hear buzz words like Lambda and Firecracker and really have no idea where the usage is. I get docker, containers, barely understand k8s but why do you need to spin up a VM only to tear it down compared to where you could just spin up a VM and use it when you really need to. Always there, always when.
Is it purely a cloud experience, cost saving exercise?
and always charging you :)
Allows you to build a compute plane where any node in the plane can service the traffic for any application.
Any one application can dynamically grow to consume the available free compute of the plane as needed in response to changes in traffic patterns.
Applications use no resources when they aren't handling traffic.
Growing the capacity of the compute plane means bringing more nodes online.
Can't come up with a use case for this beyond managing many large-scale deployments. If you aren't working "at scale" this is something that would sit below a vendor boundary for you.
As it turns out, a lot of APIs for phone apps fit this category. You don't want a machine sitting around idle 99% of the time to answer those API calls.
If you made a iPhone game and wanted the high scores to sync to a global scoreboard when the game is over, how would you build that?
What if you only expected 10 players a day?
Firecracker has a much amaller overhead compared to regular VMs - which makes the (time and compute) costs of spinning up new VMs really low. This can be an advantage, depending on how chunky your workloads are - the less chunky they are - the more they can take advantage of finer-grained scaling.
Sometimes you just want to slap some lines pf code together and run them from time to time, and don’t need a whole server (physical or virtual) for that.
Sometimes you have no idea if you’ll have to run a piece of code 100 times a day or 10’000’000 times a day.
Sometimes you don’t feel like paying a whole month for and maintaining a whole instance for a cronjob that lasts 20 seconds, and maybe it runs once a week.
Actually you can do virtualisation on any instance type afaik, but only with .metal instances you can use hardware acceleration.
A huge thank you to Colin Percival for sharing this.
Particularly love the "Once the low-hanging fruit was out of the way" line... which to Colin means custom bus_dma patch(es).
Now anyone can now enjoy for free:
"with 1 CPU and 128 MB of RAM, the FreeBSD kernel can boot in under 20 ms"
If you're used to devops with k8s clusters or lots of docker, this is absolutely amazing.
Oh, my... how could I achieve the same on real hardware without VMs ? ;)
It's everything else that's slow. For example, this is my machine
Startup finished in 14.552s (firmware) + 2.885s (loader) + 741ms (kernel) + 23.116s (initrd) + 11.191s (userspace) = 52.488s
I've had libvirt bog standard qemu-kvm (not a microvm) creating a new Ubuntu VM from a disk image & booting to a login prompt in under 10 seconds for more than a decade. This is without fiddling with virtio, doing hardware scans for PCI, VGA, SATA and such, and booting via Grub (your "loader"). Those should be pretty comparable!