Kata Containers: Virtual Machines that feel and perform like containers
katacontainers.io
katacontainers.io
> Mapping files using DAX provides a number of benefits over more traditional VM file and device mapping mechanisms:
> Mapping as a direct access device allows the guest to directly access the host memory pages (such as via Execute In Place (XIP)), bypassing the guest kernel's page cache. This zero copy provides both time and space optimizations.
> Mapping as a direct access device inside the VM allows pages from the host to be demand loaded using page faults, rather than having to make requests via a virtualized device (causing expensive VM exits/hypercalls), thus providing a speed optimization.
> Utilizing mmap(2)'s MAP_SHARED shared memory option on the host allows the host to efficiently share pages.
From https://github.com/kata-containers/kata-containers/tree/main...
Having gone through an evaluation of Firecracker's security my main conclusion was that sandboxing the processes in the guest is the highest 'bang for your buck' way to reduce escapes.
Here's one option I came across recently that's pretty nice and simple, using virsh and cloud images, and making it easy to spin up a VM and ssh into it. [1]
One use case I'm looking for is to provision system images including ZFS pools/dataset, and not sharing the host kernel module.
[1] https://earlruby.org/2023/02/quickly-create-guest-vms-using-...
Are these still going? They seem a bit dead to me.
Firecracker and firectl themselves are maintained of course, but lower level.
Last time I looked (a few months ago), the documentation was pretty sparse or outdated. A lot of documentation I found stated something like "we broke that in the new version, you can't actually do that right now", like using it with Docker, which I would much prefer over setting up Kubernetes.
I still think it's an absolutely great thing, but the on onboarding could have been a lot less rocky.
FWIW, I eventually kind of DIYed it with QEMU MicroVM and virtiofs - never did anything with it though.
It still is, though it works somewhat seamlessly when installing with https://github.com/kata-containers/kata-containers/blob/main...
Though only one of the hypervisors works well.
Well don't leave us hanging, which one?
(My money's on QEMU)
Kata gives you a few different options for what/how you'd like to boot including firecracker.
This isn't exclusive to firecracker but if you stay lightweight you can have vm's booting under a half second if you're using slim images.
https://jvns.ca/blog/2021/01/23/firecracker--start-a-vm-in-l...
I honestly think for a lot of people, vm's with the convenience/orchestration tools of containers make more sense for a lot of general use cases simply because of the security benefits. The convenience still needs some work though.
Compare that to a docker container where there's basically 0 additional work that has to be done to be up and running.
For most cases I'd be really tempted to work on hardening the docker container than on setting up a VM. Things like Apparmor and seccomp in particular would likely go a very long way.
The big problem with Katacontainers is not whether or not they are slightly faster or slower than containers, but the fixed memory allocation which means you must first know and then allocate the maximum amount of memory they might ever need up front. This can practically limit the number of Katacontainers you can run to something much smaller than is possible with ordinary containers, since RAM is the constrained resource on most servers.
Nevertheless, with confidential computing coming along, it's likely that at some point in the future many containers will really be VMs, since current CPUs implement confidential computing on top of existing VM primitives (and that's basically necessary due to the way the guest RAM is encrypted). It's likely that any workload that touches PII, finance, health, etc will be required to use confidential computing.
Yup, that's always been the big reason to use containers for me. Startup time and runtime performance are nice benefits, but the memory usage is the giant win. Freeing memory in response to the apps need and also not needing extra memory for running the various OS parts and pieces.
The down side is, of course, security. But that was always the case with containers.
We've always had 'compile once, run anywhere' but there's always been caveats and gotchas.
edit: don't shoot the messenger. I was merely highlighting the main difference between native and webasm in the context of the discussion.
I don’t think it will replace Docker files since they let you package up such a wide variety of existing server software and WASM is more limited. But if your software does compile to WASM then maybe you don’t care about that.
I think of WASM more like a plugin format, but I expect there will be a lot of engineering effort put into optimizing it, like happened with V8 for JavaScript. Not all web standards win, but betting against one that’s well-established and has a lot of support seems like a mistake.
For the stuff people run on their Kubernetes clusters I have more mixed expectations. Containers are more universal, but I can totally see a microservice architecture running as a lot of WASM runtimes with a handful of containers.
Conversely the problem with containers is that memory allocation including the OS page cache is not guaranteed. That's bad for a lot of applications, especially databases. It seems Docker has some support for shared page cache but it's not in the Kubernetes pod spec as far as I can see. [0] You would probably need some kind of annotations and a specialized controller to make this work.
https://github.com/kata-containers/kata-containers/blob/d50f... uses virtio-mem
There is still a benefit to ballooning support even if it's not exposed to userspace within the VM, because VMs aren't always used purely to host a single infinitely-long-lived application without outside intervention.
virtualization adds very overhead, a Windows VM running with a dedicated GPU can get 95% of the host's score on 3dmark.
the biggest issue on these cases is IO which can be handled in a few ways.
Missing word?
"Very little"? 5% is enough to turn this year's high-end machine into last year's model.
Nested virtualisation is also a thing now, and 5% per layer adds up fast.
There's usually overhead in the places where the communication requires an additional hop. If you want your host filesystem isolated you're going to need a translation layer and it will be slower. If you're willing to open up your host OS's filesystem, you can basically get ~0 overhead.
I suspect that nearly all containerization software is insecure. Especially with timing attacks like side-channel attacks (Row Hammer, etc):
https://en.wikipedia.org/wiki/Timing_attack
https://en.wikipedia.org/wiki/Side-channel_attack
In the end, the only way to "prove" container security is to be able to point to the fact that nobody has broken out of it yet. It's ..remarkable that our entire cloud infrastructure runs on containers that have never been audited by brute-force in this manner.
Nah, it’s 2023, brute it
This seems, well, naive?
You think there aren't a lot of people that have tried to break out of cloud containers?
Both EC2 and gcloud have had issues over the years with container breakout and leaks.
Complex (ie literally layers of operating systems) software has bugs. Yes. But so does the non-containerised base-case.
We have the best style of honeypots you could ever ask for already running- payment infrastructure on the internet. Go get 'em.
Here's work we did to exploit Firecracker. We had a known, promising vulnerability, and still failed to break out.
https://web.archive.org/web/20220927150915/https://www.grapl...
Exploitation is extremely hard. No one has ever exploited RowHammer in the wild, to my knowledge, and there are a lot of reasons why - but even still, RowHammer isn't magic, you can mitigate against it by increasing your refresh rate, or by limiting the attacker's execution time. Not to mention that vendors have deployed numerous methods of reducing the likelihood of an attack and it is quite complex to pull off these days.
The CPU has a key it uses to decrypt memory on access - and that key is never known by the operating system or any software running within it. If you use RowHammer to access the "wrong" memory location, the decrypted values would be random garbage.
Pure hardware protection, no changes to memory chips, almost no runtime performance cost, effective security.
And I have no idea why you think no one is trying to break them or audit them.
Hey that's an interesting solution you have there, do you think that emulation might help guard against side-channel attacks? Do you have any plans to mitigate them? I wonder if anyone has gamified container security, maybe with a honeypot somewhere where people could try to break out.
So if it’s a dedicated kernel, can this fool game anti cheat systems into thinking it’s not in a VM? Or still the same problem?
Kata VMs are especially VM-y because they use a lot of VM-only features that wouldn't work with real hardware to enhance performance by sharing work between the guest and the host.
The kernel still exists in a VM. The “dedicated kernel” bit of Kata distinguishes it from typical containers, not from VMs.
Kata is an abstraction atop existing hypervisors that are themselves just an abstraction over KVM. There’s nothing new here w.r.t. VM detection evasion.
> fool it into thinking it’s installed in the right space
Most VM detection is about observing devices, drivers, timings, or other side-channel type data that is often only seen in a VM.
https://www.redhat.com/en/blog/red-hat-openshift-sandboxed-c...
Also, AWS has Fargate, which is basically containers-as-VMs. Every ECS task or EKS Pod is launched as a separate EC2 instance under the hood. This obviates a lot of the need for a solution like Kata there.
By contrast, running attacker code in Firecracker is much safer - for one thing, the attacker needs to escalate within the VM in order to then expose the necessary primitives to then escape the VM. So it immediately adds an additional required vuln. But also, instead of the entire Linux kernel being the attack surface (it is for the first vuln, but...) you have to attack a much smaller codebase that implements the VM.
In the case of Firecracker this is really hard and, depending on your exploit, if you end up controlling the Firecracker VM itself you're actually in yet another sandbox - so you now need another vuln to escape (although that sandbox is not amazing imo).
When AWS requires isolation they never use containers.
2. So you can run in Kubernetes and use the Kubernetes API, alongside any non-VM Pods and many other features in your Kubernetes cluster
3. Because traditional hypervisors have a higher overhead cost than the microVMs used by Kata and Firecracker
4. Because traditional hypervisors have a longer startup time (e.g. Firecracker is used by Amazon to run Lambda functions)
Then you just use it with:
docker run --runtime io.containerd.kata.v2 --rm -it hello-world