[1] https://unit42.paloaltonetworks.com/breaking-docker-via-runc...
People need to stop looking at containers as a cheap way to get security. They might be a more convenient way to get lots of apps running on a single machine, but they're not very secure.
2. It sounds like this paper is mainly about covert channels not side channels. Covert channels assume cooperation between both sides, so they're only relevant if one of the sides can't communicate trivially (e.g. via network)
agreed. AWS gets a lot of flak, but open sourcing firecracker was really great. I'd really prefer to see us move toward vms instead of containers, even if we kept the same k8s abstractions.
> .. covert ..
thanks for the catch, should have taken more time. Here's a better paper:
1. For me containers are one of those abstractions, defined by exposing an application controlled userspace. Containers can be implemented by different isolation technologies, from simple chroot/cgroup/namespaces... to VMs.
2. I'd still use chroot&co to partially isolate containers within a pod, while using VMs to strongly isolate pods from each other. This enables features like shared block-devices, unix-domain-sockets and monitoring the processes in an application container from a separate diagnostics container.
Containers absolutely are intended to be a security boundary.
VMs on the other hand actually are designed as a security boundary, but even then there are still attacks you can do against other VMs on the same box.
A key feature of OS virtualisation is the strong segmentation boundary between
1. Guests
2. Guests and the hypervisor.
For this reason, VMs are seen to provide a stronger security boundary than containers and are used in preference where that aspect is critical owing to environment, multi-tenancy, business context.
See also https://searchcloudsecurity.techtarget.com/tip/VMs-vs-contai...
There's a world of difference between the amalgamation of hacks that comprise cgroups and something like BSD jails, which are and afaik always have been intended to be a security boundary, which implements real first-class kernel isolation for jailed processes, not just another subtree under proc that provides some direction to the kernel around resource consumption/priority and relies on UID/GID hacks to control access.
A minimal VM, like firecracker has a small attack surface, so I'm willing to trust that privilege escalation/VM escapes will be rare.
A process restricted by cgroup/namespace/etc. still has access to the huge API surface exposed by the kernel, so privilege escalation is common, and I'm unwilling to trust this mechanism to isolate malicious code.
They didn't start out at the design phase that way, but they absolutely are today.
This was back in the early 2000's, and still seems to work rather well for the basic webhosting.