Leaky Vessels flaws allow hackers to escape Docker, runc containers
bleepingcomputer.com
bleepingcomputer.com
Does container-selinux limit this container escape vulnerability?
I will continue to stick with VMs as my security boundary.
> gVisor provides a virtualized environment in order to sandbox containers. The system interfaces normally implemented by the host kernel are moved into a distinct, per-sandbox application kernel in order to minimize the risk of a container escape exploit. gVisor does not introduce large fixed overheads however, and still retains a process-like model with respect to resource utilization.
https://news.ycombinator.com/item?id=38609105
kata containers: https://github.com/kata-containers :
> Kata Containers is an open source project and community working to build a standard implementation of lightweight Virtual Machines (VMs) that feel and perform like containers, but provide the workload isolation and security advantages of VMs.
I even resorted to posting instructions on stack exchange demonstrating how to walk /sys to find device numbers and use mknod to read the hosts root volume and now as most servers don't mount their efi partition, how you could mount it rw inside a container with little effort.
Containers are namespaces and purely depend on privilege dropping for the security they provide.
Part of the problem is that container breakout is narrowly defined.
The fact that a privileged container can upload firmware or access private keys by reading the hosts root volume didn't count.
While the ability to disable privileged mode wouldn't have solved this issue it still would have reduced the attack surface. But will the projects refusal to even take that step I gave up.
Deciding that the only safe option was to consider containers as namespaces and nothing more.
Unfortunately adding persistence and other functionality tends to result people running it as uid0, which means that you have to consider as anything that can launch a container as having superuser privileges.
Buildkit doesn't look too useful(?) with podman as some of its better features are built in to podman.
In many circumstances you can make some tradeoffs, like running multiple customers or projects on the same vmware installation, because the risk is low or the other tenants are well known. I don't think Kubernetes is quite there yet where I'd trust multiple tenants on the same cluster.
As an example, a VM from a reputable cloud provider is secure enough for most purposes, likewise I'd trust managed containers from a reputable cloud provider (except Azure, they have a dodgy track record when it comes to container security)
"Security settings that you specify for a Container apply only to the individual Container, and they override settings made at the Pod level when there is overlap."
https://kubernetes.io/docs/tasks/configure-pod-container/sec...
While there are other Byzantine faults like this CSV, that should give you a hint.
The problem is that containers depend on good actors dropping privlages on a shared kernel.
Where VM's have a natural abstraction and their own kernel, containers require proper configuration of seccomp, apparmor, selinux, dropping privileges, etc... to reduce the attack surface.
An easy demonstration of this is to launch a privlaged container and run insmod with some random kernel module on the host OS.
In namespaces, you start with no isolation, and you add whatever you want or more realistically remember to add. They are not jails where you start with a reasonable secure baseline.
In fact the k8s 'baseline' pod security standard still has MKNOD caps which can result in issues like CVE-2021-25741. While /proc now has filtering /sys doesn't and you can find the major minor numbers of the hosts root filesystem and read it from a container.
It is not that containers can't be made reasonably secure, but that they aren't inherently secure. They are default allow all because they are inherently just namespaces implemented basically through value remapping.