Containers are not a security boundary.
Containers are not a security boundary.
Any system that treats them as such is inherently compromised.
Containers are not a security boundary.
Containers are not a security boundary.
Any system that treats them as such is inherently compromised.
- Networking is the obvious one in this scenario. By default (on Docker/LXC/others) all containers are on the same virtual bridge and can communicate with each other. Even with some additional configuration and isolation it is possible to MITM attack other containers on the same host.
- It is very easy to DDoS adjacent containers e.g. by spamming signals, forking new processes, creating files. There is again no default safeguard against this.
cgroups can prevent containers from using too much memory or cpu
If a process's network namespace contains a network device then it is not an escape to use it by definition. If the network namespace for a process contains no devices then being able to use a network device would be an escape.
services:
your_service:
image: your_image
networks:
isolated_network:
another_service:
image: another_image
# This service is on the default network
I'd be curious how a vanilla (actual) networking setup is full of holes...Does that also apply to Kubernetes workloads? And does that then require an encrypted service mesh (e.g. Linkerd) or TLS between services?
Still a good idea to use tls.
That's silly. Of course containers are a security boundary. They have advantages and disadvantages. Treat them as tools and not slogans.
I have so far only used it for hosting some gameservers which I don't trust, i.e some simple containers, but I really want to try it in a new k3s cluster once I get it setup and move some services there. I like the idea of putting internet facing ones into it as an additional layer of separation and could imagine it being useful in production.
> supposed to keep whatever is inside trapped unless you poke holes in that protection
As far as I know that was never a design decision for containers on Linux; certainly not in the early days.
EDIT: For that matter, they're clearly being used for security; the features in Linux that are used by runc et al. are the same features used by eg. Chrome to isolate components in order to contain vulnerabilities.
>Turns out the ufw firewall I enabled and diligently kept on a strict allowlist with only my internal servers didn’t work on a new server because of Docker. When I containerized MongoDB, Docker helpfully inserted an allow rule into iptables, opening up MongoDB to the world. So while my firewall was “active”, doing a sudo iptables -L | grep 27017 showed that MongoDB was open the world. This has been a Docker footgun since 2014.
Story was previously discussed on HN[1]. Sure, you could argue the author should have done more to secure the endpoint, but this was 100% a failure mode due to how Docker prioritizes convenience over security.
[0] https://blog.newsblur.com/2021/06/28/story-of-a-hacking/
so use "-p 127.0.0.1:5432:5432"
- https://github.com/docker-library/postgres/issues/770
- https://sysdig.com/blog/zoom-into-kinsing-kdevtmpfsi
- https://sysdig.com/blog/cloud-defense-in-depth/
- https://thenewstack.io/kinsing-malware-targets-kubernetes/
- https://stackoverflow.com/search?q=kinsing
- https://github.com/search?q=repo%3Adocker-library%2Fpostgres...
-----------
https://docs.docker.com/network/packet-filtering-firewalls/
"On Linux, Docker manipulates iptables rules to provide network isolation. While this is an implementation detail and you should not modify the rules Docker inserts into your iptables policies, it does have some implications on what you need to do if you want to have your own policies in addition to those managed by Docker.
If you're running Docker on a host that is exposed to the Internet, you will probably want to have iptables policies in place that prevent unauthorized access to containers or other services running on your host. This page describes how to achieve that, and what caveats you need to be aware of."
I think the hardware can help bridging the gap between containers and VMs by enabling userspace processes behave as VMs, which is more or less what QEMU+KVM try to do, except that it still comes with some overheads and less flexibility.
Maybe I am wrong. We can wait for a security professional to comment.
Also, one can achieve similar effects with containers as well, just think AppArmor, capabilities, permissions, etc., all layers of administrative privileges between some untrusted code and the host.
But I guess that doesn't mean virtual machines aren't easily escapable without extra work, same as containers.
That said, it’s hard to get right at all times.
VMs easily give a false sense of security especially with any kind of network-based trust.
who's everybody? There's special kind of VM hosts for that, containers is like your kitchen jars, if someone is vomiting with Ebola in your kitchen - your jars will not help you
I think the general advice is that a single container can never be a robust security boundary because the OS surface area they involve is so large that the isolation layer is ripe for possible vulnerabilities. You also really have to avoid screwing up, there are lot of fiddly little security mistakes you can make when attempting to use a container to run untrusted code.
Typically you might use something like gvisor, or a VM. Systems where isolation is simpler to reason about and the attack surface is smaller.
In any case a single isolation boundary can have a vulnerability and my understanding is that more advanced systems typically involve multiple layers of isolation to sandbox untrusted code.
I find it super frustrating that we're stuck with kernels with inherent weaknesses to their security approach that we have to re-implement them in userspace in one way or another (gVisor, Firecracker, etc.) just to get the hardware-provided userspace/kernel boundary to work properly.
edit: typo
1. You get file, process, and network namespaces, which are a security boundary
2. You get a seccomp filter, which is a security boundary
The "containers are not a security boundary" meme needs to die.
Elsewhere you mention that "containers are not sufficient for untrusted code" but that's a very specific and very niche threat model. Most people don't say "send me a binary and I'll execute it", or have arbitrary RCE + multitenancy concerns.
Containers aren't sufficient for multi-tenant RCE because the RCE is by design so 100% of your security pressure is on the container at that point. In the vast majority of cases you're dealing with servers that don't intend to allow arbitrary code execution, and containers are an extremely easy way to drive up the cost of an attack given that the attacker has already spent a lot of time and money on the RCE.
SELinux is also not sufficient for the "RCE by design" threat model - is SELinux not a security boundary?
Further, containers can limit the impact of remote vulnerabilities like path traversal attacks, since they have file isolation by default.
edit: I see elsewhere that there's a real lack of clarity here.
First off, a security boundary can be meaningfully defined as a limitation on an attacker that does not have a way around it without additional exploitation.
So the main reason why people have said "containers are not a security boundary" is because:
a) Very, very early on, escaping a container was trivial - like you could just ask to leave and you'd be out.
b) There were some blog posts basically saying "containers aren't sufficient for multi-tenancy" where arbitrary users can run arbitrary code on the same host. This is still the case today - but it's also an extremely rare threat model.
Why would containers not be sufficient for (b) ? Because the majority of the Linux kernel is still exposed to the attacker within a container - the vast majority of system call interfaces are exposed (but seccomp removes a number of these, which is nice). The Linux kernel is not at all sufficiently hardened against attackers who can make arbitrary system calls, therefor containers are not sufficient against those attackers. If you give the attacker RCE by default (ie: your service is "Send me a binary and i'll run it") then an attacker can spend all of their time and money just on a local privesc, which isn't crazy difficult.
Since the cost of RCE in an RCE-aaS is 0 the consensus is that containers aren't strong enough for RCE-aaS threat models. In that case use Firecracker or gVisor or a dedicated host.
Otherwise, RCE costs tend to be pretty high and having to develop an additional LPE on top of one is, at minimum, quite a pain for many attackers.
Containers are extremely easy to deploy software into, something like a Firecracker VM is not. Containers are basically just processes, so you can monitor them and manage them trivially. Monitoring and managing VMs with processes inside of them is obviously harder. So I think the 'bang for your buck' with containers is extremely solid.
The name is at least misleading if not wrong, then? What do they "contain"?
Under the hood, they're a fancy wrapper around a pile of tar files.
Tar files certainly contain other files and are also not a security boundary.