How to Escape a Container
panoptica.app
panoptica.app
2 needs Docker socket, it's explicitly meant for running other workloads
3 needs shared PID namespace in addition to SYS_PTRACE, neither is granted by default by any container runtime
4 needs SYS_MODULE, again no one has that
5 and 6 need DAC_READ_SEARCH, no one grants that, no one uses that
None of those seem like vulnerabilities or things that would be available without the admin taking explicit steps to specifically want to allow escaping. Being root in the container would not be enough to get any of those capabilities.
Not really unfortunately. A novice sysadmin granting someone's request to add some permission called SYS_MODULE may not know that it could be equivalent to full root access. That's why posts like this are important for education.
...is a social engineering exploit, not a technical one.
> A novice sysadmin granting someone's request to [be able to run some tool/software that says it needs this, but is in fact secretly malicious]
or perhaps more common
> A novice sysadmin granting someone's request to [be able to run some legitimate tool or software, then later a virus in the container taking advantage and breaking out]
for a real world example, the Nvidia Nsight docs specifically direct you to add SYS_ADMIN; do you think your everyday ML engineer would know that doing so poses a security risk if they're doing this on Docker?
You still can't run untrusted code but at least you can isolate stuff like Compose environments
It confuses the heck out of me because red-team-type people will shout from the mountaintops about how insecure containers are and how absolutely trivial it is to break out of them, but when I go look for an example of that trivialness, all I find is stuff like this. Not that these aren't at all useful techniques - there are plenty of containers with --privileged, or a Docker socket mount, etc. But surely this doesn't apply to >90% of containers out there, especially ones that are exposed to the Internet. Your average Redis or Nginx container, or some container running a Python or Node webapp, is not going to have Docker mounted or some weird capability added. Sure, misconfigs happen, us sysadmins get lazy, but this is really common-sense stuff. It feels almost unfair to "blame" it on the container.
Of course, as mentioned, there are 0-days that allow for container breakout, and those are the truly scary stuff. But they seem to be few and far between, and they get publicized semi-widely and patched pretty quickly when they are found.
So to this day I don't really understand all the security folks who act like they (as in any decent attacker, not just a nation-state with 0days up the wazoo) could break out of a container with their eyes closed, while the only material I can find on the open Internet is stuff like this. Am I just looking in the wrong places?
A part of me still feels like that's somewhat separate from a "typical" scenario where you want to harden a container that's exposed to the outside world in some way; like I said, your average container running a webapp or whatever. Nvidia Insight seems to be a performance analysis tool, so I would hope nobody is running it with untrusted input or exposing it to the Internet. (Yes, I know someone out there totally is.)
Similarly, it feels near-obvious to me that adding a privilege called "SYS_ADMIN" to a container will make it more or less equivalent to root on the host. It's not like Docker hides this info from people, last I checked it's explained pretty prominently.
You're totally right overall, it still matters, it's just something about this framing that rubs me the wrong way.
I think your response betrays a broken line of communication in our industry. People tend to assume most production environments are configured by people who knew what they were doing, were deploying a well-behaved application and had time and management support to do a good job. That's almost never the case, even in well-funded, well-regarded tech companies.
So these ”escapes“ are basically giving a container full access to the host and using that. None of them are enabled by default.
Generally, it’s not the stuff the container is intended to run that’s doing the escape, though. This is usually the second step, after getting local execution through a vulnerability in the application running in the container.
This, by the way, is why I think Google had it right by splitting most products into SWE and SRE groups, so you had a group of people who could focus full-time on the production environment. SWEs are rewarded for building and deploying stuff, so they're going to build and deploy stuff.
unfortunately cannot find it anymore.
on edit - maybe it was write permissions to the file system, but largely the same level of "once you have this it is game over"
https://devblogs.microsoft.com/oldnewthing/20060508-22/?p=31...
And in this case. If you escaped the container, you are literally nobody. All you allowed to read is public files that any user can read.
By the way. If you use rootless container, this is the default settings you will get. Since you don't have access to uid0 at first place.
Judging by the number of HOWTOs I see on GitHub and YouTube that tell you to check the Privileged box, this assumption isn't worth what you think it is.
- You need to mount a filesystem, use a raw socket, etc. and Stack Overflow / ChatGPT tells you to enable a capability, so you do. It works, so you check it in.
- You're deploying something legacy, that assumes it has more control of the OS than it does inside a container.
- You're a data engineer. All of your tools assume they run as root, because that's just how data science is.
- You're building a "sidecar" for monitoring, control over other containers or something similar. The extra privileges are needed to do your job.
- You're trying to access something you can't, and it's generally easier to overprovision access than to do least privilege.
These problems aren't unique to containers and the cloud, mind you. I saw the same problems, e.g. when working on mobile device security. In general, it's a lot easier to just turn off half of SELinux than to learn how to configure it to do what you want, especially if you have a deadline.
It feels like "breaking out of vanilla containers" and "breaking out of misconfigured containers" are two different topics, two different threat models. And while the second absolutely matters in the real world, the really scary stuff is obviously the first (and usually involves 0-days, kernel exploits, etc?). But people seem to talk less about the first.
I also wrote my own docker-like containerization code for educational purposes a while ago, so container has both these meanings for me. Yet, me brain was expecting a physical escape story. Brains are funny!
I, however, have not written any containerization code.
Also many of the cappabilities described in this article aren't compatible with a rootless user deployment scénario.
Not sharing the kernel with the host os (or other containers) is a huge security boundary.
A bit of debootstrap, a few apt-get commands, and copying in config files, and you have a lightweight VM, minimal image.
Something people have been doing for 20 years.
There are also sorts of tricks, such as having two images, one for the app layer, one for the OS, which makes the deploy for app updates faster.
I'm not even sure why people care about image size all that much. You copy it to your local cluster, then deploy from there.
But seriously though, it’s so you can write exploits or satisfy that curious itch when working with a cloud service.
But honestly, I was hoping before I clicked it that this was going to be about how to escape from the inside of a shipping container.
You don't, they cannot be opened from inside once locked. Also they're airtight, so bang on the walls and hope help arrives before you suffocate.
That's more nightmare inducing than I was hoping.
Really?! I know that these things are built to be very stable, but they never gave me the impression to be air-tight. Not bad. I'm sure the world of shipping containers is a huge rabbit-hole to read into.
Related video :) https://www.youtube.com/watch?v=-trd_f6j3eI
Containers are not a security boundary.
Containers are not a security boundary.
Any system that treats them as such is inherently compromised.
I have so far only used it for hosting some gameservers which I don't trust, i.e some simple containers, but I really want to try it in a new k3s cluster once I get it setup and move some services there. I like the idea of putting internet facing ones into it as an additional layer of separation and could imagine it being useful in production.
> supposed to keep whatever is inside trapped unless you poke holes in that protection
As far as I know that was never a design decision for containers on Linux; certainly not in the early days.
EDIT: For that matter, they're clearly being used for security; the features in Linux that are used by runc et al. are the same features used by eg. Chrome to isolate components in order to contain vulnerabilities.
>Turns out the ufw firewall I enabled and diligently kept on a strict allowlist with only my internal servers didn’t work on a new server because of Docker. When I containerized MongoDB, Docker helpfully inserted an allow rule into iptables, opening up MongoDB to the world. So while my firewall was “active”, doing a sudo iptables -L | grep 27017 showed that MongoDB was open the world. This has been a Docker footgun since 2014.
Story was previously discussed on HN[1]. Sure, you could argue the author should have done more to secure the endpoint, but this was 100% a failure mode due to how Docker prioritizes convenience over security.
[0] https://blog.newsblur.com/2021/06/28/story-of-a-hacking/
so use "-p 127.0.0.1:5432:5432"
- https://github.com/docker-library/postgres/issues/770
- https://sysdig.com/blog/zoom-into-kinsing-kdevtmpfsi
- https://sysdig.com/blog/cloud-defense-in-depth/
- https://thenewstack.io/kinsing-malware-targets-kubernetes/
- https://stackoverflow.com/search?q=kinsing
- https://github.com/search?q=repo%3Adocker-library%2Fpostgres...
-----------
https://docs.docker.com/network/packet-filtering-firewalls/
"On Linux, Docker manipulates iptables rules to provide network isolation. While this is an implementation detail and you should not modify the rules Docker inserts into your iptables policies, it does have some implications on what you need to do if you want to have your own policies in addition to those managed by Docker.
If you're running Docker on a host that is exposed to the Internet, you will probably want to have iptables policies in place that prevent unauthorized access to containers or other services running on your host. This page describes how to achieve that, and what caveats you need to be aware of."
I think the hardware can help bridging the gap between containers and VMs by enabling userspace processes behave as VMs, which is more or less what QEMU+KVM try to do, except that it still comes with some overheads and less flexibility.
Maybe I am wrong. We can wait for a security professional to comment.
Also, one can achieve similar effects with containers as well, just think AppArmor, capabilities, permissions, etc., all layers of administrative privileges between some untrusted code and the host.
But I guess that doesn't mean virtual machines aren't easily escapable without extra work, same as containers.
That said, it’s hard to get right at all times.
VMs easily give a false sense of security especially with any kind of network-based trust.
who's everybody? There's special kind of VM hosts for that, containers is like your kitchen jars, if someone is vomiting with Ebola in your kitchen - your jars will not help you
I think the general advice is that a single container can never be a robust security boundary because the OS surface area they involve is so large that the isolation layer is ripe for possible vulnerabilities. You also really have to avoid screwing up, there are lot of fiddly little security mistakes you can make when attempting to use a container to run untrusted code.
Typically you might use something like gvisor, or a VM. Systems where isolation is simpler to reason about and the attack surface is smaller.
In any case a single isolation boundary can have a vulnerability and my understanding is that more advanced systems typically involve multiple layers of isolation to sandbox untrusted code.
I find it super frustrating that we're stuck with kernels with inherent weaknesses to their security approach that we have to re-implement them in userspace in one way or another (gVisor, Firecracker, etc.) just to get the hardware-provided userspace/kernel boundary to work properly.
edit: typo
- Networking is the obvious one in this scenario. By default (on Docker/LXC/others) all containers are on the same virtual bridge and can communicate with each other. Even with some additional configuration and isolation it is possible to MITM attack other containers on the same host.
- It is very easy to DDoS adjacent containers e.g. by spamming signals, forking new processes, creating files. There is again no default safeguard against this.
cgroups can prevent containers from using too much memory or cpu
If a process's network namespace contains a network device then it is not an escape to use it by definition. If the network namespace for a process contains no devices then being able to use a network device would be an escape.
services:
your_service:
image: your_image
networks:
isolated_network:
another_service:
image: another_image
# This service is on the default network
I'd be curious how a vanilla (actual) networking setup is full of holes...Does that also apply to Kubernetes workloads? And does that then require an encrypted service mesh (e.g. Linkerd) or TLS between services?
Still a good idea to use tls.
That's silly. Of course containers are a security boundary. They have advantages and disadvantages. Treat them as tools and not slogans.
The name is at least misleading if not wrong, then? What do they "contain"?
Under the hood, they're a fancy wrapper around a pile of tar files.
Tar files certainly contain other files and are also not a security boundary.
1. You get file, process, and network namespaces, which are a security boundary
2. You get a seccomp filter, which is a security boundary
The "containers are not a security boundary" meme needs to die.
Elsewhere you mention that "containers are not sufficient for untrusted code" but that's a very specific and very niche threat model. Most people don't say "send me a binary and I'll execute it", or have arbitrary RCE + multitenancy concerns.
Containers aren't sufficient for multi-tenant RCE because the RCE is by design so 100% of your security pressure is on the container at that point. In the vast majority of cases you're dealing with servers that don't intend to allow arbitrary code execution, and containers are an extremely easy way to drive up the cost of an attack given that the attacker has already spent a lot of time and money on the RCE.
SELinux is also not sufficient for the "RCE by design" threat model - is SELinux not a security boundary?
Further, containers can limit the impact of remote vulnerabilities like path traversal attacks, since they have file isolation by default.
edit: I see elsewhere that there's a real lack of clarity here.
First off, a security boundary can be meaningfully defined as a limitation on an attacker that does not have a way around it without additional exploitation.
So the main reason why people have said "containers are not a security boundary" is because:
a) Very, very early on, escaping a container was trivial - like you could just ask to leave and you'd be out.
b) There were some blog posts basically saying "containers aren't sufficient for multi-tenancy" where arbitrary users can run arbitrary code on the same host. This is still the case today - but it's also an extremely rare threat model.
Why would containers not be sufficient for (b) ? Because the majority of the Linux kernel is still exposed to the attacker within a container - the vast majority of system call interfaces are exposed (but seccomp removes a number of these, which is nice). The Linux kernel is not at all sufficiently hardened against attackers who can make arbitrary system calls, therefor containers are not sufficient against those attackers. If you give the attacker RCE by default (ie: your service is "Send me a binary and i'll run it") then an attacker can spend all of their time and money just on a local privesc, which isn't crazy difficult.
Since the cost of RCE in an RCE-aaS is 0 the consensus is that containers aren't strong enough for RCE-aaS threat models. In that case use Firecracker or gVisor or a dedicated host.
Otherwise, RCE costs tend to be pretty high and having to develop an additional LPE on top of one is, at minimum, quite a pain for many attackers.
Containers are extremely easy to deploy software into, something like a Firecracker VM is not. Containers are basically just processes, so you can monitor them and manage them trivially. Monitoring and managing VMs with processes inside of them is obviously harder. So I think the 'bang for your buck' with containers is extremely solid.