Understanding Docker Container Escapes
blog.trailofbits.com
blog.trailofbits.com
Also a good demonstration of how runc container security is really just Linux security, and the isolation provided by that type of container, is dependent on underlying Linux features.
I could see circumstances where this would be of practical use, where you've landed in a privileged container and are looking for a relatively easy/universal breakout to the underlying node.
The ToB breakdown is cool too as the initial tweet was a bit hard to parse, unless your Bashfu is good.
(ofc the post from ToB I'm really looking forward to hasn't been released yet, which is what they found in the Kubernetes audit)
1. Docker is not synonymous with Linux containers. LXC for example has always had much better security than Docker despite using the same infrastructure.
2. You don't need to interface the kennel directly. There is for example a Google project (gvisor?) that acts as a syscall proxy in userland to combat these types of problems.
LXC main dev is on HN. I'm sure he can explain this much better.
The Docker daemon runs as root (which is kind of inevitable, if you don't then you end up needing hacks like slirp4netns to hook up networking)
However, there's nothing innately inside Docker that requires contained processes to run as root.
Indeed it's a standard part of recommended Docker security guidance not to run your contained processes as the root user.
Docker also provide the facility to enable user namespaces at the daemon level, so that root inside the container != root on the host.
For example user mappings were handled very differently in lxc compared to (earlier versions of) docker.
LXC also supports unprivileged operation (which I named "rootless mode"). Docker gained support for this very recently as an experimental feature in 19.03 (still not released), but LXC has supported it (and defaulted to it whenever possible) for years. Though, the Docker one is arguably better in some respects of lack-of-privilege (thanks to great work from Akihiro Suda and Guissupe Scrvano) but it's still new.
LXC also has put a lot more work into fundamental security work (both in-kernel and within LXC).
Disclaimer: I maintain runc, the runtime Docker uses. There is no question that LXC has better engineering in this department. I collaborate with them quite often, but they have more engineers working on fundamental problems within containers.
1. Docker does support user namespaces today, correct? Your reply seemed to imply that it doesn’t.
2. Once rootless mode is released in Docker stable, the only difference in available security features between lxc and Docker will be the more flexible uid mapping for user namespaces, correct?
3. The flexible uid mapping feature, compared to user namespaces with static mapping as implemented by Docker, is an additional protection against container-to-container attacks, but not against container-to-host attacks. Did I get that right?
4. User namespaces, with or without flexible uid mapping, are considered a less secure containment method than seccomp and selinux/apparmor, all of which Docker/runc and lxc support equally well, correct?
Yes (though I don't agree my comment implied that Docker doesn't support user namespaces at all), but it doesn't support having different mappings for individual containers. This has both usability problems (--volume is painful to use) and security problems (inter-container attacks are still possible if you can "break out" of the container or otherwise disrupt the other container).
> 2. Once rootless mode is released in Docker stable, the only difference in available security features between lxc and Docker will be the more flexible uid mapping for user namespaces, correct?
Security features, (arguably) yes. But I would still argue that LXC has more security hardening work put into it than Docker. Of course they've had their own security issues but there definitely are arguments to be made that it isn't identical. I outlined some examples here[1].
Also the default configuration is still going to be run-as-root-without-user-namespaces with Docker (meaning the vast majority of users are running hideously insecurely). LXD and LXC defaults to using user namespaces. To be fair, both use seccomp and AppArmor/SELinux policies by default -- but depending on seccomp and AppArmor/SELinux is a much worse security position than
> 3. The flexible uid mapping feature, compared to user namespaces with static mapping as implemented by Docker, is an additional protection against container-to-container attacks, but not against container-to-host attacks. Did I get that right?
Yes.
> 4. User namespaces, with or without flexible uid mapping, are considered a less secure containment method than seccomp and selinux/apparmor, all of which Docker/runc and lxc support equally well, correct?
That's not quite true. User namespaces are arguably a much better containment method for containers. There are hundreds of user-namespace related hardening checks within the kernel (as well as the obvious "the euid space is different" protections) which you don't end up taking advantage of if you run in &init_userns. In fact, most kernel developers working in this space (namely Eric Biederman) don't consider security issues to be as serious if you can't exploit them without disabling user namespace protections. CVE-2019-5736 and CVE-2016-9962 were both blocked by using user namespaces.
But yes, there are some breakouts that user namespace support in your kernel have historically caused (and we have seen that many times) -- but that's why both Docker and LXD block unshare(CLONE_NEWUSER) with seccomp. But you can have all three! And (once Docker is configured) then all three support them all equally effectively.
I appreciate that you have a nuanced position on the topic of Docker security, based on deep expertise. Sadly, that nuance is lost on 99% of the people I see shouting that "Docker is insecure", the same people who presumably downvoted my original comment into oblivion. They are calling Docker insecure not because they understand what you explained (they don't), but because they have heard half-truths or outright fabrications, and are repeating them with absolute conviction, without bothering to argue their point or check even the most basic facts. As someone who has a lot of actual first-hand experience with Docker I find that very frustrating.
So, although I agree with everything you said, and appreciate that you took the time to write it down; I believe that your answer has unintentionally vindicated the many people lurking on this site who hold the widespread, almost cult-like belief that Docker is very insecure - insecure to the level of gross negligence, in a way that you and I understand it isn't.
Disclaimer: I'm a maintainer of runc, the runtime Docker uses.
It's how Docker doesn't have a good reputation for security.
ofc happy to be corrected, as I'm aware you'd know more about this :)
If you compare the defaults, LXC wins overall because they have rootless containers and user namespaces by default (runc has them too -- I implemented them -- but it's not the default in Docker). To be balanced, LXD's isolation of individual containers is not on by default either (because of backwards compatibility requirements) -- but Docker doesn't have an equivalent feature. If you configure a Docker setup to be as-close-as-possible to an LXC setup, then it's much harder to give a definitive answer. Generally, the containers we set up look almost identical from the kernel's point of view so we have similar kernel 0day problems. So it comes down to the security of the runtime in particular.
I am currently working on solving several pretty fundamental security issues that exist both within LXC and runc (and many more programs generally)[1], so it's not like either is perfect (though LXC does have more code to defend against the attacks I'm working on fixing). LXC does make use of more of the kernel hardening work that we (both the LXC folks and myself) have worked on. A trivial example is that LXC uses TIOCGPTPEER (a feature I originally implemented that allows you to avoid certain theoretical attacks by container processes against the runtime) but Docker doesn't use it (and because runc doesn't have a container manager by design we can't implement it in runc). LXC also supports using pidfds (a new feature in Linux 5.1 that Christian Brauner has been working on for a while) which allow much nicer methods of avoiding PID recycling race conditions -- with runc we still use the old pid+starttime method which is prone to well-known (though usually harmless) attacks.
Funnily enough, I'm actually giving a talk about this topic at the end of this week[2] and was writing slides when I saw this thread. :P
[1]: https://github.com/openSUSE/libpathrs [2]: https://2019.container.camp/au/schedule/securing-container-r...
LXC had unprivileged container support since 2013 so that part is fairly mature now. 'Unprivileged' in this case means the container process itself is running as a normal user.
[I maintain runc, and collaborate with the LXC folks.]
(as a mitigation)
okay, but even for official images, figuring out the provenance of a build on docker hub is totally impossible
I challenge you to start with an image sha and tell me what git version (or even what repo) was used to create it
docker needs to get better at supply chain
The number of people who do it in production is correlated with the complexity of the thing they are trying to do and how much the "fuck it, just do <some terrible idea>" relieves that complexity.
Because of this escape, giving access to /var/run/docker.sock to regular users is the same as giving them root access.
Also as the article says mounting /var/run/docker.sock is (now, because of this escape) the same as giving that container access to the host system.
docker run -itv /:/host ubuntu chroot /hostThe idea, of course, is to minimize the surface area for attacks. That container has a need for it to run that way; the vast majority of our other containers do not.
Ironically, attackers would have access to more valuable data (and more freedom of movement) on your personal/dev box than on a monitored production host.
The title should not in this case be editorialized; relative to the topic, it's about as boring as titles come.
But on the occasion that they're necessary, isolating from other pods with a network policy and having no public ingress is enough right?
Assuming of course that you trust the container process(es) - or is that the issue?
For example if you're running Kubernetes, then if you have a user who has RBAC rights to exec into a privileged pod, but doesn't have rights to create new privileged pods, then this could be a privesc risk, as they can use something like this to escalate to the underlying node after exec'ing into the privileged pod.
It (as with all things security) depends on your threat model :)
As far as I can tell, as long as there is no service/ingress on the privileged container, and a netpol blocks those that do from accessing it, it's less than ideal but 'ok' that this privileged container is running behind the scenes.
This is like acting surprised that a Linux root user can do immeasurable harm to the underlying OS.
Yes, the container runs in the "unconfined" profile after a restart.
Right, what you want is “privileged except for XYZ”, which is not supported by Docker. That’s a missing feature which is not the same as a bug. Calling it a security bug is even more misleading.
> Sometimes that happens, sometimes it does not. And when it does not, you have no way of knowing.
Right, it should fail every time. That is a bug. But it’s not security bug, and fixing that bug won’t give you the feature you want, it will just make it clearer that the feature is not supported.