Don't expose the docker socket, even to a container
lvh.io
lvh.io
You mean Linux containers are not designed for this, nor are they designed to be secure. What a sad design failure.
Some [1] container technology was developed with security as a first principle.
[1]: http://us-east.manta.joyent.com/jmc/public/opensolaris/ARChi...
https://www.joyent.com/blog/triton-docker-and-the-best-of-al...
I'd be curious if the Joyent folks believe Triton to have the security properties that everyone seems to want out of Docker. It doesn't involve the standard Docker daemon, so it might.
Yesterday at http://containersummit.io/, Bryan's talk "Going Container Native" (video coming soon) was in the same vein, SmartOS (illumos) directly benefits from its Solaris lineage -- Solaris Zones were engineered for security. We, Joyent, absolutely do believe Triton has the security properties that everyone seems to want out of Docker, because Joyent has been running OS containers in multi-tenant production since ~2006.
Open source SmartDataCenter [1], rebranded Triton, is the exact same code that is run in Joyent's Public Cloud. This truth in code, product and service contributed to me leaving an OpenStack company to join Joyent two years ago.
Which is totally expected, because docker tries to be useful to the largest user-base possible. It would be quite harder to use if didn't support directory mounting (via -v). And it would be a total nightmare for almost every user if you had to specify a list of allowed syscalls for every container.
This reminds me the situation with SELinux a lot. It has improved a lot but I still see "disable SELinux" almost in every tutorial I read on CentOS, Fedora or RHEL.
Because security is hard.
Just look at the Mac OS X users running as root and disabling Gatekeeper.
Or the developers that stay away from the sandbox model.
One of the nice things of mobile OSes is that there isn't a way around the container model. Although the history with permissions kind of messes it.
Defense in depth. With KVM, you'd have to exploit the guest kernel, then exploit KVM's hardware emulation to get to host userspace, then escalate to host root/kernel. (And KVM can use seccomp to limit itself, making that last step harder.) With Linux containers or Solaris Zones, you run directly on top of the host kernel, so you just need a single kernel exploit.
In any case, Zones and containers have the same security properties (running directly on top of the kernel), so the original comment about Zones being more secure than containers makes no sense.
http://www.cvedetails.com/cve/CVE-2008-5689/
I also know of only one FreeBSD Jail escape CVE in its entire existence.
https://www.freebsd.org/security/advisories/FreeBSD-SA-04:03...
These were years ago. Jails and zones are everywhere. You're kidding yourself if you think people haven't tried to attack them.
Solaris got hit by the one in 2012 which is the non-canonical RIP which hit a bunch of operating systems.
FreeBSD has had numerous kernel exploits which could be used for zone escapes.
edit: well i guess if your attacker knows you're on FreeBSD they could override to prison0 in the exploit. But it's a good thing we dont have a ton of local priv exploits, eh?
edit2: and that was a pretty ignorant statement of mine about root exploits anyway :)
https://docs.sandstorm.io/en/latest/developing/security-prac...
That said, there is a cost in compatibility. Most apps can be made to run just fine in this constrained environment, but it does sometimes require tweaks. Docker is more interested in compatibility than in security, so naturally they don't do this kind of attack surface reduction by default. (You can configure it manually, but realistically if it isn't mandatory then few people will bother.)
The meaning of "Linux containers are not secure" is that untrusted code should not be given root privileges within a container. Google is generally not doing that. They have e.g. trusted Gmail code running on the same machine as trusted YouTube code, handling untrusted emails and untrusted videos. But the Gmail team is not worried about the YouTube team hacking them, or vice versa. The security mechanisms just need to keep honest people honest.
And when they do have untrusted, third-party code to run, for Google Cloud Platform, they use VMs or actual sandboxes: see section 6.1 of the paper you linked.
http://csl.stanford.edu/~christos/publications/2015.heracles...
Basically, if someone breaks into Gmail, you're already in deep trouble. You're not significantly better off because they didn't break into Hangouts. Same in the other direction. So you can run those on the same hardware.
If you think YouTube is less sensitive than Gmail (I don't know if Google actually does), you can separate those, sure. But there are enough applications that are no worse to break into than YouTube, like Photos. You can run those on the same hardware.
At Google's scale it's not too hard to say, this is the stuff we're ridiculously paranoid about and this is everything else, and get 80% CPU utilization on both infrastructures.
https://blog.docker.com/2013/08/containers-docker-how-secure... concludes,
"Docker containers are, by default, quite secure; especially if you take care of running your processes inside the containers as non-privileged users (i.e. non root)."
They also recommend using SELinux with Docker to beef up security (see https://blog.docker.com/2014/07/new-dockercon-video-docker-s...)
As long as people are using containers as a security boundary it makes sense to pay attention to things like this one about the control socket.
People find privilege escalation bugs in the Linux kernel often enough that you can't really claim that running arbitrary native code is "secure" unless you're doing something to mitigate these. Things like:
http://www.openwall.com/lists/oss-security/2015/07/22/7
Note that SELinux won't do anything about this kind of bug. You have to use seccomp and similar to disable large swaths of the Linux kernel API, particularly exotic parts that aren't well-tested or reviewed.
https://docs.sandstorm.io/en/latest/developing/security-prac...
Admittedly, though, "only minimal modification" is probably too much modification for Docker's goals. That's fair. But I don't think that gives them a pass to pretend the problem doesn't exist.
To paraphrase Theo de Raadt:
‟You are absolutely deluded, if not stupid, if you think that a worldwide collection of software engineers who can’t write operating systems or applications without security holes, can then turn around and suddenly write virtualization layers without security holes.”
[1] http://blog.valbonne-consulting.com/2015/04/14/as-a-goat-im-...
However my main frustration is I'm not really sure what the article is advocating for here. Instead of not giving access to the docker daemon to containers (which is legitimately needed for complex deployments where one container needs to dynamically start up "sibling" containers, e.g. a CI service), wouldn't it make more sense to talk about not viewing Docker container security the same as VM security in the first place?
Sure if you're going to do that anyway it makes sense to disable access to the socket, but then there's a million other things you'll have to do because docker containers are currently not primarily intended as a replacement for the isolation security of VMs. Their security is more like a useful extra layer, rather than a full blown replacement.
They can boot a new VM with Docker images in 200ms, which is very close to LXC, and perfectly isolated by hypervisor.
The problem of "Virtual Machine" is not "Virtual"/Virtualization, the problem is the full blown guest OS, aka "Machine".
If you want Docker + security isolation, I'm intrigued by Clear Containers, which is a lightweight KVM-based virtualization thing:
https://lists.clearlinux.org/pipermail/dev/2015-September/00...
Isn't that what a process is? The containers in Linux are based on Jails from FreeBSD and zones from Solaris. They are absolutely there for security.
Regarding the remaining part of your post, I understand what you are trying to show but python is a really bad example. You absolutely can have python 2 and 3 side by side, or even different minor versions. And with virtualenv or pyvenv (that came with 3.4) you can even have multiple installation of the sane version. If you add setuptools to your application you can easily generate single file package (I personally like wheel) the deployment is as simple as writing pip install myawesomeapp-1.0.py2.py3.whl it downloads all dependencies. There is not much that Docker would help, it only makes things more complex.
Still Docker the company's core value proposition is a hosted registry, something many savvy corporations will never go for. Docker the product could probably do just fine if the company were to fold.
and with everything moving to services, I see the utility of actually using components diminishing fast (well except for those providing those services)
but for everyone else, docker solves no actual problem that can't already be solved now.
But what you really want in that case is to link the application statically. If you don't want to have benefits of shared objects:
- smaller binary - memory savings (if multiple programs are using the same library, it is loaded once) - less files to patch to fix a security vulnerability
The share objects have these features but it comes at price of lower performance, so by putting all .so files into a single docker file instead of statically compiling your application you're getting worst out of both worlds.
(Containers are still good, for isolating concerns and management. Multiple versions of the same library is just not it.)
It's the level of isolation of a process, yes. Just as two processes can use their address spaces as they see fit without bothering each other, even loading different versions of the same library, under a Linux container, two applications can use their filesystem as they see fit without bothering each other, even using different versions of the same binary applications.
But the security isolation between two processes running as the same user account is extremely weak. While it's true that one process can't write to another one's memory directly, it's not a fundamental breach of the security policy if it can do so indirectly. There may be things to increase defense-in-depth (like Yama) but fundamentally if you're the same UID there is no security boundary. The same rule applies to containers.
> python is a really bad example
Yeah, agreed. I was just trying to come up with something quick. If your app works with v(irtual)env, by all means just use that and stop messing with containers. However, if you've got some large closed-source app with a portion in Python, and it expects /usr/bin/python to both work and be some exact version, you need to virtualize the filesystem.
I'm sorry for disagreeing, but that's what chroot() or even chdir() supposed to do. It's not for security (process can fool it), but they do provide isolation assuming there are no malicious actors.
Containers were created to provide security, perfect example is FreeBSD Jail which precedes Solaris zones. It supposed to be secure version of chroot() which should not be escapable. It was successfully used in early 2000 before VMs to provide shared hosting.
> But the security isolation between two processes running as the same user account is extremely weak. While it's true that one process can't write to another one's memory directly, it's not a fundamental breach of the security policy if it can do so indirectly. There may be things to increase defense-in-depth (like Yama) but fundamentally if you're the same UID there is no security boundary. The same rule applies to containers.
Agreed, except the last sentence. With processes the isolation is weak because same UID represent the same user, if you use a different UID the isolation is enforced. The containers (assuming they are correctly set up) allow you to actually have two root accounts that can't interfere with each other.
> Yeah, agreed. I was just trying to come up with something quick. If your app works with v(irtual)env, by all means just use that and stop messing with containers. However, if you've got some large closed-source app with a portion in Python, and it expects /usr/bin/python to both work and be some exact version, you need to virtualize the filesystem.
Assuming these are not malicious, you can just do:
chroot /app1_root python myapp1.py chroot /app2_root python myapp2.py
And as long as they don't use any tricks to get out, they won't step on each others teas.
To the best of my knowledge, Docker (the official implementation) does not do that. rkt does, as mentioned at the bottom of this blog post mentioned elsethread:
https://coreos.com/blog/rkt-0.8-with-new-vm-support/
(The Linux implementation of this is somewhat poor, in that you need to have a separate UID reserved in the global namespace, and you can only do 1:1 maps in containers. A nicer implementation would treat the user principal as a (container, UID) tuple. I recall that Linux tried that, but gave up for backwards-compatibility reasons.)
> chroot /app1_root python myapp1.py
Yeah, I think 80% of what Docker actually gets people in practice is a system for managing and running things in chroots. Containers also let you give them separate networking setups, track PIDs properly, and apply resource controls. But I've seen homegrown approximations that preceded Docker, based on stuff like schroot.
and we are encouraging this now, instead of fixing the damned apps?
sheesh this makes me feel old.
Most apps need very little from the underlying OS if you actually take the time to e.g. set up a toolchain with a build container that you then move the build artefacts out of to install into the final container. Instead you see a lot of containers that in effect include all the build dependencies and a nearly full OS pulled in by that.
It is improving, though slowly.
Plus they always base it on images from the Internet, so bascially we trust some stranger with root privileges to all our data. Not always on the same image of course.
I'd be inclined to agree. The reductionist sum of mechanisms that make a Linux container have always been about detaching, multiplexing and partitioning kernel resource subsystems. Docker was the first program that really hyped it into the idea of being about application deployment, but I fear this gives people wrong impressions and makes the mistake of treating an emergent property as if it were a fundamental.
I have no idea what point you're trying to convey with the second part. That applications are not business logic over kernel resources is a curious argument to make.
There are a lot of benefits to containers and they don't have to be insecure. More efficient resource utilization and orders of magnitude faster allocation and launching to name two.
Google runs a significant portion of its internal operations in a container infrastructure and has for quite a while.[1]
They're perfectly capable of deployment into production environments.
I won't comment on docker as I haven't spent the time to fully grok all its warts.
They use containers inside virtual machines. Virtualization for security, containers for deployment.
If your goal is strong isolation, then VMs are definitely better today. The purpose of Docker and similar container technologies is not that kind of isolation. It's to package up and distribute applications in a way that's more decoupled than simply installing them all on the same system.
Google is using containers instead of VMs. This still provides security isolation and allows them to use resources more efficiently (VM has overhead where you need a whole OS for every instance).
This approach does not make much sense in public cloud, where you already run inside of VM and the overhead is really for Amazon not you. So I see Docker is now pivoting to be a package manager, but there are already tools that do that. You can argue that Docker is simpler but so was rpm when it started. As Docker will grow it will become more complex in order to support all functionality package format already provides. There might be an argument that you can run multiple Docker containers on a single host, but that's what processes are for.
There is change happening, and looks like cloud companies want to create "cloud os", I guess Docker is step toward that direction, but at current state in don't see it offering anything valuable to the organization that uses it.
If it's a binary, why not compile it statically? That way it's a single file and it also will perform faster (using shared objects has overhead)
I 100% agree with the compile it statically thing. This is what many major internet properties likes Google do and I would argue part of why Go is so popular. But, it is hard for the vast majority of applications that expect to open files for assets and lack build systems for static compilation.
1) Easy global mirroring of software using "boring" protocols like http and ftp
2) Cryptographic signing of the software so we can trust mirrors and systems that put users in control of who to trust.
3) Human significant package names that are easy to pull onto a host e.g. `apt-get install $name`
Where package management broke down:
1) Package collisions e.g. If I want to install a new custom build of python it replaces the host version and everything may break. The python3 v python2.6 problem.
2) Dependency namespacing. e.g. If I want rely on a non-official mirror to ship me whizbang project X they could also replace my libc by adding it to the repo because the names and versions collide.
Making sure that we hold onto the good three properties of package management while fixing the two problems is important for Linux moving forward. The last 15 years of Linux was dominated by the centralized package management system and a ton of hacks have developed to work around it when you need a new package or want to install custom software. This is why I spend so much time working on container image specs like appc and oci; I hope we can arrive at a good container image format for the next 15 years that everyone can rely on.
For examples see slide 6 forward here: https://speakerdeck.com/philips/rkt-and-the-need-for-the-app...
I'm sorry, but you hit my pet peeve when you used python as example. Python was written in a way that you can have multiple versions installed side by side and they all can work without problem. If you have a python3 package that is uninstalling python2.6 then its author screwed something up and I would be afraid to use it at all.
> 2) Dependency namespacing. e.g. If I want rely on a non-official mirror to ship me whizbang project X they could also replace my libc by adding it to the repo because the names and versions collide.
Is that still an issue? Zypper in OpenSuSE is quite smart and for each package it remember which repo it is from. It tries to satisfy dependencies without changing vendor, and prompts if only way to satisfy dependencies is to change vendor.
That said if an application requires different version of glibc much better would be to compile it statically, but even then glibc supposed to match your kernel version so you're still risking some incompatibilities.
Which is the right thing to use is entirely dependent on your use case.
You're either being entirely too prescriptive or interchanging containers and Docker freely when they are not equivalent, though one is a particular implementation of the other.
People have been running containers that emulate a full system for quite a while (see FreeBSD jails, illumos/Solaris zones, Linux OpenVZ, LXC, etc.).
However, we also do believe the isolation containers provide isn't sufficient for multi-tenant usages. This is the main motivation behind Hyper, which run groups of container images (Pods) as Virtual Machines.
[1] https://hyper.sh
*We're in a limited beta of cloudmonkey.io, and we want to run unrelated/untrusted containers side by side securely.
It seems the main point is if there is any way to exploit code running within a container that has unfettered root access to the host system via the docker socket, an attacker would then have complete control over the host system.
Exploitation is often mitigated in layers, where if Service A is exploited, an attacker can only rwx what and where Service A has been granted priveleges to rwx. That should be as little as possible, the bare minimum access that service needs to operate. There's no reason your web server or database should be able to install new programs, create users, etc.
If Service B is running in a container and is given access to write to the docker socket, suddenly any exploitation of that service opens a door to immediately have full and unfettered root access to the host system.
> [0] FTA "... ended up making a screencast to unambiguously demonstrate the flaw in their setup..."
With LXC containers you start them as root and there is no lingering background LXC process running. Docker also starts containers as root but also has dockerd hanging around presumably so non root users can interface with it. But the container process is still running as root so dockerd seems a bit redundant and unnecesary.
This is because untill recently you couldn't run chroot as non root users and needed to run containers as root. But 'user namespaces' (> kernel 3.8) changes this and allows users to run processes in namespaces as a non root user. LXC has supported unprivileged containers for some time now [2] so you can run LXC containers as non root users, as in the entire container process is unprivileged. Docker and Rkt are working on this but its not simple to implement for container managers as non privileged users cannot access networking and mounts. But when it does presumably dockerd can run as an unprivileged process.
But Linux kernel namespaces have not been designed for multi-tenancy for instance cgroups are not namespace aware, and untill this changes in the kernel, containers will not provide the level of isolation or security required for multi-tenant workloads.
And containers managers like LXC or Docker that take these capabilities and merge them with networking and layered filesystems like aufs or overlayfs cannot work around this. Parallels OVZ is designed for multi tenancy but the kernel patch it appears is too large and invasive and doesn't look it will be merged.
So user namespaces is one level of security and isolation, you can also use seccomp, app armour, selinux or even grsec. But you have to find the middle ground between security and usability and given the relative confusion about containers, namespaces, and container managers it will take time to mature.
[1] https://www.flockport.com/how-linux-containers-work/
[2] https://www.flockport.com/lxc-using-unprivileged-containers/
I mean, I know it's popular to pass the socket in for automatically re-configuring proxies... but I haven't seen any serious use outside of that.
Doesn't really describe Docker, does it?
I know you're trying to make joke, but it's not funny.
This is a blog post about how docker is insecure when you can get root perms over a socket the docker daemon exports when you explicitly map it back into a docker container... as if it's unexpected when it's explicitly docker's design and the command line you provided.
Oh well, I did think it was funny.