Do we? I'm using them, but so far I'm not that impressed.
Do we? I'm using them, but so far I'm not that impressed.
(Five point releases after promising this wouldn't be the case anymore, Docker still doesn't use user namespaces, so I really hope your Dockerfiles aren't using root! Have fun.)
(I haven't ruled out that it might be a lack of a culture of security, too, on the Docker team's part; there was that whole "oh, yeah, put desktop applications in Docker containers, never mind that now you're running Chrome as root and putting the unprivileged X socket in the container, letting it pump messages to any other application" thing from a Docker core contributor. Maybe they just don't know, themselves? Woof.)
If you feel there are further actionable steps to move best-practices into default behaviors in Docker today, or improve security please feel free to submit a PR or open a thread on the mailing list. (If it's a vulnerability, however, email security@docker.com)
Definitely, but, in my experience, it gets worse than that, and why this is really important and not just important. Most containers are built on top of full OS images, with a full OS image's worth of vulnerabilities (some inert because services aren't running, but setuid's still respected--and I wonder how many setuid'd tools are actively tested in a containerized environment!), and it's rare in my experience that those base containers actually get updated after the end user selects one and a version to go along with it. Which means that even a properly-built container that doesn't run its internals as a non-root user may--and a pessimist might say "probably does"--have user escalation bugs squirreled away somewhere that the user not only doesn't know about, but isn't capable of auditing or is even aware that they might exist.
This, not "users don't read docs", is why I am so very, very salty about this in Docker, and while I'm glad that now Docker users (and, full disclosure, I'm not one anymore in part because of the poor security story--monolithic do-everything daemon running as root, no user namespaces--and in part just not wanting to be the stooge for somebody's platform play) don't have to "expect to wait long", this was promised in something like Docker 1.4, a year ago.
This is one of the problems DockerSlim (http://dockersl.im) is trying to address. You take those containers built on full OS images and you remove everything your app is not using reducing the attack surface.
Have you considered using rkt as an alternative to docker? It tries to avoid a lot of the security related problems docker has. Specifically there's no daemon which runs as root, instead rkt is invoked directly, so the only time rkt requires root, is the actual execution of a container, not downloading or verifying for example. Next, rkt by default wont run unsigned or untrusted images meaning by default you must trust the image/author, and finally rkt has the ability to run your container in a light weight VM using lkvm, getting you all the benefits of VM level isolation, but you have the ability to use the same container tooling for it all, and decreasing the overhead by optimizing for the container use-case.
Recently rkt also got support for logging different events into the TPM audit log, making it possible to have tamper-proof audit trails of what containers have run on your system. User namespaces are also implemented, but I'm not sure how well tested they are.
Solving the problem of using, and creating large full OS images is quite difficult. As a stepping stone we have also created a tool which scans container images on our Quay.io registry looking for images effected by CVEs. This should hopefully help until we can properly solve creating functional minimal containers easily, but that's unfortunately not quite as easy as just telling people to use buildroot (you still need a way to update those images and know when apps need have security updates).
Hope this helps.
I have, and I respect it a lot more--and CoreOS in general has struck me for quite some time as being the adults in the room in this space, I think the value prop remains very shaky but I appreciate that CoreOS seems to give a damn--but I don't have a lot of use for rkt in general. The only place I have any multi-tenancy, I'm using FreeBSD and jails. (I'd like to use Illumos or OpenIndiana, as I better understand that stack than I do FreeBSD, but cperciva has done a great job with FreeBSD AMIs on AWS and I don't have the bandwidth to maintain an OpenIndiana AMI.)
Sure, you don't need to have the process running as root in the container, but you need to have root-equivalent access to start the container. For a (large) subset of use-cases, this just isn't an option.
I should be able to allow any given user of a system the ability to start a docker container and be confident that they won't be able to break the host system. Until this is the case, you don't have security in Docker.
[0] https://gnu.org/software/guix/manual/html_node/Invoking-guix...
Seccomp and user namespaces are in the Docker experimental build (https://docker.com/experimental) and should land in 1.10.
Docker 1.9 supports all other namespaces, cgroups, pivot_root, cap drop, selinux, apparmor, uid/gid drop.
You can also sign and verify all images with a built-in Notary/TUF implementation, and we partnered with Yubico to support hardware signing out of the box. I'm hoping we can make image signing the default in the near future, and make it mandatory within the year.
At this point I'm comfortable saying Docker's security story is strong (although not perfect of course). But if you have specific suggestions for improvements we are interested!
EDIT I got it wrong userns is still experimental and will land in 1.10
But yeah, I guess we should totally fund container startups, because… something.
(They seem good enough for local development, at least.)
To be honest, I don't think containers are even particularly good for local development. I'd rather run something that acts like my production environment, which means a virtualized system that's version- and patch-equivalent to my prod environment. I use Vagrant and scaled-down virtual machines for that.
Then if you throw nanobsd into the mix you can create server images that are read only except for the application containers. Then you have a single server image for your application already setup, that you can just boot from or upload to some cloud service.
And now that FreeBSD has a 64bit linux emulator and docker ported everything just get's better.
[1] NanoBSD servers: https://2010.asiabsdcon.org/papers/abc2010-P4A-paper.pdf
[2] NanoBSD servers: https://lwn.net/Articles/387405/
[3] Docker: https://wiki.freebsd.org/Docker
Containers democratize dev environments.
This is different from running the Docker from a user account. With Docker you are interfacing with the Docker daemon from a non root account but the container is still running as root. With Unprivileged containers thanks to user namespaces containers are launched and run from the user account.
But unprivileged containers need to be a simple no fuss experience and this can only get addressed in the kernel. On top of that the most popular user land container implementation Docker chooses to run containers without an init and since most apps you want to run in a container are not designed to work in an init less environment and will require daemons, services, logging, cron and when run beyond a single host, ssh and agents, just managing the basic process of running apps and their state adds tons of additional complexity for users. Integrating user namespaces on top of this is going to be non trivial.
Contrast that with LXC containers which have a normal init and can manage multiple processes enabling your VM workloads to move seamlessly to containers without any extra engineering. Any orchestration you already use will work obviating the need for reinvention. That’s a huge win but if you listen to the current container narrative and the folks pushing a monoculture and container standards it would appear there are no alternatives and running init less containers is the only ‘proper’ way to use containers, never mind the complexity.
A lot of problems related to unprivileged containers are in kernel namespaces and cgroups which are not going to be solved in user land. Cgroups are not namespace aware and only root users can manage them. Access to resources like mounts and networking require privileges and that won’t change. These are probably not 'sexy or fundable’ so the problems remain unsolved, and instead complexity deriving from niche use cases whether security or micro services is foisted on everyone.
A container is just a Linux process in its own namespace that anyone can create with unshare, iplink and chroot or pivot root. An ‘immutable container’ is nothing but launching a copy of a container enabled by overlay file systems like aufs or overlayfs, a ‘stateless’ container is a bind mount to the host. Using worlds like stateless, immutable or idempotent just obscures simple underlying technologies and prevents wider understanding of core Linux technologies that need to be highlighted and supported. But we choose to focus on wrappers and the narrative becomes about funding and standards. Without more focus and support on the filesystems, namespaces and other critical enabling technologies end users will not get a consistent experience. How much support do they have beyond the occasional article on LWN? These devs and projects work in obscurity with little support. This does not seem to be sustainable development model.