Things to avoid in Docker containers
developerblog.redhat.com
developerblog.redhat.com
1. The upstream base image maintainers aren’t doing their job of keeping the image up to date, so it is huge. If the upstream maintainer had this base image updated, there would be no updates. Not the developer’s fault. Also, if the developer puts it into production, it’s their responsibility to make sure there are no CVE’s in their code, and hence their responsibility to do a yum update. You touched it, you own it.
2. The Docker tools don’t have a way of flattening the image easily (without doing a save/load). This can be fixed down the road.
Telling developers not to do yum updates terrifies me as a security person….
Layered images such as Apache and Nginx still may be maintained by operations or even Middlware teams.
Each team is responsible for making sure they update their piece of the supply chain. If the operations team (or image supplier) doesn't, the "yum upgrade" (or apt-get) downstream (by the developer or middlware maintainer) will bloat the environment.
The problem is, it's a symptom, not the problem. Telling people to NOT do yum updates, just promotes security problems because people upstream aren't keeping up with updates.
All of that said, the ONLY sane way to do all of this is with Automation. Hence, the reason for ImageStreams [1], BuildConfigs, DeployConfigs, and Triggers inside of OpenShift.
When any piece of the spider web detects a change, all of the dependent layers kick off jobs to rebuild themselves. Also, having good tests is critical. Yeah, most people on the planet haven't thought through this. Yes, Docker images make your life more convenient, at the deploy stage, with the downside that you need to do a LOT more thinking (automation and testing) at build phase.
Like shipping containers, we pack them at the factory, but can move them around at will after that. Instead of having dock hands load the ship at the port (which took 45 days in 1906). The supply chain problems are the same either way, but at least Docker gives you a good format to solve them.
[1]: https://docs.openshift.com/enterprise/3.0/architecture/core_...
Perhaps the article could have been clearer, but the advice is more subtle than it first appears. It ties into point 4 ("Don't use a single layer image"). You need to `yum update` at the right time to make good use of layering and to be able to control updates separately from builds.
Imagine you build both Nginx and Apache containers off the same distro image. Rather than having `RUN yum -y update && yum -y clean all` as the first step (which is easy to mess up or forget), you can instead build an "update" layer as a named image with it's own Dockerfile consisting solely of the `yum update`. Now your two application Dockerfiles can simply `from base-update-date`. This means you control when you get updates explicitly orthogonal to the application build and lets you run the two apps off slightly different versions, if there is indeed an issue you need time to resolve regarding backwards compatibility. This also means you get the most out of the caching mechanisms in Docker.
Full disclosure - I am both a hatter and a sec guy. [edit, spelling, grammar]
Though, I agree, it's very important to highlight that too much branching can cause massive environment bloat and way too many permutations.
If we just tell everyone to do yum updates at their layer this problem goes away. Ahh, recursion...
If I'm not using latest, and I don't know that the Apache and Nginx images have been updated, then it's very likely that "yum update" might do something.
It's a balancing act. It's not easy. But, if you use a package exactly twice in your environment, then putting it in the base image actually makes sense.
Think at scale, not about individual images.
Nothing is free, nothing. Building a cool new C library (musl) is fun. Maintaining one for 10, 20, or 30 years (glibc) is work (and fun)...
I hate ham fisted advice like this that will always end up being dogmatized and multiplying complexity needlessly.
For example, I run gunicorn inside of a container with worker processes, and I had to do some trickery (writing to /proc/1/fd/1) to redirect the logs of the worker processes to the stdout of the master process so that they would show up in `docker logs`.
With all the hoops people jump through so they can say "we use Docker, we're cool" I think I shall refer to Docker fan-boys as Olympic Gymnasts from now on.
I really really want a LOG directive for docker where I can just setup some pipes which the daemon will collect and tag into it's log files (would make dealing with old Java processes a lot easier).
But actually finding somewhere to hook this logic in - huge task. A lot harder then grabbing stdout/stderr of the init process in the container.
This exact job is what Syslog (over a network) does.
I use containers for 10 years (they are called VServer), and I newer had a problem with standard setup: some less important logs are stored directly in container and rotated with logrotate, some more important logs (and error messages from less important logs) are forwarded to central server for analyzing, some are just disabled or discarded.
But couldn't Docker provide additional file handles for extra logging streams?
Programs which understand they live in a Docker world could hook into an special system call directly. Non-Docker-Logging-Enabled apps could have a helper program (conceptually similar to 'ip netns exec') remap STDOUT & STDERR.
I have loads of containers logging to stdout and another container which collects all stdouts and send to a logging-as-a-service. Why is it a bad idea ?
Let's compare logging via stdout, to say syslog, with a simple test.
How do you identify the difference between a warning and an error or an emergency, in a string sent to stdout? What happens if the message contains a newline. How do you know if what follows is a continuation of the log entry, or is a new log entry?
If your application dares to have two processes in the same container (i know, crazy concept) how do you separate them.
The "literally one process per container" idea is what both causes the issue, and causes the obvious solution to the issue to be even harder to solve for the cases where software doesn't natively support the syslog protocol.
Anyways, this is dealt with in the logging-as-a-service (LAAS) that receives the logs... say you are using the python logger, you have different types of logging (DEBUG, INFO, WARN, ERR). If you just start the line with one of those statuses, your LAAS will know it should email you or something like that.
The print, or file, or whatever, it's just how your data is getting somewhere for the proper analysis.
You know if it's a new line or not because every log message starts with the same format, something like {APP, TIME, SEVERITY} and then the message...
> Were you the one who downvotted me because you disagree? o.O
Down voting on HN is broken[1], so I don't use it, ever[2].
1. In the sense of, it does the wrong thing, not in the sense of it doesn't do anything.
2. Except when i hit the wrong thing because of the complete lack of usability on a mobile device.
I currently have logs from the django app and nginx and was pretty easy to set up.
It might not be the best solution but I don't think it's stupid
I'd say they are for isolation before they're for scaling out
> […] will always end up being dogmatized and multiplying complexity needlessly.
If more than one process helps you get your job done, great. If it's all one self-contained application, go for it. You start feeling the pain when you have unrelated or loosely-coupled applications in the same container.
More specifically, uid 0 on the container == uid 0 on the host, so if container breakout occurs then you have root on the host.
As you mention, user namespaces and remapping the root user should solve this issue, but initial support for this was only released 7 days ago in docker 1.10.2 > https://github.com/docker/docker/commit/ed9434c5bb64f49db442...
[edit]: Preliminary support for user namespaces was added in 1.10, above commit is addressing a bug in userns support.
But why would you want to run things as root anyway? You probably didn't before Docker, so why do it now? With Docker, there's even less reason to muck about with capabilities for binding low ports etc. It's absolutely an unnecessary risk to take.
There have been and will continue to be parts of the kernel hard-coded to respect UID 0. Until these are all found and fixed (likely many years from now) using usernamespaces to remap root will not provide all the safety one might assume. It is super handy for other users though.
Full disclosure - I am both a hatter and a sec guy.
Almost no one has the resources to stand up to a targeted attack from even a mildly capable hacker. Therefore, the risk profile most administrators are working against are attacks of opportunity. In that race, as in the old joke about outrunning your friend instead of a lion, you need only be be some small amount more difficult to crack than the next potential victim.
My understanding is that the most important reason to have root in the container is to install software through the standard measures, but obviously, we don't want to have to run our build process for containers as root on the real host.
Given the comparatively restricted behavior, is this a good practice or are there implications of using a remapped root during container build time that would linger on to run time?
And I should have been clear that a remapped root that drops drops privileges and/or transitions into another user is still better than an unremapped root doing the same. It's just that the remapping is not a panacea.
If you don't use one of those, you can still have supervisor run outside of your containers, and manage a set of "docker run" commands.
Not sure if I understand correctly. I'm using ECS. How would you provision the Host with supervisord (automatedly)? ECS and the other systems are taking care of restarting the containers anyway, but they don't necessarily have any logic attached (like restart it 5 times, then stop and notifify if the container doesn't come up again).
From my POV it's definitely better to directly have it deployed within the container. Maybe it's depending on the use case.
The problem is that it often happens that the container cannot restart because the main process (invoked by ENTRYPOINT or CMD) cannot start up anymore, e.g. because of corrupted (config) data which has to be loaded on startup or other suddenly unsatisfied dependencies.
Restart policies are a thing since 1.2 https://docs.docker.com/engine/reference/run/#restart-polici...
I am fine with services. I am also fine with larger applications under duress. I think the format still makes operations lives easier. Containerize everything, is better than some things...
https://github.com/projectatomic/oci-register-machine
and
https://github.com/projectatomic/oci-systemd-hook
Also, ping Dan Walsh as he is leading work on our (Red Hat) end to get this working.
I sent email to Dan.
Example: a Django web app... after deploying the new version container you need to run migrations etc
Similarly for say RabbitMQ you may need to ensure your vhosts are defined etc... commands that need to be executed against a running server, can't be done in the Dockerfile itself.
So far I've ended up with a bash script in each container that does the necessary init tasks before exec-ing the main process. Or in some cases (Rabbit, Mongo) they first run the server in a limited state, do the init tasks against it, then stop it and exec again as the container process.
Have started to feel it might be cleaner to instead just start the main process and then provision it with Ansible afterwards. But then I lose the ability to have a system up and running just via docker-compose...
Whatever process is orchestrating your containers should take responsibility for running setup chores in a separate, ephemeral container from the eventual application container.
docker-compose is often not enough, unfortunately.
If I can keep credentials and data outside of the container, is this still a bad idea? Would I run into trouble with other containers needing to use SMTP?
https://github.com/tomav/docker-mailserver
I'm afraid I can't account for it - having never tried it, but it looks like the sort of thing you might be interested in!
PS If you are a newbie, do try something simple first - it took a while to get my head around the differences in the docker tooling (for example this project uses docker-compose which is another tool on top of the docker cl tool).
EDIT: Ooh - a tutorial
Currently setting up a mail-server myself, the majority of clients are in China but I want reliable sending to foreign servers: it's quite difficult. Like finding a host, when so many Chinese IP ranges are blanket blacklisted by western mailservers, and the China/foreign link is often slow/overloaded/lossy. Taking the time to update Gentoo postfix documentation en-route, for the other DIY'ers with an NCH (not compiled here) problem. Not using docker.
But with such image it is useful to add few tools so docker exec -ti container /bin/bash creates a shell that one can use for debugging or verification.
You could manage your own data volumes or else let docker do it for you. I don't see the difference. In the end, you will still need a backup policy. I snapshot all data every 30 minutes for storage on another device. Nothing would change to that policy when using external data containers.
> Don’t ship your application in two pieces
Internally -and externally provided software have different dynamics. Issues are different. I update my own software with bug fixes and new features quite frequently. I don't need to do that for externally provided software. If that happens, I will indeed rebuild the container. In all practical terms, my own software sits in a host folder where I can update it, without rebuilding the container.
> Don’t use only the “latest” tag
I use debian:latest. I don't see a problem with that. It may lead to trouble some day, but then I just change the tag to the latest but one version before rebuilding. The problem has not occurred up till now.
> Don’t run more than one process in a single container
I have containers that internally queue their tasks. The queue processor is then a second process, besides the main network listener process. There are many reasons why it could be meaningful for a program to use more than one process. A categorical imperative in this respect is misguided. Other engineering concerns will take precedence.
You see, if I can reasonably split a container into two, I will. It's just like with a function. If it is possible to split it in two, I will most likely do so. But then again, such design recommendation should never be phrased as a categorical imperative.
> Don’t store credentials in the image. Use environment variables
Both are pretty much the same problem. If the attacker can read files containing credentials, he will also be able to read environment variables with them.
> Containers are ephemeral
A network-based service uses a long-running listener to process requests. Why shut it down? That would just disrupt the service. Containers may very well be long-lived. They are not necessarily ephemeral.
This subject is not part of a domain such as morality where categorical imperatives are the norm. There are pretty much no categorical imperatives in software engineering. Software is mostly subject to just an Aristotelian non-contradiction policy. If your choices are non-contradictory, feel free to go with them. Furthermore, unmotivated, categorical imperatives simply have no place in this field.
It doesn't matter how long they are up, it means you should expect them to go down at any moment and lose all data stored inside. Plan for it by storing your data and config somewhere else.
* Custom volumes and system mounts are persistent.
* Images are persistent-ish, they get replaced by updates.
* Containers are ephemeral, intended to scale up, down, sideways or whatever.
> Don't store data in containers
As stated above, you should expect your containers to puff out and into thin air at any moment.
Also remember that docker auto-created volumes get auto-deleted when no more containers are using them. Use custom created volumes or system mounts if you don't want to wake up to a nasty surprise.
> Don’t use only the “latest” tag
It will bite you in the ass if you do. It may not today, nor tomorrow, but it will some day, the exact day that a bigger version upgrade happens. It's not a "hasn't happened to me yet, so it may not happen" thing, it will happen to you if you keep using the "latest" tag, no doubts about that.
So just plan accordingly and don't use the "latest" tag (for deployment).
> Don’t run more than one process in a single container
This could be more of a "really try not to" than a "don't", but the truth is the more stuff you put in a single container, the harder it becomes to manage and scale. It is much better to have self-contained containers (see what I did there?) that you can match and mix with stuff like docker-compose.
> Don’t store credentials in the image. Use environment variables
No, they are not the same problem, and it has nothing to do with security. When you want to scale and spin up a new container... now you can't because your credentials are baked into the image.
Don't store credentials in the image, period. If you don't like environment variables, you could mount a per-container config dir, but it's usually much easier to manage config passed as environment variables to the container.
Another reason - once image is built - why rebuilding it? There is always a chance that you will end up with different versions of os or your packages. Immutable image which doesn't store credentials can be used on multiple environments with different credentials / configurations without changing / rebuilding the image.
FROM image:staging
COPY ./production-credentials/ /
When implemented in this way, staging image won't be able to run in production code until it will be tested and accepted (sealed) using formal process.If you will use environment variables, you (or your deployment team) will be able to run staging images in production at any moment.
You're essentially treating a container like a vps, so why not just run your app directly on a vps?
That's not the advice he was giving. He was talking about tagging the images you create. If you always use latest, you're not keeping a history of images in your registry that allows for easy rollbacks. Latest is an entirely appropriate tag to use so long as it's always a synonym for some other longer-lived tag.
> There are many reasons why it could be meaningful for a program to use more than one process.
There are, but there are people (like Phusion) that want to run lots of processes inside a container. That is, IMHO and that of the story author, a bad idea. I may be more lenient than the author, but I believe a single container should represent a single logical process, whether that makes to a single OS process or not. But when you have many distinct logical tasks, they should be separated into their own containers.
> Both are pretty much the same problem. If the attacker can read files containing credentials, he will also be able to read environment variables with them.
This is very, very wrong. Credentials stored inside the image are less flexible and stored in more places than credentials that are supplied as environment variables. Images are designed to be copied around easily and trying to add authorization to the copying process is never going to be particularly successful. In contrast, credentials supplied as environment variables only live on the machine that's running the container and are only visible to those with access to the machine.
> A network-based service uses a long-running listener to process requests. Why shut it down?
Because software is constantly changing. At my last job, we released many times per day. Trying to patch a running container is just a recipe for pain. It's much simpler to start new containers and kill the old ones. You can even do it without downtime if you've got a system to route to the correct container (something like a long-lived HAProxy process that watches etcd/consul for the container it should route to. More and more, the notion of a non-distributed system is becoming obsolete. And the one lesson I've learned about distributed systems is that state is what makes them hard and you have to be incredibly intentional about how you manage your state. This also means that any time you can avoid adding state, be it in code or infrastructure, you should avoid doing so. Ephemeral, immutable containers follow that philosophy.