Dockerfile Best Practices
github.com
github.com
There's no mention of multi-stage builds, labelling images, using env variables to pass ephemeral config settings, setting up tmpfs folders to avoid cluttering containers with temp data, the importance of rebuilding images periodically... In short, the basics.
Edit: this stuff is already covered by Docker in it's official Dockerfile best practices guide.
https://docs.docker.com/develop/develop-images/dockerfile_be...
I was expecting far more from an article on HN about Dockerfile best practices, particularly one which is currently being listed so prominently. I mean, does the HN crowd need a top ranking link to be reminded that we should not be running external-facing services as root?
This info is quite literally made available by Docker itself in the form of it's official best practices guide, which exists for years now.
https://docs.docker.com/develop/develop-images/dockerfile_be...
How is anyone publishing a so-called best practices guide when they aren't even aware of the stuff covered by the official best practices guide?
:-)
I see the point but this kinda looks like extreme paranoia.
I actually advise developers that are not so strong at docker and containers to use his 1000 in their container. Such uid is usually the default first uid in Ubuntu and thus makes running containers in their development machine easier (since now they don't have to deal with file permissions and uids without a corresponding user on their system).
Someone will boo this and maybe even downvote, but it really helps developer "think operationally", and gets their dev environment closer to prod.
At that point, you're gonna to need to change your strategy and properly map the user/group outside the container to the user/group inside the container anyway.
That allows an attacker to potentially append some nasty code outside the container and get password-less sudo. Unlikely? Very. But security is all about layers.
I've seen full-fat Ubuntu containers with uid 1000 which is in docker group on the host and ~ mounted to run some python flask app, in production.
So then running your containers as that UID without user namespacing (docker’s default) opens you up to more attack surface than if it was uid 1001.
If you mount ~ in your containers you have bigger problems than uid 1000.
This is probably wrong and lazy. I think the rigorous approach is to use namespaces
https://www.jujens.eu/posts/en/2017/Jul/02/docker-userns-rem...
https://www.objectif-libre.com/en/blog/2020/06/30/securiser-...
> Writing production-worthy Dockerfiles is, unfortunately, not as simple as you would imagine. Most Docker images in the wild fail here, and even professionals often[1] get[2] this[2] wrong[3].
And yet, in some of the linked URLs, people are presenting reasons for why that approach was used in particular, instead of the supposedly safer alternatives.
For example, https://github.com/caddyserver/caddy-docker/issues/104
> We actually originally did run as non-root by default, but simplicity we decided to drop that (see #24, and also #103 for some other related discussion).
> If your Dockerfile works for you, that's great. In most cases where users want to run as non-root, they also don't need to listen to :80/:443 in the container, so the setpcap magic isn't necessary at all.
> It's also worth noting that caddy is an official image, and as such needs to be similarly-shaped to other official images of the same type. At a quick glance, none of nginx, traefik, httpd, or haproxy support running as non-root out of the box either.
> Finally, it's worth considering why you want to run as non-root. What attack vectors are you trying to avoid? Container escape vulnerabilities are pretty much the only real risk, but anyone running a modern Docker version is immune to many of them. It's also worth considering user namespace remapping as a mitigation. In my experience the main reason for running as non-root is to pass compliance checks - not a bad reason, but it's also worth recognizing that non-compliance does not automatically equal decreased security (and vice-versa).
If larger projects, like Nginx, Traefik, Httpd and HAProxy were all creating containers like that, it makes you think about the reasoning behind it. Is it easier to just run containers as root and not worry about the permissions inside of the container? If so, wouldn't it really make more sense for the container runtime to have some sort of mechanisms in place to allow people to do what's easy within the containers while also making sure that it has no harmful impact outside of them?
Because to me it seems like people will continuously take the path of least resistance and from where i stand, it should be up to the creators of the container technologies to make sure that this path is safe by default.
That does not sound like an adequate excuse, does it?
I'm afraid none of your remarks makes a reasonable point. Even if you believe that running random stuff as root is ok, if you do not have any reason to do that then why should we mindlessly follow bad practices? If we aim at running a safe system.and there is absolutely zero drawback in following best practices, then why should we continue to make the mistake of intentionally using poor practices?
Also, all the sentences that begin with '>' were comments from https://github.com/caddyserver/caddy-docker/issues/104, and not from the user gbrindisi
If the official Docker image did that, then most users would be very confused and we would get lots of support complaints. A cost-benefit analysis told us it was not worth the headache to run as non-root since the Caddy project values highly user experience.
You either specify the user when creating the container (which can have certain implications, like being able to bind to < port 1024... solveable but still something to deal with) or you can for everything to run with a uid mapping such that uids in the container are mapped to higher uids on the host.
For example, look at SSL/TLS certificates - before Let's Encrypt ( https://letsencrypt.org/ ) and tools like either Certbot ( https://certbot.eff.org/ ) or even web servers like Caddy ( https://caddyserver.com/ ) which automate both provisioning and renewing certificates, people used to simply run HTTP. But now, it's easier than ever to use them for transport level security, and the stats seem to vaguely back this up, for example: https://www.welivesecurity.com/2018/09/03/majority-worlds-to...
Why should users inside of containers be any different? What are the factors that prevent safe defaults from being implemented? That's what i don't understand.
Disclaimer: i'm not advocating for running things as root, but rather my claim is that if things are hard to do, they simply won't be done unless absolutely necessary. Any tech vendor should acknowledge this and make sure that doing things the "right way" is as easy as possible.
The issue is complexity and lack of appropriate abstractions.
The "Let's Encrypt" of dockerfile safety would be something that makes it trivial to 1) create a user 2) chmod/chown an appropriate spot on the fs, 3) ideally let the author defer these actions to always finalize the image in user mode. That way, you declare at the top that you will do these user actions, but RUN stays as root. Or just provide an SRUN, SCOPY,SADD directive which acts like running with sudo. Then, you can easily extend a layer or base image without being concerned with the details of how user space is implemented.
Also there is no standard or idiomatic protocol for setting up user space in a dockerfile.
But the argument is, should it be a bad practice?
Are you really asking if running external-facing services as root should be considered a bad practice?
Maybe not a bad practice, but it should be redundant.
This is correct in principle, but very hard in practice. This is because kernel support for containers were kind of "tacked on" and more-or-less scattered across the code-base. And although they're getting a lot better, there's still no easy way to reason about their security. So a lot of the advice around permissions management and access control are a kind of defence-in-depth.
I'd love it if we get to a point where containers can make strong statements around security.
I don't like the idea of running my (internal, not meant to face the internet) containers so that they are publicly accessible or running their processes as root. Yet this seems to be the default with docker and I'm sure a lot of people don't bother fixing that.
There is so much confusion regarding users and user namespaces. I think something needs to change in the way docker documents those things and also in the way defaults are chosen for various configuration options.
Containers are just namespaced processes that share the same kernel as the host. A host has access to all container processes, uids, gids, file systems, and networks. Cgroups are used to limit resource access.
To run containers securely you need to understand how to protect running processes. You need to use unprivileged users where possible, drop all kernel capabilities not required, run Linux Security Modules (AppArmor, SELinux) to prevent processes from doing things they shouldn’t; and, run containers based on the smallest image possible, since a container should only have files that are absolutely required to run a process, and nothing more.
Even when you do it all right, in a multi tenant environment, it’s not safe to run all containers on the same hosts.
The point about multi-tenancy is absolutely understandable. Isn't this an old story from the PHP world with multi-tenancy? I think a good generalization is: don't run on multi-tenant systems if you do anything (!) critical (e.g. authentication or payments)?
But that of course disregards the fact that when people _can_ do something, they _will_ do it even though they shouldn't (like running E-Commerce systems in multi-tenant environments).
Another thought regarding isolation: aren't VMs essentially just running on one host as well? Is that why you said "VMs are _more_ isolated"?
But there are other vectors. With a VM you get a whole linux distribution, which of course increases the attack surface, but at the same time you also get much better isolation and that distribution's team of maintainers looking over your software, providing security patches, advisories, a simple way to update the system and so on. On the other hand there exist 'docker best practices' tutorials (not the posted one) that recommend not updating your base system at all in the name of reproducibility. Docker's solution to update management is manual image tagging and manual updates, possibly with help of external tooling. I don't think that's a good solution for that problem.
Imo the overall best solution is to run stuff in VMs and pick a lightweight distro for that.
That not updating part is of course just plain and simply bad advice.
What solutions for update management would you recommend in the VM space?
It is more secure to run a process with seccomp filters than it is to run it without.
It is more secure to run a process with seccomp and vm isolation than just one or none of these.
Running as root is lazy and equals container escape, especially when running on anything other than scratch and read only file system.
The only reason Nginx and Traefik run as root is to bind to privileged ports (80,443). There is no reason to do that inside of a container, since you can remap exposed ports outside of the container.
Containers are not VMs and must be handled differently. You are always one RCE away from having your entire container platform compromised.
If the container is running as root permitting it is redundant, since the kernel doesn’t filter root for kernel capabilities anyways.
If a privileged user sets CAP_NET_BIND_SERVICE on an executable binary using setpcap to allow a non-root user the ability in a container to bind to a privileged port, elevated privileges are still required for execve to create a process that is permitted to use the kernel capability. Think sudo but for processes.
The argument with containers is that binding to a privileged port isn’t necessary, so you shouldn’t do it. And by not doing it you improve your security posture.
Docker compose has the key word as well.
No. From Docker's official reference:
> The main purpose of a CMD is to provide defaults for an executing container. These defaults can include an executable, or they can omit the executable, in which case you must specify an ENTRYPOINT instruction as well.
https://docs.docker.com/engine/reference/builder/#cmd
ENTRYPOINT should point to the entrypoint, and CMD should store default command line arguments. This allows a container image to be executed as a command line application.
[Edit] This is what I use for local development docker images.
Note: I disagree with this article, none of the things listed are "best" practices in my opinion... more like random practices somebody prefers. Use what you like.
> UIDs below 10,000 are a security risk on several systems, because if someone does manage to escalate privileges outside the Docker container their Docker container UID may overlap with a more privileged system user's UID granting them additional permissions.
> [...] there may sometimes be reasons to not do what is described here, but if you don't know then this is probably what you should be doing.
Your version:
> Use what you like.
> UID=1000
Surely, if "X" has posed a security flaw in the past, and "Y" achieves the same as "X" without being vulnerable to the exploits (e.g. the UID advice), then "Y" is objectively better and not a matter of style, is it not?
https://docs.docker.com/develop/develop-images/dockerfile_be...
The basic idea is you create a user in your Dockerfile, switch to that user with the USER instruction and now future instructions in your Dockerfile will be run as that user.
Also when COPY'ing you'll want to add --chown myuser:myuser too.
The above Dockerfile shows examples of all of that.
I'm not a fan of customizing the UID / GID because then in development with volumes you can get into trouble. Technically you could set UID + GID as build arguments but in practice I never ran into a scenario where this was needed because 99% of the time on a dev box your uid:gid will be 1000:1000 and in production chances are you are in control of provisioning your VPS so your deploy user will be 1000:1000 too. Also you probably won't be using volumes, but if you did they will work out of the box with the above example.
If you are interested I’ve collected more best practices to prevent more common security issues: https://cloudberry.engineering/article/dockerfile-security-b...
...and that's why no one uses it that way, and instead uses Docker's built-in support for healthchecks.
Check out the HEALTHCHECK entry in Docker's reference
Note that adding the healthcheck isn’t enough. Docker won’t actually do anything if a container is unhealthy. You need a seperate process to restart the unhealthy container.
I'm not sure you fully grasp the issue or understand how Docker works. Docker's healthchecks are not "container measuring its own health". Docker's healthchecks are a standard interface that was designed to allow container orchestration services to poll containers to check if they are still in working order.
From your own description, it sounds like you tried to reinvent the wheel, and did it poorly.
And I'm sorry to break it to you, but if you have developers faking health checks in production then your choice of container runtime or container orchestration system is not the problem you need to worry about.
> Unfortunately, although Docker did add it natively, it is optional (you have to pass --init to the docker run command). Additionally, because it is a feature of the runtime and e.g. Kubernetes will not use the Docker runtime but rather a different container runtime it is not always the default so it is best if your image provides a valid entrypoint like tini instead.
I have noticed that after moving away from tini, some of my containers take a lot longer to shut down, despite the init flag. I think this might be related.