Docker Security Cheat Sheet
cheatsheetseries.owasp.org
cheatsheetseries.owasp.org
* minimizing the difference between `docker restart` (preserves overlayfs changes) and container re-creates (resets overlayfs back to the image state)
* surprise data loss on container redeployments because data was unexpectedly being written to the overlayfs instead of a volume
* unexpectedly running out of disk space in `/var/lib/docker` because data was being written outside of a volume
* performance issues caused by excessive overlayfs writes (storage drivers and /var/lib/docker not necessarily designed for IO performance)
$ systemd-analyze security service-name
It prints out a long list of hardening flags that can be applied inside of your service, like so: NAME DESCRIPTION EXPOSURE
PrivateNetwork= Service has access to the host's network 0.5
User=/DynamicUser= Service runs as root user 0.4
CapabilityBoundingSet=~CAP_SET(UID|GID|PCAP) Service may change UID/GID identities/capabilities 0.3
CapabilityBoundingSet=~CAP_SYS_ADMIN Service has administrator privileges 0.3
...
Here's what I typically use for a .NET 5 application: WorkingDirectory = /opt/appname/app
ReadWritePaths = /opt/appname/data
UMask = 0077
LockPersonality = yes
NoNewPrivileges = yes
PrivateDevices = yes
PrivateMounts = yes
PrivateTmp = yes
PrivateUsers = yes
ProtectClock = yes
ProtectControlGroups = yes
ProtectHome = yes
ProtectHostname = yes
ProtectKernelLogs = yes
ProtectKernelModules = yes
ProtectKernelTunables = yes
ProtectSystem = strict
RemoveIPC = yes
RestrictAddressFamilies = AF_UNIX AF_INET AF_INET6
RestrictNamespaces = yes
RestrictRealtime = yes
RestrictSUIDSGID = yes
SystemCallArchitectures = native
ProtectProc = invisible
CapabilityBoundingSet =
SystemCallFilter = ~@clock @module @mount @raw-io @reboot @swap @privileged @cpu-emulation @obsolete
ReadWritePaths should be replaced with a combination of DynamicUser + writing local persistent data to $STATE_DIRECTORY, but I'm too lazy to do that yet.See systemd.exec(5) for more.
I hear you but fighting systemd in 2021 is like pushing water uphill. With a fork.
I have ported a few docker-compose.yml-s this way, because I don’t understand docker enough to troubleshoot issues with getting my firewall rules to apply to docker traffic. Dependency hell is not an issue for these projects, so I feel happier with systemd.
The isolation provided by unit files is orthogonal with running containers.
nspawn is not even running a dedicated daemon. Plus, it's no secret that docker was not designed with security in mind and its isolation is bolted on. [1]
Furthermore, systemd is already installed and running on most systems (like it or not)
When I build new images and it fails because the pinned version is not available anymore, I have to dig into Debian or Ubuntu packages websites to find the new ones as they don't keep the old packages online.
I know I could ask Hadolint to ignore this rule but I don't like this and I think it's important to stick to a certain version of a package to avoid problems. I'm just trying to find any tip that could make me use pinned version and avoid this manual search every time I have to. Does apt-get install allows wildcard for example?
[1] https://nicolasbouliane.com/blog/nextcloud-docker-upgrade-er...
You can do `apt-get install my_package==4.1.*`, for instance and that will pass the validation and still be close to the ideal of having reproducible builds.
It comes with a very handy tool as well https://github.com/docker/docker-bench-security
I wonder where all this complexity ends. If a human can't fully grok the systems we work on, then there is no way we can hope to not be misled and taken advantage of. Does anyone else share these concerns?
But do consider the list when some junior says they can get the company billing system running on Docker on prod.
Try SELinux and see if you think this is still complicated ;)
ADD vs COPY was new to me, though.
...More in general, I think it’s not the systems that are complex but our minds/brains that are limited to deal with it. Just look at the complexity in nature.
If you're only using docker for development to bundle things in a different OS base image then there's not that much need for paranoia. You probably would trust the other OS vendor with your host system too. Running docker without security may be perfectly fine in that case.
For random scripts from github running them in docker is mostly a precaution so they don't screw up your host system or some malicious dependency makes a low-effort attempt at exfiltrating your ~/.ssh dir or whatever. In that case it makes sense to understand the basics, e.g. not just running docker --privileged willy-nilly just because some README asks you to.
If you're running some PaaS container service where users can execute arbitrary code on shared hardware then you better have a full-time security engineer or two who will be paid to understanding all the details involved.
Yes, linux security is complicated. So why not work on improving this instead? Optimise and simplify what we already have.
Docker hasn't made the complexity of linux security go away. It has just added a whole other dimension of potential security issues people now need to manage in addition to the security of the base system.
Programmers need to shift their mindset. We already have far too much complexity. Stop thinking about what new things you can create, start thinking about how you can improve and simplify the software that we already have.
"Run something with no -- or very specific -- network access" was a really annoying problem to solve (LD_PRELOAD?) in the before times.
>Run something with no, or very specific, network access
If this was such a problem in linux then why didn't people focus on improving this instead? We could have solved this problem in linux and made the whole system better for everyone.
Instead people left the problem there and just piled more crud ontop. We added an entire new layer of abstraction that everyone now has to spend weeks learning how to use instead of just fixing the original problem. The whole system is now far more complex, and the original problem is still there.
The art of software engineering is about managing complexity. The best way to manage something is to reduce the amount of it you have to worry about. For the rest, paying really close attention, accepting ownership and directly engaging is key. Reducing the number of parties you have to trust is typically a happy side-effect of reducing complexity.
I look at containerization as an attempt to hand-wave away critical responsibilities around ownership of complexity in an application. There are tons of examples that illustrate this throughout the ecosystem, but I think this one is most apt -
One day you discover your application is a difficult mess to reconstruct from source each time. You have reached a fork in the road because management is complaining that it takes 2 weeks to configure a new QA environment from scratch. Do you either:
A) Review fundamental assumptions about the problem domain, technology choices and teamwork. Potentially consider rewriting your application from scratch using fewer tools & computers, and with more focus on driving the actual business value equations.
Or,
B) Decide that Tom's computer is the new golden image of production and bless it as such. Now let's find a way to manage a whole farm of these things!
And the Dockerfile format is a good way of capturing what it actually takes to reconstruct your application/environment from source in a manageable, version-control-able text file.
I do agree that the rush to containerization is generally a rush to draw expansive abstraction boundaries, but this particular story isn't what people are doing. (At least not with containers - it's certainly what people were doing with VM images ten years ago!)
Linux is not very good at security boundaries anyway, just run one thing in each VM and don’t leave anything else there to privilege-escalate into.
No tech is perfect, but docker has dramatically reduced the surface area of tunables that my application devs have to care about in order to get our product shipped. Even if it's not brought it down to zero, it's a step in the right direction in terms of UX complexity exposed to the developer.
Host-users and guest-users must be explicitly mapped by whomever starts the container, so this issue is not a security threat to outside of the container. That said, if a guest-os is running as root and then someone compromises it, they have at their disposal all the powers of root in that container.
OWASP has also mixed good container guidance with "Docker" as well as some Docker specific things, which I'd like to maintain a delineation with now and into the future because they are different. I understand most uneducated people are looking for "how do I do x in Docker" but subtly educating the user matters for broader discourse.
I understand the motivation of marking scripts and binaries as non-writable inside the container as an extra layer of assurance (along with a non-root user that can only execute). But it’s a disservice to developers if you don’t explain why. A lot of people walk away from this thinking they’re protecting the host OS and wind up cargo-cult-creating a container user with full write/execute permissions.
https://chiselapp.com/user/rkeene/repository/bash-drop-netwo...
Of course, if you're running containers in a ephemeral VM, no biggie.
If your docker is running in a private network and someone got in, it already means your whole system is compromised.
If you exposed application within a docker that has code0 vulnerability, it really does not matter, your system is already compromised..
Know what you are doing!
Don't just run an unknown public container without checking it's Dockerfile first. If you don't understand what's happening, chances are high that you don't know what you're doing.
Correct me if I'm wrong, but by default, containers can't communicate with each other even if ICC isn't disabled because the daemon gives them unique "default" networks. Only if you specify the same network for different containers can they communicate...
Edit: this behavior is specific to docker-compose. If you do docker run without specifying a network, it does use the docker0 bridge.