Docker container security cheat sheet
blog.gitguardian.com
blog.gitguardian.com
Even if you use a cloud firewall it's worth avoiding 8000:8000 for the sake of being explicit with your intentions.
The reason you'd want to avoid that is because you'll probably have your services reverse proxied by nginx, in which case only 80/443 need to be published because the internet will be hitting nginx, not your internal service at example.com:8000 or whatever port it's running on.
This topic and many other security gotchas / best practices were in my DockerCon talk from a few months ago at: https://nickjanetakis.com/blog/best-practices-around-product..., it goes over patterns how you can use a more restrictive and secure 127.0.0.1:8000:8000 value in prod but still use 8000:8000 in dev so you can check it on multiple devices on your local network, all with the same docker-compose.yml file.
[0] https://news.ycombinator.com/item?id=27359081
edit: wrong link
https://github.com/rootless-containers/rootlesskit/tree/v0.1...
Otherwise for general dockerized applications, you won't notice any difference.
You may find some quirks, but these can all be worked around easily as described on the rootless docker page.
We run it in production with no issues so far.
`sudo sh -c "echo 1 > /proc/sys/kernel/unprivileged_userns_clone`
Basically user namespaces let you do the same thing as pid or network namespaces, but with users. This means you can have a "fake" root user. The Linux kernel, marvel of software that it is, is easily confused by this and basically hands you trivial privescs if you're this "fake" root user.
This problem is pretty much only getting worse because user namespaces are becoming more powerful whereas kernel security is staying the same (ie: not moving at all).
That's why it's gated in most distros by default.
Not sure about most distros gating that sysctl. Ubuntu works fine with rootless Docker with no changes and looking at their install instructions, there's only mention of setting that sysctl on debian and arch.
Yep. I just was answering the question, which is what the tradeoff is.
> Not sure about most distros gating that sysctl. Ubuntu works fine with rootless Docker with no changes and looking at their install instructions, there's only mention of setting that sysctl on debian and arch.
Interesting. I haven't checked in a while. I'm also not on the latest debian though.
Debian used to gate them behind a sysctl, but that's changing in the upcoming Bullseye release:
"The previous Debian default was to restrict this feature to processes running as root, because it exposed more security issues in the kernel. However, as the implementation of this feature has matured, we are now confident that the risk of enabling it is outweighed by the security benefits it provides."
https://www.debian.org/releases/bullseye/amd64/release-notes...
Ubuntu has allowed user namespaces by default for years. Which distros are still holding out?
But am I totally wrong?
managing k8s yourself on bare metal is hard. Managed k8s on any provider is a real value.
I'm more concerned with moving to Docker images, with the associated increase in needing to manage the whole container, vs. Lambda/S3/DynamoDB. I'm especially worried about misconfiguring the images and suddenly some security threats that were handled by Amazon's services are suddenly my responsibility and I fail.
I'm concerned with foregoing AWS's security hardening and redundancy more than global deployments and alarms (at the moment)
The first reason is that it's all running on Linux anyway, and Linux is (in general) swiss cheese, security-wise. Even with a billion container tweaks, there are still holes that can be exploited from the container to escalate to the host OS.
The second is that attacks don't need to privilege-escalate to the host to cause havoc. If the attacker can read memory, they can get credentials for other network services and exploit them from the container. Or they can drop malware from the container to any users of a service, or upload it to a service. Or they can just scan the network looking for another vulnerable service. Or it could be something like EC2/ECS Metadata Service was left accessible and they can start enumerating your cloud account(s). More than enough for the average attacker.
Just assume that a Docker container is exactly the same as running a regular process on the host OS, and it will be much simpler to identify attack surfaces and mitigate them.
I feel like i can understand this point of view, since following all of that advice indeed would be cumbersome. However, at the same time you definitely have to consider what it is that you're running on your infrastructure. A small internal system or even an ERP that's not exposed to the internet will probably give you more leeway in regards to being able to ship stuff now, without spending bunches of time locking everything down, especially if you build all of the containers yourself. On the other hand, a large finance application that is publically accessible and needs to weather thousands of attacks daily will probably need a rather different approach.
Overall, however, i'd say that it's good to have lists of tips like these, because figuring out all of it alone would take a whole lot of time. That said, even in the more relaxed environments, it's generally a good idea to consider at least some of them, for example:
> Unless you are very confident with what you are doing, never expose the UNIX socket that Docker is listening to: /var/run/docker.sock
Being an early adopter of Docker, i once made this mistake on a throwaway VPS. It took less than 24 hours for it to be mining crypto. That said, the socket can be a good option for tools like Portainer (which implementations like Podman miss out on), yet it definitely should never be exposed publically.
As always, security isn't a boolean of on/off, but rather is a sliding scale of sorts - figure out the risks that you're likely to be facing and choose the appropriate means to combat them. Of course, it would be better if Docker provided safer defaults, too.
Is the isolation provided by Docker worth nothing, from a security perspective... also no :)
Hardening containers is a good element of an overall security strategy. It needs to be combined with other controls, both preventative controls at things like the network layer, and detective controls to spot when a preventative control fails and allow for rapid mitigation.
It really depends on the context. Docker already supports defining env cars in container images, so it makes no sense to sneak a .env file into a container image. If all you're doing is setting env cars locally to run a container then if those env cars don't include secrets then it's pretty safe. However it would be preferable if those env cars are handled by the container orchestration system. For instance, docker compose files also support specificing env variables, as well as Kubernetes.
However, if the container doesn't contain any processes running as root, there doesn't seem to be any benefit (besides defense in depth) to marking the code as read-only.
In general that's a good piece of hardening advice, you just need to mount an empty volume for any temp files that are needed by the app.
For not running as root, the main benefit (to my view) is that there have been multiple CVEs in container stacks where the issue fully or partially mitigated if the container was running as non UID-0
CVE-2021-30465 was partially mitigated, and CVE-2020-15257 had a requirement of the container running as UID-0.
So whilst it's not a panacea, in general not running containers as root is a good layer of defence.