File Permissions: A painful side of Docker (2019)
blog.gougousis.net
blog.gougousis.net
Make the build script use local $USERID and $GROUPID as args during the build process.
In docker-compose.yml (or, if using docker directly, using --build-arg):
build:
context: ./build
args:
USERID: ${USERID}
GROUPID: ${GROUPID}
So you're passing the local uid and gid as variables to the build process.(1)In build/Dockerfile:
FROM image:tag
WORKDIR "/application"
ARG USERID
ARG GROUPID
RUN if [ ${USERID:-0} -ne 0 ] && [ ${GROUPID:-0} -ne 0 ]; then userdel -f www-data ;fi \
&& if getent group ${GROUPID} ; then groupdel www-data; fi \
&& groupadd -g ${GROUPID} www-data && useradd -m -l -u ${USERID} -g www-data www-data -s /bin/bash \
(1) $USERID and $USERID might not be available as an environment variable on your system. To do so, place this under .bashrc: export USERID=$(id -u)
export GROUPID=$(id -g)Our approach so far was to add yet another layer (a script to pass uid/gid to Compose), but if we don't need the script that would be fantastic.
EDIT: Ah, I just saw the bashrc wrinkle you mention. Yeah, that's why we had the script, and it's a damn shame Docker can't do this natively. It has been a major hassle.
Yep, it's because the build args get read in from a .env file by default and then from there Docker Compose sends those build args to Docker when it builds the image.
This was one of the topics from my talk at DockerCon last week (creating a production ready Docker Compose set up). The video and 6,000 word blog post for it will be coming out tomorrow. Both things will be added to the talk's reference links at https://github.com/nickjj/dockercon21-docker-best-practices.
I run all of my containers as a non-root user and create the user in the image with its default values of 1000:1000 for the uid:gid. I haven't bothered to expose the uid:gid as build arguments because it's pretty much never an issue in development or production.
With a uid:gid of 1000:1000 built into the image any bind mounted files end up being correctly owned by the Docker host's user under the following conditions:
- Docker Desktop on macOS
- Docker Desktop on Windows using WSL 1
- Docker Desktop on Windows using WSL 2 and native Linux (as long as your dev box's user is set to 1000:1000)
IMO it's really rare that your dev box's user wouldn't be 1000:1000 on native Linux or WSL 2.
In production you also have full control over the uid:gid of your deploy user.
The only time where it kind of stinks is CI, but it's super easy to get around this by simply not using volumes in CI.
I have a bunch of examples of this pattern at:
- https://github.com/nickjj/docker-flask-example
- https://github.com/nickjj/docker-django-example
- https://github.com/nickjj/docker-rails-example
- https://github.com/nickjj/docker-phoenix-example
- https://github.com/nickjj/docker-node-example
- https://github.com/oleksandra-holovina/docker-play-exampleAny company-wide (GNU/)Linux deployment that uses LDAP or some other centralized user directory will not have devs with UID/GID 1000:1000. Hope is not a strategy.
You can go the extra mile and turn the UID:GID into build args like the original parent and you're good to go. No hacks necessary, and since it's all self contained into a .env file there's nothing extra you need to run since you're likely using an .env file already for other vars.
Alternatively you could do this: https://news.ycombinator.com/item?id=27344491
In either case you can solve the problem without too much effort.
That doesn't help you if you're attempting to use pre-built/existing Docker images that are not built internally and make the assumption that “1000:1000 is good enough”. You then not only have to hack around Docker limitations, but also around someone else's broken assumption.
Most pre-built images that I've come across don't require bind mounts to function.
Images like PostgreSQL aren't affected by this because you can use a named volume, and most pre-built applications that are shipped as images tend to store their state in a database and don't require bind mounts to function.
Any major company using LDAP/AD or other forms of centralized user management won't be able to make that guarantee.
> In production you also have full control over the uid:gid of your deploy user.
If you're running in an un-managed environment, yes - managed hosting of any kind generally doesn't provide these guarantees.
This method works well until it doesn't work at all, and I think I would prefer one that works slightly less well but also had an easier way to override it. Then again, I might try this and see if we ever hit an issue, thanks!
1. Images are still pre-baked with a given UID/GID pair, so you can't distribute them as something universal and reusable.
2. This requires workarounds / extra steps on a local workstation, so it doesn't work for everyone unless they follow a given project's unique quirks setup.
Shell/compose duct tape like this doesn't make for a great experience, this really should be solved by upstream projects themselves as it's an extremely common issue when attempting to use Docker.
The only tedious thing is you have to adapt this for every image type you run.
In fact, docker-compose up -d takes care of the build thing by itself. It's a five second tradeoff for the lifetime of the application.
In environments where vulnerability scanning of docker images used is important, running anything in production that isn’t stored in a docker registry kind of breaks things.
This approach also won’t work with container orchestrators like Kubernetes, ECS, Lambda, CloudRun, etc.
Where I can see doing a docker build of a small layer that just sets file perms potentially being useful is for container based dev environments to be ran on laptops and workstations.
The tedious thing is that this escalates into complexity whenever you have to deal with K developers using M projects developed by N teams each using a different way to handle this:
Do I need to set USERID for project foo, or UID? Does it default to 1000 or the author's UID? Oh, someone has a problem with our project, did they remember to set COMPANY_USERID in their bashrc? Oh, wait, they're using zsh, how do you do that there? Oh, but they followed this other project's readme and that set COMPANY_USERID but not COMPANY_GROUPID...
Docker is supposed to simplify this by unification and a limited API surface, and applying hacks like this on top kind of kills that whole premise.
You set it to the output of id -u and id -g. It's two lines. There are definitely lots of things more complex when dealing with docker than this.
You provide the team with a script containing those two lines and a docker-compose wrapper and you're set.
Of course it would have been better not to have to care about these things, but hey, at least you're not installing and configuring 4-5 services to bootstrap an application.
You run docker-compose build ONCE and you're set. On my machine, it takes five seconds.
Heck, you can even run docker-compose build everytime you start the application, it will use the cached build and take less than one second.
---
Correction: the docker-compose up -d takes care of the build process the first time it runs.
Literally, it takes more to complain about the issue than build the image ONCE.
I don't think the reproducibility is out. It's the same app, the same image, the same intended user, you just inject, once, the local user and group ids.
Decades of network filesystem users have had many solutions to that.
1) pass user/group names around and resolve them at the destination to UID/GID; 2) ignore them entirely; assign ownership of all newly created files to the currently authenticated user (if authorized).
Are there other ones?
services:
foo:
image: foo/bar:6.9
user: ${UID:-1000}:${UID:-1000}
On Linux with Bash it runs with your current user and most other platforms it runs with id 1000, which is setup as the default user in the Dockerfile. This is no problem on MacOS or Windows because of the way Docker-Desktop uses VM's.ZSH or other shells don't necessarily set $UID, so if you're running Linux, not id 1000 and not running Bash you might need a little .env file with `UID=1001` in it to make it work. And then the user is still nameless in the container. This is kind of rare and I only use it for dev containers where most relevant files (and permissions) are bind-mounted from the host, so it hasn't really been a problem in practice.
Remaps would be cleaner but I find it too much work to explain for normal developers just wanting to use a dev container.
See more here: https://stackoverflow.com/a/50900530/15428104
$ declare -p UID declare -ir UID="1000"
The -x option is missing.
It allows different mounts to expose the same content with different ownership, and in general to map permissions IDs between mounts in any way we like.
systemd-homed wll use that to abstract over the uids and gids of portable home directories, for example.
Wait, what? Why not install the immutable files as root and let them be readable to everyone?
http://docs.podman.io/en/latest/markdown/podman-run.1.html
Especially the "userns" option with the "keep-id" value.
"Docker and the host filesystem owner matching problem": https://www.joyfulbikeshedding.com/blog/2021-03-15-docker-an...
In my blog post I layout 2 solution strategies, how one might go about implementing them, and caveats to watch out for.
1. Matching the container's UID/GID with the host's UID/GID.
2. Remounting the host path in the container using BindFS.
There's also weird junk you sometimes need to do in order to capture file handles depending on how a container engine is running the container, which you need to do before you fork or drop privs. But it took me years to finally run into that use case, most people will never need to do this.
It's still pretty much a proof of concept and it relies on docker compose but perhaps some of you may find it useful as a starting point: https://github.com/tacone/loki
- `--user` didn't work for me because there were root permissions in my image
- I didn't dig into why `userns-remap` didn't work
- I didn't give https://github.com/boxboat/fixuid a try yet
Some notes from my experience
setfacl -dm "u:alexandros:rw" ~/alpine
should besetfacl -R -dm "u:alexandros:rwx" ~/alpine
In case:
- `-R`: There is existing content in `~/alpine` you want made avalable
- `x`: You want your container to be able to create directories
However, you can still run into problems if
- Your container copies data from outside your bind-mount to inside. It sort-of worked except somehow the mask was `r--`, making things lose writeable.
- Your container moves data from outside your bind-mount to inside. This fully preserves the permissions.
I ended up creating a `.keep` file in the bind mount and doing a `cp --attributes-only --preserve=mode,ownership,xattr .keep <target>`
nsenter -U --preserve-credentials -n -m -t $(cat $XDG_RUNTIME_DIR/docker.pid) /usr/bin/chown -R root:root /home/user/workspace
This one liner enters the namespace of Rootless Docker, and does the chown back to your normal user (root is your host user when you switch back).Useful anytime you use a filesystem mount ... (Ex: storing database on disk so docker doesn't kill it every run).
You can now do backups, rebuild docker images, etc.
More information: https://github.com/jpetazzo/nsenter#how-do-i-use-nsenter
And not only that, I think that examples given in the article ("Assume that your Apache/PHP container is mounting the host’s /home/alexandros/myapp/ application directory to the container’s /var/www/html directory.") are in fact anti-patterns. If your container depends on specific file being available at specific location on the host then you're doing it wrong. The only place where that makes sense is on developer's local environment. In shared enviornments you want something like Kubernetes ConfigMap to contain config files, and dedicated persistent volumes for everything else.
It could be I just haven't dug enough into the kernel internals, maybe there is a transparent permissions remapping thing. But something would absolutely have to map permissions. Otherwise there is no way to use filesystem ownership between execution environments without them using conflicting UID/GIDs, to say nothing of changing the file perms.
a) Many problems solvable with a volume can be solved with a bind-mount, cache-mount, etc [0].
b) In the event that you actually need to map in a user-file, wrap the docker command in a script that manages the logic. At this point you're writing a system tool that's doing things outside of the context of a container - it's not really docker's fault that it doesn't try to make this trivial.
Use supervisord to coordinate the processes inside your Docker container, as easy as that. Bonus point, you don't need to wrangle with properly handling "docker stop"/ctrl+c.
However, this is something that's basically unavoidable if you're attempting to use OCI/Docker for dev where you access a developer's source code checkout from a container running a standardized language runtime. And that's what a lot of people use OCI/Docker for...
In practice bindmount smell can also be somewhat alleviated by using things like k8s device plugins to request things at a higher level ('I want GPU access' vs. 'please bindmount /dev/drm... and use the proper modes'). It's still effectively a bindmount, but some extra security precautions can be made to ensure exclusive access and that no arbitrary mounts from the host are permitted. And things like k8s device plugins can also poke at file modes and other namespace magic at runtime so that the end user never has to worry about things like UID/GID and chardev modes. That IMO prevents the smell associated with random host bindmouts.
They're also very easy to write, so if you ever happen to run k8s and need to give workloads access to some odd/custom host hardware, implementing a proper plugin for it is quite painless and gives much better guarantees than plain bindmounts.
Would love a real solution from docker though.