Linux 3.8 introduced unprivileged user namespaces [pdf]
man7.org
man7.org
This is included in kernel 3.8 released in 2013, and the linked PDF says that's when CLONE_NEWUSER itself was released - i.e., ever since user namespaces have existed, they've been allowed to unprivileged users.
Some downstream distributors, including Debian, started carrying a patch that added a restriction back in via the kernel.unprivileged_userns_clone sysctl. This is because allowing unprivileged user namespaces exposes a lot of attack surface to users. There's nothing fundamentally unsafe about it, but there are many kernel components for which bugs are only exploitable by unprivileged users if the feature is on. See https://lwn.net/Articles/673597/ for more discussion. (I thought a version of that patch eventually got accepted upstream, but I can't find it.)
The only reference to kernel 4.6 in the document is CLONE_NEWCGROUP, which is entirely unrelated to unprivileged user namespaces.
Edit: corrected from 0.8
Edit edit: corrected from (1, 0.2)
That could fix the barrier of entry and make docker tutorials passed around more canonical. I would love it if docker "just worked" across machines when testing locally, commands and all.
Even sharing docker in open source projects, there's a learning curve for me where commands in the README won't work. Is it my docker installation? Version differences with docker? Docker compose? Did the container images I'm pulling in change some way?
Do I need sudo or not? I guess with proper group permissions I'm okay - but will the developer(s) I share instructions with have these permissions ready to go?
On StackOverflow: copy/pasting docker CLI (even given proper context fitting into a larger whole) commands and configs probably has a 50% success rate for commands, and maybe 10% if it's a config of some sort (e.g. compose files)
No, because 4.6 is pretty old at this point and this feature has been around for ages.
I've been using podman instead of docker more and more recently, it's pretty great. It doesn't require root access to the machine (that's how I use it), although there are some limitations.
https://github.com/containers/libpod/blob/master/rootless.md
It needs the setuid (or setcap) helpers `newuidmap` and `newgidmap`, plus setup in /etc/subuid and /etc/subgid, to allocate some UIDs on the host system for unprivileged users to use on the guest container. This is required so that you can have multiple users inside your container. There is a way to use user namespaces entirely unprivileged, but you only get one UID and one GID inside your namespace, because UIDs and GIDs need to have a one-to-one map across user namespaces. You can pick whether your UID maps to root inside the namespace, but then you can only use root, you can't drop privileges to anything else.
On most Linux distros /etc/subuid and /etc/subgid get automatically created these days when you create a local user account. If you have LDAP users (which we do), you need to fill them in manually somehow.
I've also been using entirely-unprivileged user namespaces at work to sandbox builds and tests so that running them on the build/test farm works just like on your local machine, and you have fewer "it works on my machine" problems. In this case I only need one user inside the namespace, and it doesn't need to be root (you can set up mounts, networking, etc. before you drop privileges).
[See also my comment about how 4.6 isn't relevant here, we have some machines at work that are on 4.1 and the functionality works fine.]
Another project in this space is bubblewrap https://github.com/containers/bubblewrap , which can run either with unprivileged user namespaces or by being setuid. It's intended to create an environment for container runtimes to use so that the container software itself doesn't need to be privileged, and the idea is that bubblewrap itself uses the privileged interfaces to set up the environment and then drops privileges before running user-provided code, so it shouldn't introduce more risk.
For an almost drop-in daemonless, rootless docker replacement for this use case, see podman (https://podman.io/). You can `alias podman=docker` and it will just work.
You do need a bit of configuration, specifically you need to create `/etc/subuid` and `/etc/subgid` if they don't exist and add subordinate UIDs/GIDs which will be used to map the containers users and groups. E.g.
usermod --add-subuids 100000-165536 $USER
usermod --add-subgids 100000-165536 $USERYes, and it's super awesome. Unless you need docker-compose!
https://developers.redhat.com/blog/2019/01/15/podman-managin...
There's a solution for that :) (Note: still somewhat beta quality)
E.g. https://github.com/containers/libpod/issues/4039#issuecommen...
To be less nebulous: regarding the problem of being “rootless,” macOS support is neither here nor there. As far as I can tell it doesn’t really matter since the containers are already in a VM anyways.
There are some caveats - specifically copy-on-write filesystems aren't yet supported but Redhat says it's on the RHEL 7 roadmap. And you need a fresh of install 7.7 to get the right behavior out of the box, otherwise there are manual steps required for upgraded systems. Specifically stackoverflow examples should if there is a usermapping to unprivileged IDs and the user aliases or links docker to podman.
[1] https://www.redhat.com/en/blog/three-new-container-capabilit...
This is not particularly surprising. Docker does two things. One is a filesystem package manager, the other is providing a common interface around dozens of low-level features. The filesystem packaging stuff works well enough. It's the other part that is inherently complex because it does so many things at once (networking, mounts, resource management, security, process management, ...). When something doesn't work you'll end up with old-school linux sysadmin work except that you now don't deal with a single network interface and a handful of iptables rules but dozens scattered across namespaces and complex mount trees.
There's no magic, just convenience as long as you stay on the happy path.
[1] https://www.freedesktop.org/software/systemd/man/systemd-nsp...
> User NSs permit novel applications; for example: > Running Linux containers without root privileges > Docker, LXC
This seems to answer yes to your question.
I learnt a lot from these talks.
kill(1, 9);https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
This is great news, because I can't switch over to OpenBSD (docker, bluetooth, etc) or more folkloric distributions like VoidOS and Qubes. Going to make a bunch of Anki cards today to remember these namespaces and how to use them!
The main use of user namespaces seems to be running stuff that wants to be root as non-root. It would seem better to simply fix all those tools to not check if they are root, and instead just try to do the thing they were trying to do.
edit: oh, it's also mentioned in the document slide 53 along with Flatpak
[1] https://unix.stackexchange.com/questions/557293/how-can-i-ma...