`sudo sh -c "echo 1 > /proc/sys/kernel/unprivileged_userns_clone`
Basically user namespaces let you do the same thing as pid or network namespaces, but with users. This means you can have a "fake" root user. The Linux kernel, marvel of software that it is, is easily confused by this and basically hands you trivial privescs if you're this "fake" root user.
This problem is pretty much only getting worse because user namespaces are becoming more powerful whereas kernel security is staying the same (ie: not moving at all).
That's why it's gated in most distros by default.
Not sure about most distros gating that sysctl. Ubuntu works fine with rootless Docker with no changes and looking at their install instructions, there's only mention of setting that sysctl on debian and arch.
Yep. I just was answering the question, which is what the tradeoff is.
> Not sure about most distros gating that sysctl. Ubuntu works fine with rootless Docker with no changes and looking at their install instructions, there's only mention of setting that sysctl on debian and arch.
Interesting. I haven't checked in a while. I'm also not on the latest debian though.
Debian used to gate them behind a sysctl, but that's changing in the upcoming Bullseye release:
"The previous Debian default was to restrict this feature to processes running as root, because it exposed more security issues in the kernel. However, as the implementation of this feature has matured, we are now confident that the risk of enabling it is outweighed by the security benefits it provides."
https://www.debian.org/releases/bullseye/amd64/release-notes...
Ubuntu has allowed user namespaces by default for years. Which distros are still holding out?
https://github.com/rootless-containers/rootlesskit/tree/v0.1...
Otherwise for general dockerized applications, you won't notice any difference.
You may find some quirks, but these can all be worked around easily as described on the rootless docker page.
We run it in production with no issues so far.