Docker: insecure opening of file-descriptor allows privilege escalation
bugzilla.redhat.com
bugzilla.redhat.com
You just need to have proc_fd_access_allowed(). I've not checked if ptrace_may_access(task, PTRACE_MODE_READ_FSCREDS) calls into SELinux hooks (it probably does, and if it doesn't then resolving further files probably does too) but neither seccomp profiles (unless you're blocking open(2)) nor blocking CAP_SYS_PTRACE can help you here.
Now, the LXC exploit used ptrace in order to stop the process from closing its file descriptors. I'm not sure how you would reliably hit the race in this issue (something with SIGSTOP presumably?).
In any case, SUSE's update has additional fixes which also fix the issue even when you give a container CAP_SYS_PTRACE (the released patch does _not_ protect containers that have CAP_SYS_PTRACE enabled). The patches will be merged upstream ASAP, but Docker didn't want them in the patchset sent to its customers (preferring instead to update their vendored runC once they are merged upstream).
It's very important for everyone to understand my advice is RHEL/Fedora specific, which is---I think---the source of the misunderstanding here.
Putting aside `ptrace` being the best way to guarantee a race win, the reason for my focus on `CAP_SYS_PTRACE` is that with SELinux enabled there is no other way to exploit having access to the file descriptors. Even if you explicitly try to pass a containerized process an external file descriptor "legitimately" (e.g. with `sendmsg`) SELinux will still ultimately block the access due to the type restrictions. This means that with `setenforce 1` you need to use something like code injection to get the external process to access the file descriptors on your behalf.
Thanks. :D
> Putting aside `ptrace` being the best way to guarantee a race win, the reason for my focus on `CAP_SYS_PTRACE` is that with SELinux enabled there is no other way to exploit having access to the file descriptors. Even if you explicitly try to pass a containerized process an external file descriptor "legitimately" (e.g. with `sendmsg`) SELinux will still ultimately block the access due to the type restrictions. This means that with `setenforce 1` you need to use something like code injection to get the external process to access the file descriptors on your behalf.
Ah okay, yeah I suspected that's what you meant (on _RHEL_ xyz is the case). Thanks for clarifying.
Totally on me. We fight against it, but it's hard not to have the implicit context of RHEL/Fed be omnipresent on the Red Hat bugzilla. In fact, when I wrote the comment in question I had just finished lighting my incense to the sīla of `systemd`... :)
Did not realize you were in Sydney. AU truly has the best hackers.
Update by Trevor Jay:
"This is an extremely difficult to exploit flaw on standard RHEL and Fedora systems.
I checked the 1.10.3 and 1.12.5 builds on Brew. Both drop the `CAP_SYS_PTRACE` capability by default. 1.10.3 blacklists `ptrace` calls under the default seccomp profile. Thus, this flaw only comes into play for containers that already have elevated privileges.
Even if `ptrace` is available. The proposed exploit scenario of quickly attaching to a process joining the container space and using its file descriptors is not possible under the default SELinux configuration. The containerized PID 1 will have a type of `container_t` or similar SELinux type and thus will be blocked by standard type enforcement from accessing accessing any resources that haven't already been made available to containerized processes."
We would all be better off if we designed systems such that some helper process, already running with the right environment / config / privileges, spawns the process for you and proxies input/output to your terminal.
And (as I mentioned in the other thread) this helper process could be literally sshd. Instead of having sudo, ssh root@localhost. No weird process trees with confusing things like effective UIDs. Instead of having runc exec, ssh root@container. No file descriptors get passed that aren't explicitly forwarded over the SSH connection.
Patching sshd to run over UNIX sockets without encryption and to use getpeername() for authentication is left as an exercise to the reader.
Host web
LogLevel QUIET
ForwardAgent yes
ProxyCommand setsid lxc-attach -q -n web01 --clear-env -- /usr/sbin/sshd -iThat design certainly lets lxc-attach permit only running sshd, though, which is a benefit. (And I forgot about sshd -i, which you could presumably point at a UNIX socket using socat or something.)
1. ping hasn't needed to be setuid since Linux 3.0 (and isn't on most distros), the kernel lets you call socket(AF_INET, SOCK_DGRAM, IPPROTO_ICMP) without privileges, which lets you send ping packets and nothing else. Looks like Mac OS X and FreeBSD, at least, also support the same interface. This approach has successfully been used to eliminate other setuid binaries like pt_chown, which fixes ownership of a tty (the kernel now just sets the ownership correctly when you call open).
2. Use something like ForceCommand to allow "ssh root@localhost /sbin/ping", and make /bin/ping a shell script that does that. This carries strictly less complexity than making a setuid binary; any attack that applies to it also applies to setuid binaries, but the execution environment for setuid binaries is more open to the attacker's control.
3. Make a little ping server that you can request to conduct pings for you. For ping in particular this is probably silly, but for things like updating utmp (traditionally you make every program setgid utmp, or you use a helper setgid binary called utempter), there's probably some existing daemon like logind that can grow some small APIs.
# chmod 0711 /bin/ping # setcap CAP_NET_RAW=eip /bin/ping