How We Built SELinux Support for Kubernetes
gravitational.com
gravitational.com
Of course you need to buy into the whole premise of SELinux - that in addition to simple Unix permissions you can use mandatory access control to fine-grain control everything that a process can do and get access to. Obviously this is a giant boil-the-ocean project that has consumed many good people for decades. For SVirt (secure virtualization) this was pretty successful, allowing many qemu instances to run on a single server, and confining them with SELinux so each can only access its own files even when they are all running as the same Unix user.
The linked article is a good intro.
[1] But they're fixing this, yay! I can't find a link now but there's a project to decentralise SELinux policy. Perhaps it's still an internal-only Red Hat project.
+1 this is sorely needed and I wasn't sure if the design would ever evolve to include this
https://www.projectatomic.io/blog/2016/03/dwalsh_selinux_con...
These days containers have their own domains:
https://www.mankier.com/8/container_selinux
As you say, the principle is the same.
[edit] container_t is (or was?) an alias for svirt_lxc_net_t:
# podman inspect my_container_id | udica my_container
I wonder how much would the result be different in their case or how much time would they save by having the first version of the policy being generated this way instead of writing it from scratch.
It's unsafe to run untrusted workloads in the host kernel using only namespaces - SELinux does not change this, no matter how much Red Hat wants it to be true. The kernel has a massive attack surface. For untrusted workloads, you need virtualization like gVisor or Kata/Firecracker which isolates workloads from the host kernel.
SELinux can mitigate some logic bugs in the container runtime, yes, but it doesn't make your containers secure.
Also, the title of this submission on HN is inaccurate - Gravitational did not build SELinux support for k8s. Red Hat did, a long time ago. They added support for SELinux to their own product, Gravity.
Also, I could see giving a user shell access within a container and still restricting what they do with SELinux.
That said, as someone one time worked on an SELinux issue in Docker and then became the defacto "expert"... it's a terrible mess, often coming with incompatible policy changes tied to kernel changes and policies which are stacked on top of each other like a house of cards that are difficult to tease apart for other distros.
Something I notice in regards to selinux, is many people don't understand that you are not just applying context to files, but also PID's. You can see the context of a PID by doing 'ps -eZ', similar to how you would with files. This allows you to create a finer degree of process seperation.
You are right though, RedHat has been doing this with OpenShift since 2011, and it is a major part of their security model. RedHat's current push is the deployment of rootless containers, which addresses similar concerncs.
You can also match by owner in iptables. So you can have a network restricted user and sudo to it.
k8s solves this in a different way. You can write applications that are completely stateless when spun up on a k8s node and can be limited to a read only file system. Your container can be a scratch container with no other executables binaries.
I don't know what their product is, but i seriously doubt mixing these two concepts is the right thing to do.
I'm also not sure I would consider this a totally solved problem within kubernetes. In supporting gravity, our kubernetes distribution, we used to hard-disable all privileged containers. This become a common point of friction, since many pieces of common kubernetes software seemed to just use privileged containers instead of properly listing and setting needed capabilities. Although this is a separate issue from selinux, I'm just pointing out that many kubernetes clusters may not be as isolated between running processes as expected.
And to answer what the product is, gravity is a kubernetes distribution, that in this use case allows for packaging a kubernetes based application together with the runtime, for a unified installer for the application + runtime. Think something like packaging up a SAAS application, and being able to install a copy within a banks network who doesn't want their systems talking to the internet while abstracting away the kubernetes portions.
The problem we were running into, is that end user would keep trying to install the application onto selinux enabled hosts, which would break the install or runtime. And whether you believe in selinux or not, it causes significant friction to try and tell your customers we got this, just turn off that security feature in production.
Not to dunk on the work you’ve done: it’s still awesome. Just want to point out the disconnect from users wondering how this would help them.
> The problem we had was that Gravity was never designed to run on SELinux-enabled hosts and our customers kept randomly breaking their installations when trying to run Gravity with SELinux enabled.
In other words, they don't assume that a trusted container/user stays "trusted" operationally.