The single-user approach creates clashes between the network devs who want to build empires of containers they "own", and the stack-level bare metal purists who want the system to be as clean and secure as possible by isolating things where they should be isolated (to a single user instance for that purpose alone). This is not a new problem, nor a very well thought-out solution.
Containers are always a less-than-ideal implementation for people running Linux natively. The ideal way to sandbox in Linux is create a user account, download and test whatever code, see what breaks or infringes with its unique notion of "privileges", and delete the user when done.
But because you can't switch users on the same kernel when you're not running Linux natively, we have containers and all the messes they create. https://developers.slashdot.org/story/12/12/29/018234/linus-...
- under its own userid
- in its own namespace for mount, network, process-id, user-id, ipc, uts and control groups
and presto, you are running in a container.
So people thought that instead of fixing the apps it was easier to fix the kernel. But this resulted in a big complexity with namespaces, capabilities, cgroups etc.
With linux, it's because the tools are exposed piecewise. Creating a process namespace isn't the same thing as creating a virtual network device. You can use cgroups to stop a container all at once, sure, but you can also use them for other useful purposes and the code is the same. Some of these APIs are simple, some less so. But they're tools.
I know, I know, now you're asking "well, why didn't someone just put all the tools together in one box that would do it with one sane UI?!"
They did. It's called Docker.
Last I checked, they weren’t suitable for multi-tenant machines or for running untrusted code.
Beyond that it is very difficult to get things right, and there's definitely some information leakage (e.g. see the output of `mount` in a docker container... not a docker specific problem).
The main issue is, having a VM layer is always going to be "more secure" (because it's another security boundary) than not having a VM layer. Of course VM's have their own issues.
So if you have a need to run multi-tenant/untrusted workloads it's kind of a CYA situation where if you don't use this extra security boundary and there is a problem sometime down the line then you'll have to answer questions like "why didn't you do ..."
The reality is you can do a lot to lock things down with seccomp+selinux/apparmor.
No free lunch. But yeah, Linux's suite of "container" tools are basically as secure as any other OS's "containers", no matter what spin you're reading elsewhere.
> They did. It's called Docker.
Yes I see that. But I was hoping more for a toolbox like GNU fileutils, not a toolbox with a billion dollar valuation (which is what Docker is)
What do you see as the primary differences between those two projects, or types of projects?
I think just saying it's a VM is an over simplification
Admittedly jails and pledge(2) seem rather nice to work with compared to Linux's equivalents.
He has spoken out against the modern security industry and security theatre, both of which are distinct from security.