Of course you can do all these things (on Linux at least), but it involves a dizzying array of technologies.
Of course you can do all these things (on Linux at least), but it involves a dizzying array of technologies.
https://man.openbsd.org/pledge.2
https://man.openbsd.org/unveil.2
A few random examples:
https://github.com/tmux/tmux/blob/c8494dff7b6b9a996866edaf8c...
https://github.com/openbsd/ports/blob/master/www/mozilla-fir...
https://github.com/openbsd/ports/blob/master/www/mozilla-fir...
To get the best isolation you need to patch the source — the application needs to go through initial setup and then drop privileges to the absolute possible minimum. But it's easy to make custom wrappers for third-party applications — the above profiles taken from the OpenBSD ports tree are the proof.
This shouldn't be underestimated. It's really not that useful for most apps that don't have the source patched.
For example, any dynamically linked program that needs to spawn a subprocess using a command that it doesn't ship alongside itself is going to cause problems in your sandbox. That turns out to be the vast majority of programs that have been shipped in the FHS universe for the last few decades. It's only recently that sandboxing has come into vogue.
And the OS's that do a good job (like MacOS) get a bad rap from both developers and users who are mad their programs can't do what they want them to do because they don't have unfettered access to the filesystem any more.
sudo systemd-nspawn \
--directory=/ \
--volatile=yes \
--bind=/home/tmp:/home/tmp \
--private-network \
--port=tcp:80:8000 \
--property=CPUQuota=5% \
--property=MemoryMax=1G \
-- \
bash -c 'yes > /dev/null'Like all other systemd-nspawn features, this is not a security feature and provides protection against accidental destructive operations only.
$ podman run -it myimage bash
Fuse is rather irrelevant since I will mount volumes if I need to do IO. And slirp has been benchmarked to 9+ Gbps. So sure there's always room for improvement but the current situation is pretty sweet imo.
SELinux is pretty standard nowadays, most people just can't be bothered to learn it.
When I started using Linux, the security model had features like root being allowed to pwn the kernel and just about anything else. Now /dev/mem is gone, lockdown exists, namespaces exist, there are many more capabilities than there used to be, seccomp and seccomp-bpf exist, etc.
A lot of security ideas on Linux also rely on trusting the software, a model that may very well make sense in the context of Linux distributions with repositories of open source software. In this case, software itself opts into sandboxing to harden itself. This is pretty much what goes on with stuff like Flatpak. This also sounds like the idea with Pledge.
Can we do better? Maybe. I think adding true, practical sandboxing onto existing kernels not designed with this in mind is possible with tradeoffs. gVisor is an interesting approach: it's not a panacea but certainly a step in the right direction, using a usermode kernel that is itself separated into modules that are heavily locked down using ordinary Linux security mechanisms. This adds an additional layer to the moat that in theory is quite strong.
Another approach is virtual machines. And yes, this approach has some reasonable critique: typical VMMs and virtual machine software are very complex and involve needing to trust that the hardware was implemented correctly that the design isn't flawed, something that some people never trusted and others have lost some faith in with many hardware flaws and Qemu bugs coming to light. However, I still think lightweight virtual machines ought not be ignored entirely. Improvements have been made over time that make the idea of using virtual machines for this purpose a bit more within reason. For example, Firecracker and Ignite are pretty interesting tools for running software in a more sandboxed manner. And, even with virtual machines potentially having flaws, a lot of the time it may at least require root privileges or more to actually exploit some flaws, which makes this not entirely useless from a defense-in-depth standpoint.
But you just want an idiot proof way to call some binary and provide some limits, and have it be guaranteed to actually enforce those limits as a proper security boundary. Do we have it?... No. Will it be simple, something that can be done with some simple syscalls? Probably not. Could it be done with an unfortunately necessarily complicated piece of code? Probably yeah. I'd love to see a tool that can give you basically seamless sandboxing with simple options and backends for different sandboxing approaches like using VMMs or gVisor. Provide easy ways to run X11 or Wayland apps and automatically give you a sandboxed Xephyr instance or something to provide some sandboxing for Xorg apps. Will it happen? Dunno.
As for me, I'd really like a way to create a namespace where only some hosts on the network are accessible. It's possible to do today using various technologies, but it'd be a lot easier using something like the gVisor model with a usermode networking stack, I think. Maybe some day.