Systemd service sandboxing and security hardening (2020)
ctrl.blog
ctrl.blog
https://news.ycombinator.com/item?id=29976096
Or simply follow whatever `systemd-analyze security` recommends, just make sure you run it on a system with recent systemd.
systemd-analyze security
Is there a tool like `audit2allow` for systemd units? selinux/python/audit2allow/audit2allow: https://github.com/SELinuxProject/selinux/blob/master/python...I haven't seen any difference between distributions with the same systemd version. Anything with a recent one should do fine. More recent than RHEL8, mind you (which is on systemd 239): for example, a syscall allow/deny analysis is buggy there and asks you to enable some protections, and then disable them. The same unit is analyzed correctly on my desktop with v250 (I use the popular rolling release distribution).
I haven't seen anything like audit2allow. It's probably not especially necessary because of the difference in philosophies: SELinux is deny by default, while in systemd you're playing whack-a-mole anyway, and are expected to add directives one by one until the application stops working. Unit logs usually make it obvious if something was denied.
Another approach would be to set SystemCallLog= to be the opposite of SystemCallFilter= (negate each group with ~) and then you'll see the call (and caller) in the journal.
SystemCallFilter=@basic-io
SystemCallLog=~@basic-io
and then restart it, I get a bunch of audit SECCOMP log messages (and the service fails to start): audit[1185972]: SECCOMP auid=4294967295 uid=63946 gid=63946 ses=4294967295 subj==unconfined pid=1185972 comm="(t-online)" exe="/usr/lib/systemd/systemd" sig=0 arch=c000003e syscall=59 compat=0 ip=0x7f5c632f66c7 code=0x7ffc0000
This is debian bullseye and systemd 247. I wonder if your audit logs are going somewhere else?In the SELinux MAC system on RHEL and Debian, in /etc/config/selinux, you have SELINUXTYPE=minimal|targeted|mls. RHEL (CentOS and Rocky Linux) and Fedora have SELINUXTYPE=targeted out-of-the-box. The compiled rulesets in /etc/selinux/targeted are generated when [...].
With e.g gnome-system-monitor on a machine with SELINUX=permissive|enforcing, you can right-click the column header in the process table to also display the 'Security context' column that's also visible with e.g. `ps -Z`. The stopdisablingselinux video is a good SELinux tutorial.
I'm out of date on Debian/Ubuntu's policy set, which could also probably almost just be sed'ed from the current RHEL policy set.
> * SELinux is deny by default, while in systemd you're playing whack-a-mole anyway, and are expected to add directives one by one until the application stops working. Unit logs usually make it obvious if something was denied.*
DENY if not unconfined is actually the out-of-the-box `targeted` config on RHEL and Fedora. For example, Firefox and Chrome currently run as unconfined processes. While decent browsers do do their own process sandboxing, SELinux and/or AppArmor and/or 'containers' with a shared X socket file (and drop-privs and setcap and cgroups and namespaces fwtw) are advisable atop really any process sandboxing?
Given that the task is to generate a hull of rules that allow for the observed computational workload to complete with least-privileges, if you enable like every rule and log every process hitting every rung on the way down while running integration tests that approximate the workload, you should end up with enough rule violations in the log to even dumbly generate a rule/policy set without the application developer's expertise around to advise on potential access violations to allow.
From https://github.com/draios/sysdig :
> "Sysdig instruments your physical and virtual machines at the OS level by installing into the Linux kernel and capturing system calls and other OS events. Sysdig also makes it possible to create trace files for system activity, similarly to what you can do for networks with tools like tcpdump and Wireshark.
Probably also worth mentioning: "[BETA] Auditing Sysdig Platform Activities" https://docs.sysdig.com/en/docs/developer-tools/beta-auditin...
A bit of SELinux:
# /etc/selinux/config
SELINUXTYPE=targeted
SELINUX=permissive
$# touch /.autorelabel # `restorecon /` at boot
$# reboot
$# setenforce 1 # redundant
$ sudo aureport --avc
$ journalctl --system -u auditd
$ journalctl --system -o json-seq --reverse
$ journalctl --system --grep "AVC" --reverse
journalctl -fa _TRANSPORT=audit
journalctl -fa _TRANSPORT=audit --grep AVC
journalctl -a _TRANSPORT=audit --grep 'AVC avc: denied' -o json | pyline -m json 'json.dumps(json.loads(l), indent=2, sort_keys=True)'firejail does this a bit better but it also started out with a blacklist approach and it's more geared towards desktop application use, not system services.
[0] https://risky.biz/netcasts/risky-business/ [1] https://www.airlockdigital.com/
That isn't a whitelist approach.
That _is_ a whitelist approach.
On the other hand if you start with a blank slate filesystem root and only bind exactly the whitelisted paths then there is nothing to leak through.
There are other ways in which blacklist-all can fail to be equivalent to whitelisting.
Systemd also supports firewalling: it supports IP address allow/deny policies, ports, etc. For more advanced firewall policies you're probably better off using an actual firewall daemon like firewalld or ufw.
If you trust those programmers, it’s indeed way more convenient than other tools, if only because it removes the need for configuring things twice. For example, instead of configuring your web server to serve files from /foo/bar/ _and_ telling SELinux that your web server is allowed to read from /foo/bar, you only configure the web server, and it will tell the OS “I shouldn’t read from anything but /foo/bar, starting … now”.
You’ll have to trust the web server to do that, though.
$ systemd-run --user --scope --property=MemoryHigh=1G qbittorrent
It works just as you'd expect — if qbittorrent's working set goes above 1024 MiB, it pushes the least recently used page out of the page cache. Doesn't really have any effects on upload or download speeds, while helping to keep more useful data in memory.Many isolation flags are not available in `systemd-run --user`, though, so if you'd like to have some protection you either have to combine `sudo systemd-run` with `su -c`, or wrap the command in firejail.
https://github.com/netblue30/firejail/
It uses the same kernel knobs as systemd does, but is more user-friendly and has more features.
I use it for every application that handles data received from other machines: books, images, documents, whatever.
In my case I'm using Bubblewrap because Firejail was only used for Zoom, and this felt a bit of a waste considering Bubblewrap was already installed.
[Service]
AmbientCapabilities=CAP_NET_BIND_SERVICE
SocketBindDeny=any
SocketBindAllow=tcp:80
SocketBindAllow=tcp:443
These ports should be denied by the kernel because they're already taken by httpd, and all other will be denied by bpf filters installed by systemd.It feels like plugging holes in a dam, but that's what you do with popular operating systems.
1. add this to .service
AmbientCapabilities=CAP_NET_BIND_SERVICE
2. or listen on :8080 and use NAT: iptables -t nat -I OUTPUT -p tcp -o lo --dport 80 -j REDIRECT --to-ports 8080
3. or make the port unprivileged sysctl -w net.ipv4.ip_unprivileged_port_start=80
It may work for httpd too, I haven't tested it.Not everything that wants to open up a port seems to support socket activation. I tried with 6tunnel and couldn't get it to work.
[1] https://www.freedesktop.org/software/systemd/man/systemd-soc...
https://www.freedesktop.org/software/systemd/man/systemd.exe...
One approach that may work with systemd is to have two processes. One would be a broker, running as root. It would grab a port, for example. The other process would be spawned by the broker as a limited service and inherit that port from the parent, with no permissions of its own to open it, only to inherit.
IDK how to express that in systemd-land though. At that point you might be better off just writing the code to sandbox things yourself.
This allows to let systemd to manage ports using socket units which will also stay up and buffer requests when restarting a service, allow service activation on demand/incoming requests or per connection service instances, e.g. for better isolation of sshd's per connection/user.
i added support for httpd to support systemd socket activation in 2013: https://svn.apache.org/viewvc?view=revision&revision=1511033
httpd can start as non-root, assuming other configurations like the access / error logs are writable by the non-root user.
I would pretty much never ask a human being to write SELinux policies unless that was explicitly part of their job whereas I can pretty much point any developer to what systemd is providing and they'll be able to work with it.
So the overlap choice seems to be more around SELinux versus Capabilities. Where SELinux is more fine-grained and tunable, but more complicated also.
The overlap/comparison between capabilities, systemds features, and selinux features isn't really well defined in any meaningful way IMO. It's really like 5 different features being used in various ways.
The normal filesystem permissions of read/write/execute for user/group/other are among those known as "discretionary access controls," meaning that they can be relaxed.
The systemd unit security options are discretionary, at the control of the administrator.
The rules can also be adjusted, and there are a number of tunable parameters.
The intent is that it is never disabled.
That is, the same command that gets denied by SELinux through systemd will run fine (and unprotected) when started from a shell.
Do you write your own policies for individual end-user programs?
(and I say this based on both observation and personal experience, I have some stuff to harden later this year and I'm really hoping I'll be able to involve somebody who -has- that level of SElinux knowledge but plan B is almost certainly going to be 'mst does his best with the unit configs')
But if one can spare the time, SELinux can secure everything, not just systemd services.
It all depends on the threat vectors one faces.
[Service]
ProtectSystem=strict
ReadWritePaths=/some/path
ReadOnlyPaths=/some/otherpath
InaccessiblePaths=/etc
Unlike systemd, they then apply to everything.
SELinux is a policy system where policy is enforced via labels.
Labels are applied to processes which classify what the process is.
Labels are applied to files which define the what classification of process can access the file.
The application of labels happens automatically based on policy. Such policy would include the location of the file or the label of the parent process.
As an example, the default policy for httpd would prevent httpd from accessing /etc/passwd even though the process is running as (or can be) the root user. I believe you could also do interesting things like prevent httpd from opening a socket on a non-standard port if you wanted to.
SELinux is very powerful but complicated. Ideally you use this with distro packages which should have policies already configured for you.
Critically it is not one vs the other. Use both if you have it.
Caveat: it has to dig into ALL the linked libraries as well.
Then anything with systemd and security can stuff it.