Troubleshooting Systemd with SystemTap
blog.janestreet.com
blog.janestreet.com
SystemTap has implemented prototype support[1] for BPF, although I'd expect it's not yet widely used or tested.
BPF itself has had a bit of a history of security issues[2] but it appears to be stabilizing, and can theoretically be a lot safer than approaches like SystemTap which require runtime kernel-level code to be inserted.
[1] - https://sourceware.org/git/?p=systemtap.git;a=blob;f=README;...
Systemtap works beautifully under rhel/centos, but it's a pain to make it work on Ubuntu.
This has slown down adoption by a lot.
We could have had dynamic tracing in gnu/Linux a lot earlier if Ubuntu had done a better job at supporting it.
This is no surprise though: canonical has often refused to adopt stuff that have been thoroughly tested and vetted in the rhel world: think of firewalld Vs ufw. Or selinux Vs apparmor. Or cockpit, that has no equivalent in the canonical world afaik.
that it doesn't work on anything else properly is pretty clearly a feature, not a bug
Ubuntu is baded on Debian, if Red Hat stuff is not in Debian is hard to blame Canonical and not Debian but let me know this mental gymnastics of blame shifting
It's worth noting that Canonical also has a pretty well-known history of being a "bad player" in the wider community. They rarely upstream patches, break things and very much try to "do their own thing"
The more cynical of us think its Canonical attempting to segment Ubuntu off from the wider Linux ecosystem to force some kind of "vendor lock-in" by virtue of being the "newbie friendly" Linux option
Some of the more egregious examples of things they forked were "small" pieces of the stack. Like the entire fucking desktop environment (Unity) and display server (Mir). Relying on things like Ubuntu-specific forks of GTK and other critical components, which is why those technologies never appeared in other distributions
Anecdotally I've also been told by people who work at Red Hat that a lot of Canonical employees were picked up about a year ago after being terminated (allegedly) for attempting to upstream the code they worked on
Feel free to wave it off as "mental gymnastics" or "blame shifting" but I do genuinely believe that Canonical has an intentionally poisonous effect on the wider Linux ecosystem
I think unless Canonical removed a lot of RedHat tech from Debian and replaced it with their own then you are just complaining that Debian is not a big fan of RH tech.
I can see a practical argument for safety, but the theory is exactly the same here: both techniques inject arbitrary runtime code into the kernel. That code could do bad things, and so it has to be validated. The location of the validation in the architecture is different, but it's a practical difference.
As far as my personal opinion? ?!@$#!@ just pick one, folks. The time for dueling tracing paradigms was like 8 years ago.
I've pretty grim personal opinion of it, based on similiar experiences as this story. I'd rather just use a simpler approach (another init) that _works_, and doesnt involve hours of debugging. I've never had any troubles with sysv-style init for example.
I used to run my home automation stuff (Home Assistant, Node-RED, and a custom shim between my power meter interface and MQTT) in a FreeBSD Jail and recently switched to a Debian VM.
Getting everything to start properly with FreeBSD's rc scripts was a minor nightmare. It took a lot of work to accomplish the task of "Run command x, which doesn't daemonize, as user y in directory z and redirect stdout and stderr to blah.log". For quite a while, I was just running things by hand under tmux because I couldn't figure out the exact method to make it work properly.
In comparison, doing the same in systemd was incredibly easy. This type of service is really easy to set up in a systemd unit file.
I'll admit that a lot of the problems with doing this in FreeBSD were likely my own lack of skill and/or understanding. Systemd doesn't require this level of skill/understanding to set up a simple service though.
In terms of benefits of systemd, the best one for me is logging. Having stdout/stderr from these processes go to the journal automatically is quite nice and querying it with journalctl is easy now that I've figured out how to use it.
When developing simple services (like the custom shim mentioned earlier), systemd greatly simplifies things since I don't need to worry about daemonizing, switching users, logging, etc.
chdir z && daemon -u y -o blah.log x
https://www.freebsd.org/cgi/man.cgi?query=daemonI probably ended up with something like this[0], which doesn't use daemon. Home Assistant also provides a FreeNAS example[1] which uses daemon.
For comparison, the systemd unit file that I ended up with is:
[Unit]
Description=Home Assistant
After=network.target
[Service]
Type=simple
Restart=always
ExecStart=/home/hass/homeassistant/bin/python /home/hass/homeassistant/bin/hass
WorkingDirectory=/home/hass
User=hass
Group=hass
[Install]
WantedBy=multi-user.target
[0] https://gist.github.com/damoun/96add58f60572cb12c11[1] https://www.home-assistant.io/docs/installation/freenas/
Now though, I really like the systemd service approach - it's so much simpler than hacking custom bash scripts, and it's actually really fully-featured too. For example, it supports lots of different service lifecycles, and even has a watchdog included that can restart failed processes.
I'm still less sure of some of the other parts, like journald, but I'm a systemd convert for services at least!
I've had frustrations with systemd, but the few times I've wanted to write an init script I was glad it was trivial instead of cobbling something together for traditional init.
Maybe because they /could/ be run directly if/when service existed without warning or penalty, it also wasn't widely used.
In practice, I hit it many many times in sysv-style systems. I've used systemd systems for much less time but haven't hit a similar issue since transitioning.
No more crazy init scripts, no more pidfiles, better templating, cgroups integration OOTB, better handling of restarts, isolation integration (private tmp, capabilities, ...), etc. I enjoy what systemd-the-init-process provides. I see it as a standardisation of what everyone did in a slightly different custom script into a known key-value format.
The real wow moment came when discovering socket activation which allowed systemd to take care of zero downtime deployments of web applications - for me that was reason enough to support the project.
Nowadays I'll use a container orchestrator to achieve the same thing, bit it was a good couple of years whilst it lasted.
In the last couple of days, I became a fan. Not for the many reasons which anyone can look up. But for systemd-nspawn, an awesome tool for spawning lightweight containers.
For instance, you can boot an ephemeral (tmpfs) copy of your system with systemd-nspawn -D / -xb. Several more examples are listed here: https://www.freedesktop.org/software/systemd/man/systemd-nsp...
I've felt the same way, but then again I'm not really missing SysV runlevels either.
Generally I don't find myself tinkering with init systems too often, and a simple list of commands to run on startup in-order suffices for most of my needs.
Many issues with systemd seem not to be with systemd-as-init-replacement, but rather with with systemd-as-kitchen-sink. It's the tight coupling that annoys many people.
Does udevd really need to be in the same repo? While there may be some nice things about journald, does it really have to be in the same source package? (And why can't it support remote logging with the industry standard syslog protocol? Now I have to run journald and rsyslog. And why doesn't it have an ACID file format that doesn't self-corrupt at times? Why couldn't they just use SQLite or OpenLDAP's LMDB?)
Compare the upvotes on the .service answers to the alternatives.
Also earlier today I wanted to find out about systemd timers and the first post was this, which thinks they're way better to debug than cron: https://mjanja.ch/2015/06/replacing-cron-jobs-with-systemd-t... - mainly because they have an options to 'run this job right now, in exactly the same environment' - that's an excellent feature, and something cron should have had ages ago.
Getting an impression of systemd from HN is like getting an impression of politics from Twitter. You'll be exposed to a lot of views, but not a wide variety of them.
Then there is systemd-networking. So much easier, simpler to configure than netplan.io and network manager. Again, lots of information presented through journalctl and I never have to deal with networking issues.
I have systemd-analyze which gives me a clear picture of boot times and security issues. From my perspective, these kinds of things are a step above what existed before.
Don't like init style shell scripts? Replace it with something like systemd's service descriptors. It would be fine with me. Now, using this as merge god knows how many different services into one badly written piece of crap that uses 100-1000x more energy than the services it replaces is not an argument that you can win.
SysV certainly wasn't and isn't the smooth experience either and took some time to grasp.
I was a relatively early adopter to CentOS 7, and I have yet to run into an issue that I could pin on systemd across any of the Linux systems I use or maintain. Then again, my requirements tend to be pretty "normal" and not terribly taxing - at most I crank out a couple of server-specific service and timer files and get on with my life.
I totally understand where some of the critics of systemd are coming from. No matter how you slice it, more lines of code is invariably going to result in more bugs. And the development team does seem a bit myopic at times - though nothing deserving the death threats they get.
However, there were some clear disadvantages to the old-school sysv init, and I thought that it was important that Linux settle on a new standard init system and do it without yet another instance of fragmentation happening, like with rpm vs deb or Flatpak vs Snap or countless other examples. Systemd happened to be it, but honestly I didn't even care what the init system was as long as something finally "won" and didn't introduce yet another schism. In fact, I highly suspect that if something else had won people would have doubtless found an axe to grind with that one too.
Is systemd + its services and timer files, bigger than sysvinit, cron, atd, pm2/forever, previous device managers, all the various bash scripts, and everything else it replaces?
I don't know either way. I'm just not completely sure Systemd is bigger.
(It's possible other "modern" inits could also do this well, and I do continue to believe that systemd in particular is not the best code even if I agree with its architecture, but it's the best out there as far as I know of the ones that actually exist. Many years ago I ran into an unsolvable bug with Upstart and a particular complex setup that I couldn't even minimize into a useful bug report. I've had problems with systemd, same as I've had problems with, say, Linux and glibc, but they've all been quite solvable.)
The problem comes when you discover that the Systemd version of a service is missing one crucial feature you need or is buggy, so you need to run an external service instead and it starts fighting you tooth and nail. Or when something breaks and tracking down the issue is hindered by it being hidden in behind a bunch of layers of indirection.
For example, try configuring an Ubuntu 16 box to randomize its MAC address before searching for WiFi access points.
Indeed, things like start order and depedencies can be completely broken on systemd. Many problems arise from incompatibilities with legacy systems and interfaces.
I had a big headache last time I had to configure systemd to wait for dhcp before running other stuff because the systemd targets for networking were not working - part of the problem was a legacy interface for network configuration.
That said, it should be eminently possible to have something like service files, and not systemd.
And I must admit, reading the back and forth on the linked issue, I still have zero confidence in systemd upstream as maintainers of systemd.
That, more than the disturbing complexity, "eat-the-world"-tendency and tight coupling (with lip service to separation, by having tightly coupled parts of the overall monolith be "separate" as individual binaries) worries me most about systemd.
I'd imagine the maintainers are extremely busy and may have a sense (or incentive?) to respond - in any way - to issues swiftly, and that closing them may in many cases be the most time-efficient way to keep their own work under control.
That said I agree that situations like this are frustrating for users, especially when they expose what appear to be genuine and reproducible issues.
But not seing that the design seems fundamentally broken, that's more of an issue (message amplification to millions of messages, the fact that it's questionable if systemd even can use the information for meaningful ordering/dependency reseloution.... When it's not systemd doing the mounting..).
I no longer have to learn a three-dimensional grid of init systems (init system, distro, distro release).
If this isn't a positive experience...
Systemd requires a small initial study effort in order to be understood.
If you're unwilling to do that kind of things (is: reading the fine manual), you might as well go buy a macbook.
After poking the internet:
After altering fstab one should either run systemctl
daemon-reload (this makes systemd to reparse /etc/fstab
and pick up the changes) or reboot[^1].
So, the summary with systemd: it will break something for everyone sooner or later. It hopefully won't be something you can't fix, but it will make you very angry and in immediate need of a tea break.BTW, this is not the first one for me, but certainly one of the most frustrating ones, and not because of the behaviour: because of the lack of error messages on stderr or at least in syslog. The messages in ~~syslog~~ journald were casual.
[^1]: https://unix.stackexchange.com/questions/169909/systemd-keep...
/etc/fstab is parsed by systemd and converted int .mount unit files, so that they can be mounted when a) asked for it (auto/noauto) and b) as soon as possible.
think of the _netdev flag: it signals systemd that a mountpoint is network-based, and thus it will attempt to mount it only after the networking target (instead of waiting for it to timeout at boot time, fail , and possibly letting processes start reading and writing from/to an empty directory).
Running daemon-reload is generally system(d)-wide though: whenever systemd configuration is altered (think adding/removing/updating an unit file) you should let systemd know about that by running systemd daemon-reload (think of apachectl reload).
"Long time" is also relative. I've been updating /etc/fstab directly for over 25 years.
Yeah... there's always a "good reason" with systemd, isn't there?
Still doesn't account for the lack of error message.
In the end I gave up. I can't put the blame on one part specifically but it was not a good experience.
Other than that I don't see much difference on defining a service with upstart or with systemd, for my basic use.
I don't like the feature creep though.. we have now systemd, systemd-udevd, systemd-journald, systemd-networkd, systemd-timesyncd, systemd-resolved, systemd-homed and who knows what else. worse some of these are kind of irreplaceable in the systemd ecosystem.
SystemTap and other kernel instrumentation stuff are of course invaluable sometimes, but I just find strace easier, even if it requires looking at and divining meaning from a lot of mystical output lines. :)
That said, the amount of processes, DBus messages, and system calls that happen on even on an idle system these days is crazy. (On Windows it's even worse, something is doing I/O all the time. Diagnostic that, ETW this [event tracing for windows], you can't stop most of them.) There is just too much broadcasting going on. And it's especially maddening on platforms that provide subscription for events. (Don't broadcast the same signal over and over. Let new clients get the current state and wait.)
I enabled various debug levels for it (it can spew out an ungodly amount of data, but exactly what you mentioned is missing, the "what are we waiting for" info) - I had a problem with shutdown taking too long.
That beeing said is there a good a writeup / deep dive into systemd concepts? Really cool would be something like a book - the official docs are quite spread out and it's either rather simple high-level stuff or basically: read the source.
Too bad it basically doesn't work in Ubuntu.
systemtap is really like heart surgery on kernel internals that are not stable. So it is expected that it breaks from time to time.
Systemtap is developed by RH people (unless something has changed). So in RH affiliated distros it should just work. But as said, my experience is that they are willing to help with other distros, too.
Or is there a reported breakage in Ubuntu these days that nobody is reacting on?