Systemd: Enable indefinite service restarts
michael.stapelberg.ch
michael.stapelberg.ch
> I’m not sure. If I had to speculate, I would guess the developers wanted to prevent laptops running out of battery too quickly because one CPU core is permanently busy just restarting some service that’s crashing in a tight loop.
sigh … bounded randomized exponential backoff retry.
(exponential: double the maximum time you might wait each iteration. Randomized: the time you want is a random amount, between [0, current maximum] (yes, zero.). Bounded: you stop doubling at a certain point, like 5 minutes, so that we'll never wait longer than 5 minutes; otherwise, at some point you're waiting for ∞s, which I guess is like giving up.)
(The concern about logs filling up is a worse one. It won't directly solve this, but a high enough max wait usually slows the rate of log generation enough that it becomes small enough to not matter. Also do your log rotations on size.)
Especially that service startup failure is usually not something that gets fixed on its own, like a network connection (where exponential backoff is (in)famous). A bad config file, or a failed disk won’t recover in 10 minutes on its own, so systemd’s default makes sense here, I believe.
At my work, we have a simple philosophy for this. The tester is allowed to (on the test system): toggle servers' power; move around network cables; input bad configuration; etc; in any permutation he wants. So long as at the end of the exersise everything is setup correctly the system should function nominally (potentially after a reasonable delay).
There should, of course, be a system level dashboard that notifies someone there is a problem; but that is unrelated to the server internal retry logic.
See the following figure, where I plot the # of requests over type: https://i.imgur.com/PNUFhjc.png
And that is why you should use 0.
---
I am not an expert, but I got nerd-sniped.
I have made some simulations[1] and done some napkin math[2], and I would summarize as follows:
I don't think there is any globally optimal, I think it depends on your exact circumstances, and which of {excess waiting time, excess load} you want to optimize. Defaulting to 0 is a lot easier to implement, and is at most a factor 2 worse than the optimal lower bound w.r.t. waiting time vs. load. You may also consider [0.5 * maximum, maximum], which trades some excess waiting time for less load. Your suggestion is a similar heuristic one might use depending on their exact circumstances.
[1] https://gist.github.com/Dryvnt/1984d9389ae7386127f5e8998bf52...
[2] Consider a family of random bounded exponential back-off strategies with ultimate upper bound U as follows:
Strategy X(z): Random of [z * U, U], where z is constant, 0 <= z <= 1
There are other families of back-off algorithms with other characteristics, and I am not considering or comparing those, just this family. Note that strategy OP suggests is X(0). Consider that X(1) is non-random, which is undesirable.Simple probability tells us
avg(X(z)) = (1 + z) * U / 2
Server load ~= frequency, frequency is inverse duration, so load(X(z)) ~= 1 / avg(X(z)) = 2 / (1 + z) * U
The difference in load between any two X strategies is load(X(z1)) / load(X(z2)) = (1 + z1) / (1 + z2)
Since z1 and z2 are constant, this relative load is constant. Due to the bounds on z, the largest possible relative load is load(X(1)) / load(X(0)) = 2 / 1 = 2
Now consider your suggested strategy Strategy Y: Random of [L, U], where L is the last choice
Note that Y = X(y), where y = L / U.
In fact, Y approaches X(1) exponentially fast, since the difference between L and T is halved each step, on average. So your suggestion still falls within this at-most-factor-2 difference. Exactly where just depends on the outage length.As always, https://xkcd.com/356/
Yes it does a lot of stuff for you and in others I had to write custom scripts but it was much more understandable and maintainable long term. Sadly systemd won and now i build my own OS without it.
I'm glad systemd "won"; it's much more maintainable IMO than shell scripts written once and forgotten about (until they break).
With one exception... I tried to set up a inetd-like service where a script was run on TCP connection with the output going out to the socket. It was tricky to get set up, but eventually I got it working... or so I thought. A few days later the system was super sluggish and I found that every connection that had been made to this service was left open and consuming resources. Switched that to an socat running under systemd and it's been fine since. Never could figure out what exactly the deal was.
And I would guess sysadmins also don't like their logging facilities filling the disks just because a service is stuck in a start loop. There are many reasons to think a service failing to start multiple times in a row won't start. Misconfiguration is probably the most frequent reason for that.
I'm sure there are exceptions to this. For those, set Restart=always. But it's an absolutely terrible default.
This is one of those things.
The 'After=' and 'Requires=' directives address this.
Depends on a mount? Point those directives at a '.mount' unit.
Depends on networking, perhaps a specific NIC? Point those directives at 'systemd-networkd-wait-online@$REQUIRED_NIC.service'
Point being: declare these things, don't wait for entropy to eventually become stable.
Back on point though: don't expect the 11th restart to work when the last 10 didn't.
Contrived examples are contrived, it's solved. Declaring dependencies.
With the requirements properly laid out we've avoided restarting in a loop and a bit of robustness
There's also 'PartOf=' which can help make the relationship bidirectional
Starting up, noticing that the environment doesn't have what you need yet and dying quickly appears to be The Kubernetes Way. A scheduler will eventually restart you and you'll have another go. Repeat until everything is up.
The kubelet operates the same way afair. On a node that hasn't joined a cluster yet, it sits in a fail/restart loop until it's provisioned.
(You can guess how we noticed the problem…)
Also logrotate. (And bounded on size.)
Before=systemd-user-sessions.service
This means that as long as systemd is trying to (re)start the service, nobody can log in. Which is a problem with infinite restarts.
It's still pretty easy to accidentally set up an infinite restart loop with the default settings if your service takes more than 2s to crash.
When the given (transient) condition goes away (either passively, or because somebody fixed something), then the service comes back without anyone needing to remember to restart the (now dead) service.
By way of example, I've run apps that would refuse to come up fully if they couldn't hit the DB at startup. Alternatively, they might also die if their DB connection went away. App lives on one server; DB lives on another.
It'd be awfully nice in that case to be able to fix the DB, and have the app service come back automatically.
This doesn't mean that you don't investigate, it just means that you have an additional guarantee that the system can automatically eventually recover.
If you set a limit on number or time or restart, what's a reasonable limit? That will be context dependent, and as soon as it's more than a few minutes, it may as well be infinite.
(And if you want a dumb unit system, there are plenty of options which will run just fine under systemd as a single unit so you never have to actually use systemd for your own services even if you're forced to use systemd for the overall system for whatever reason)
My favorite new misfeature is PulseAudio. These geniuses actually built code for a multi-user, multi-tasking OS...which will only run for ONE user, and then only if that user is logged in. So forget running cron jobs, and sounding an alert if something needs attention.
This is all code produced by FreeDesktop[.]org. Thanks to them, your industrial strength, mission-critical server OS is now only suitable for single-user desktop systems.
Nobody accidentally runs "systemctl restart" too fast, when such a command is issued it is clearly intentional and should be always respected by systemd.
# Get the number of restarts for a service to see if it exceeds an arbitrary threshold.
systemctl show -p NRestarts "${SYSTEMD_UNIT}" | cut -d= -f2
# Get when the service started, to work out how long it's been running, as the restart counter isn't reset once the service does start successfully.
systemctl show -p ActiveEnterTimestamp "${SYSTEMD_UNIT}" | cut -d= -f2
# Clear the restart counter if the service has been running for long enough based on the timestamp above
systemctl reset-failed "${SYSTEMD_UNIT}"Then you could have the default be 100ms for one-time blips, but (after a burst of failures) fall back gradually to 10s to avoid spinning during longer outages.
That said, beware of failure chains causing the interval to add up. AFAIK there's no way to have the kernel notify you of when a different process starts listening on a port.
You can use mandatory access control for this.
AppArmour or SELinux are examples.
Unfortunately they are hard, not sexy and sysadmins (people who tend to do not sexy hard things) are a dead/dying breed
Would the ExecCondition be appropriate here, minimally, with a script that runs `lsof -nP -iTCP:${yourport} -sTCP:LISTEN`?
Obviously if systemd opens the port for you it's easy enough (in this case, even across machines), but otherwise you have to do a sleep loop. And I'm not sure how dependency restarts work in this case.
ExecCondition must moves the spin to systemd, and has more overhead than doing in your own process. There's no point in gratuitously restarting after all.
Available since systemd 254, released July 2023 (only 1 release since then). Huh, has release rate severely slowed down?
For startup, I’d argue the proper way is for the process to bind the socket before forking as a daemon.
With such a design, one can launch a list of dependent processes without worrying how long they each take to start up. No polling loops needed!
Of course, that requires some careful design — “fork and forget” is too appealing. The process would be responsible for creating its PID file after forking but the socket before…
Alternatively, an IPC notification could be used but would require some sort of standardization to be generally useful.
The last thing I need is emergent behavior out of my service manager.
Hence.. why I called it a fragile mechanism.
Yes, there are ways to deal with anything, but less is more, this is unix afterall. I can't love having a bunch of options, in several places, that may override each other, most of which I would never use for any practical reason anyways.
Reminds me of the type of systems I abandoned in favor of linux.
I think you can even have systemd reboot or move the system into a recovery mode (target) if an essential unit does not come up. That way, you can get pretty robust systems that are highly tolerant to failures.
(Now after reading `man systemd.unit`, i am not fully sure how exactly restarts are cascaded to requiring units.)
OnFailure makes it easy to implement more complex restart or notification logic.
That violates everything I ever enjoyed linux for, I left Windows because it thought it knew better than me.
Though I agree the default is poor, if it bites you more than once, write a standard template to reuse for your unit files and/or write a wrapper that calls reset-failed for you. Systemd is far from perfect, but this is a minor nuisance.
You missed the other direction of the relationship.
I posted elsewhere in the thread on this, don't rely on entropy. Define your dependencies (well)
After=/Requires= are obvious. People forget PartOf=.