Systemd auto-restarts of units can hide problems from you
utcc.utoronto.ca
utcc.utoronto.ca
This is just another reason to track and alert on error rates.
It also depends how sensitive you are to failures, and required availability.
But in general this sounds more like: “you need observability on error rate”, more than anything else, and that includes notification and alerting on set thresholds that seem aberrant.
Any supervisory process, systemd, supervisord, kubernetes, etc should absolutely be making those restarts visible to the administrator so that they can resolve it.
The restart is just in hopes to keep the service available (symptoms in check) until the problem is actually fixed (disease is cured).
And we should be thankful. We don't have to be aware of every single hiccup a service experiences. Even if (and specially if) you are the one responsible for keeping it up. Modern distributed systems are far too complex for us to care about every minutia.
If something failed over, and nobody noticed, is it a problem? The answer is _maybe_. How often does that happen? Is there a pattern? Is it getting more frequent? Is that happening more often than predicted?
At work, we blow up entire VMs if they fail their health checks. They can fail for many reasons, mostly uninteresting ones. And the customers don't even notice and SRE doesn't usually care. They only become relevant once there are anomalies. When you are managing thousands of instances, self-healing is required.
Error rate may not even be impacted in a significant way when those systems restart, it is often a tiny increase in a deluge of requests. So you also need observability on self-healing events to catch trends. When self-healing fails, it tends to do so pretty catastrophically.
> and alerting on set thresholds that seem aberrant.
I wish we would remove 'thresholds' from our alerting vocabulary. Very often we end up setting simple and completely arbitrary thresholds that don't actually mean much - unless it's based on SLOs and SLAs. Generally we don't know what those thresholds are supposed to be and just guess, unless it's capacity. But even for capacity: 0% disk space free is obviously bad; what people will do is set some threshold like ("alert when 80% disk is used"). Then you get page once that threshold is crossed. Is that an emergency? I don't know, maybe it took 5 years to get there. However, if you set an alert that says "Alert me if the disk will fill up in <X> days at the current rate", you can then tell if it is an emergency or not.
I suspect that you would agree with this given the use of the word "aberrant", which to me implies anomalies. But in many contexts, people think that "alert every single time CPU utilization crosses 90%" is what we are talking about.
Somehow I've never come across this before, and it's really unlocked something for me with how I will now approach alerts
You really care about time to fill (speed), and whether your system has gone nuts and is logging faster and faster than normal (acceleration).
Disk usage is X says do little.
Disk utilization rate is Y therefore will fill by YYYYMMDD is great.
Disk utilization rate change has gone off the chart so a process crapped itself and will fill the disk in 5 mins is life saving.
We have computers to use their CPUs, RAM, and disk. Using it isn’t weird.
Detecting actual unwanted or dangerous behavior is bad.
For bonus points auto-remediate it (eg take the node offline so the workload moves to another node) and alert when your overall capacity of your service is potentially going to be impacted or is near a limit, for some version of near. Being human I much prefer “you have hours to fix this” than “I’m broken fix me”.
RED, Golden Signals, etc. there are plenty of nice frameworks for how to think about utilization and performance of your service that you can use to manage capacity and health.
CPU at 90 is not one of them.
As soon as someone suggests it I always ask for how long? On which processes? on which boxes? What is the “perfect” CPU utilization? Then make them think about why we wouldn’t want to use the resources we paid for.
Of course not. Log it, and restart.
> in the future I'm going to want to remember to check for this
Whenever you think that sentence, you should notice this as a red flag and re-think your approach. You will forget. And if not you, then somebody else in your team. You need automation for things you can forget, otherwise your mental checklists will grow too large to handle and are just a distraction.
Away from things that can be automated and are important you have a checklist and tick things off with a pen. Add $this to the list.
Atul Gawande is worth reading in general and on this topic. [1] He turned it into a book I haven't yet read.
[1] https://www.newyorker.com/magazine/2007/12/10/the-checklist
Yes, and to have alerts, I have resorted to write my own small tools [0][1] at the end of the day. Railgun proved to be very useful, but for smaller things, I'm writing the second one.
I'd argue that unexpected restarts should alert beyond a threshold. Alerting on every occurrence is too noisy. If an individual unit failure causes a service disruption architecture improvements are needed.
my preferred way to alert on this is to use the process start time that's included by default in most process-level Prometheus metrics (and is trivial to implement yourself, if you need to):
> curl -s localhost:9100/metrics | grep process_start_time_seconds
# HELP process_start_time_seconds Start time of the process since unix epoch in seconds.
# TYPE process_start_time_seconds gauge
process_start_time_seconds 1.68957669109e+09
on a stable system, this metric will be very close to static. you can feed it through the PromQL changes() function to get a count of how many restarts have happened in a given time window.in my experience, for anything with an "alert if it's down for X minutes" rule, you probably also want an "alert if it restarts N times in Y minutes" rule.
"That process that restarts spuriously once in a while has restarted again" stops being noteworthy very quickly if it's deemed low priority, and often even might lead to alerting thresholds being change because it becomes a nuisance.
"That process that restarted every now and again now restarts 5 time as often as it used to" on the other hand is a lot more likely to get attention.
So, I can see two different aspects to this problem:
* Lack of alerts.
* Lack of RCA automation.
For the first one -- it would've been nice if systemd had a "brother" program that could be used for alerts, so that no custom solution was necessary and that services could properly report intermittent problems etc.
The second is contingent on several factors: re-envisioning error handling, declarative debugging and general popularization of the concept. There are several major problems with error handling in system programming languages today. Due to poor language runtime design, programmers learn to think about any and every error as essentially fatal. Recovery is typically seen as impossible, and if there's any sort of recovery code put in place, it's usually the one that tries to do the cleanup and start fresh. It's never really about fixing the problem. So, popularizing something like Common Lisp restart system with the ability to traverse the program stack back to the failing frame would've been a good first step in this direction.
Declarative debugging, on the other hand, could be taking another step forward, where it could be made into a separate program which describes a complex recovery scheme. I.e. the idea of declarative debugging is that the programmer needs to describe the program in terms of constraints that should be checked when the program fails to identify the problematic place. The step forward would be to add automation for the cases when the failed constraint is discovered.
Database management systems don't bother with it anymore, your almost stateless deamon has no hope of gaining anything from it.
But wide logging and error identification are things too; and very useful ones. And those plug on the same spots that were created for error recovering. Because of this, those are also not very common on our current system software.
Nobody in system (storage / network / compute) or any other specialized highly-reliable system -- unlikely.
In highly reliable systems this is already being done, except it's usually an afterthought, ad hoc, functionality added without any system or framework, and so it's hard to judge its performance because it's often partially interactive, and very uneven across different components.
Any system that advertises itself as "HA" (for highly-available) will have some sort of a protocol for complex error recovery, except it's usually not a single program that's explicitly written with the purpose of definition and documentation of recovery process, instead different bits of this functionality are embedded in many different places across the entire surface of a program.
A very common example that I had to implement myself multiple times: HTTP errors. These are usually very poorly reported by the HTTP client libraries. TLS handshake failed? -- you get some numerical exit code instead of proper error report, and so once you realize that this error happens every so often, you start writing code that digs up the internal state of the TLS library at the time of the error, finds the certificates involved and tries to trace down the chain of events leading to the problem.
Similarly, various parsing and de/serialization errors... this stuff is usually atrocious in popular libraries. And in order to figure out why some JSON didn't parse the way you wanted, you need to write a lot of code to investigate and report the error properly. The most common cases I've encountered is when a client program expects a server to reply with a structured error message, but the message comes back botched, or an HTML page is sent instead etc. A lot of these things are quite predictable after some time of using the client, but in a situation where you wanted to properly report an error that was supposed to be encoded in some JSON message and instead you got an HTML page -- typically you end up reporting a wrong (and useless) error of failure of parsing HTML as JSON. To prevent that, you'd need a complicated recovery scheme that would anticipate such errors.
This is actually part of the self-healing aspect. One way to see if a service restarts is using the following command:
sudo systemctl show [servicename].service -p NRestarts
I prefer to have this set as Restart=on-failure
Also note, there are OnFailure to trigger an additional log message, or recovery service, and FailureAction to allow actions to happen, such as the suggested 'reboot' ;-)For reference: https://www.freedesktop.org/software/systemd/man/systemd.uni...
sudo systemctl show [servicename].service -p NRestarts
any way to do that on all services aside from looping ?You also don't want to deal with all of them. For this you could otherwise just use journalctl. You also have OOM issues, and this has its own facility.
...except when that causes even more problems! one example i remember from a hobby project was a worker service that would make requests to an external service's API. I'd implemented a throttle in the worker to limit the number of external requests per second made against the external service, where the throttle state was stored in process memory and not persisted anywhere -- seemed to be a pragmatic design tradeoff. you probably see where this is going.
there was an exciting interaction where an external API response caused the worker's API client to throw an unhandled exception, which took down the worker service. then systemd would diligently restart the worker service immediately per the restart policy. the throttle working state had been lost, so upon being restarted the worker service immediately fired another request to the external API, which then issued the same API response, which caused the worker service to fail in the same way, ... luckily noticed by systemd and immediately restarted.
combine with a lack of monitoring + alerting and you get a mechanism where your worker service can make about 100,000x as many external API requests over a few days as it was meant to.
You can then also set "OnFailure" to trigger another unit if the failure state is reached, e.g. to trigger a notification.
E.g.:
[Unit]
...
OnFailure=notify-failure@%n.service
[Service]
Type=simple
Restart=on-failure
RestartSec=5
..
StartLimitBurst=5
StartLimitIntervalSec=300That said, there exist options to determine that restarts are happening too quickly, indicative of a misconfiguration or that it just won’t work, period. See Timeout*.
If this is not what you want, you can always change the restart policy, or specify the maximum number of restarts.
Could be, though, that some software ships with crappy defaults which make it crash all the time and systemd just always restarting it and reporting everything is ok. Not a problem of systemd per se, though.
It was so common to not restart that many services now have a "clear error logs on startup" policy.
I think we either started demanding more from our services or just accepting worse quality. And with the mainstream accepting silent restarts it’s only going to get worse.
You should have Telegraph running on every system you have collecting stats and sending it to InfluxDb with sample rate around 15s and +- 5s of jitter.
Each critical process should have an uptime counter. You should create both Deadman alerts (No contact for 5m) and Uptime alerts (uptime falls below 60s).
s/Telgraf|InfluxDb/YourMonitoringSystemOfChoice/
Bonus: If you're running JVMs, you can expose _all_ of the JVMs stats to Telegraf via Jolokia, which can run as a Java Agent and requires 0 code changes. The JVM has very very detailed stats about it's health available by default, making problems very visible (if you look for them).
a) It's highly likely that your monitoring is not going to spot an occasionally restarting service
b) It's probably a good choice not to waste cycles on that sort of monitoring
An uptime counter counts the number of milliseconds a process is alive. So if it resets to a value lower than the previous recorded one, it's a blue star thats impossible to miss.
Are the restarts needed because there's malicious over-size content being sent to it, trying to exhaust the server resources?
If so, then "http.MaxBytesReader" might be helpful:
https://github.com/sqlitebrowser/dbhub.io/blob/5c9e1ab1cfe0f...
I went a step further and all of my servers do staggered restarts at the start of every month (previously daily/weekly, but once I worked most stuff out, a month is enough now).
No more update related issues catching me off guard when I suddenly NEED to do a restart in short order, no issues regarding manual processes or a service that doesn't start automatically after a restart/crash, or other things that you might not catch otherwise, load balancing etc. Even stuff like restarting the Swarm cluster leader/followers close to one another, yet expecting them to eventually communicate with one another properly and schedule all of the containers, pull them and launch them as needed, join them to the networks and actually communicate properly to them.
Knowing that restarts are inevitable (sometimes after a crash), might as well make sure that they work properly.
And when you discover a new failure mode and something indeed doesn't work, it's nice to test your monitoring and alerting, as well as learn how to deal with that particular failure.
Sometimes there are external demands that require the system to always be up or to minimize any downtime. In an ideal world, that requirement should be designed into the system architecture from the beginning, but in the real world that doesn't always happen.
Then the solution is redundancy. Especially with todays frequent downtime for security updates you can’t deliver a singular system that’s always up.
Sure, but you can fix that bug after it's been auto-restarted. Assuming this is a service that, say, customers rely on, your customers don't care about the details. They just care that they're able to use your service and get their work done.
Not having some sort of auto-restart or auto-healing system only serves to prolong downtime. Obviously you need to be alerted of these restarts so you can find the bug and fix it, but overall I'd value uptime over the fairly low probability that allowing it to auto-restart might expose some sort of security issue.
Of course there may be some circumstances where reliability is by far the biggest concern, so allowing something to crash and stay down may be preferable. I just think that's going to be a small minority of situations, for most people.
Well, obviously so if running under OpenBSD's rc system.
Pretty much all fault-tolerant systems feature recovery by strategies like auto-restarting things and raising alarms, so the OpenBSD people must know some secret that has eluded researchers in that area.
My solution: (Shameless plug) Create a small daemon that monitors the systemd-journal and email me a summary every 5-minutes (if any and not on blacklist).
If you're interested (it's BSD2-licensed): https://github.com/mpdroog/hfast/tree/master/contrib/deltajo...
I always thought a good monitoring system had three components. Polling, streaming data and log analysis. A lot of times folks don't bother with the logging and miss issues like what is described in the post.
Not everything needs to be monitored in a single way via a single platform. A cron job that greps/awks syslog for relevant strings and sends a slack message, email, sms, something would be a step function in the right direction.
if $syslogseverity <= 3 then {
action(type="ommail" server="127.0.0.1" port="25"
mailfrom="rsyslog@localhost"
mailto="root@localhost"
subject.template="mailSubject"
template="mailBody"
action.execonlyonceeveryinterval="3600")
}
While I generally like to dump all logs into elasticsearch/opensearch, it's also really handy to have this set up for one-off machines so that critical/error logs that usually come from things breaking & restarting still get seen.There is no builtin spreading so */5 job on 500 nodes will get you nasty spike in CPU load. Logging options are pretty much "parse emails and hope for best" and some syslog spam that's mostly noise if everything is working fine. And fucked up environment vars that trip many script-makers. No metrics whatsoever
The solution in our case was to build a tiny launcher for anything cron-related that delayed the script executions randomly within a range of milliseconds.
Usually when a service conks out, operations will just kill it and roll out another instance. Nice, but you have no way of finding out what actually happened then.
Sometimes this is caused by someone looking for exploits (and actually finding one), or other stuff you'd really rather know about.
Operation teams (and automation) have usually a primary mandate of availability above all, not to investigate any possible failure.
If you can't have that (and budget allows) keep unhealthy replicas alive and pull them off the load balancers.
In my experience option one works best, and makes option two redundant. Unknown unknowns are a thing, but that's true even when the unhealthy replicas are kept around.
Wouldn't journald have details? Of course those might not be hooked up to the main reporting flow, but the debugging info is available once you discover the problem
I know there are problems in the software I am running, but I don't what to think about it. I want it hidden from myself.
If you need alerts, you should config monitoring for the logs.
No offense to the dude if he happens to be reading this, but less quantity and more quality please. I'm sure you're very smart.
Alternatively, maybe HN can let us blacklist domains?
Although you do lose state when a process dies, if your processes are sufficiently lightweight and/or stateless, it's possible for the problem to be irrelevant until the next software release can garden/amend/curate it.
The win that you get in erlang/elixir/etc. is that you can just code the happy path most of the time, and then figure out which sad paths are the ones you need to deal with once you've seen production under load. If you've ever programmed in go, you understand how freeing that is.