My Philosophy on Alerting: Observations of a Site Reliability Engineer at Google
docs.google.com
docs.google.com
Think about how people react to smoke alarms versus car alarms. When the smoke alarm goes off, people mostly follow the official procedure. When car alarms go off, people ignore them. Why? Car alarms have very poor specificity.
I'd add another layer of car alarms are Not My Problem, but that's just me and not part of Dan's excellent original talk.
Smoke alarms are loud and typically near you (think about the one that goes off in the kitchen because of burning pizza in the oven). So you have to react or you are paralyzed. You fan it almost immediately. You have to.
Car alarm typically you can't do anything (unless it's your car) and the noise, while annoying, is not as close or as loud as a smoke alarm.
[1] Edit: Distance from and the decibels of the alarm and what you can do about the alarm. Of course there are also building alarms (follow procedure) and alarms in your own house "get the noise to stop!"
In corporate/organizational environments - often in buildings made entirely of concrete - I've had them go off very frequently, but almost never as a result of a dangerous condition - usually either a drill, a small fire that was immediately put out, someone pulling the fire alarm as a prank/protest or something else like someone burning popcorn or something.
I still think the point stands.
My point was that in a (nonadversarial) context where you have known triggers for false alarms, P(problem|alarm) falls dramatically once you have identified the presence of such a trigger, and it's entirely reasonable for our actions to reflect this.
(Though I think you're in a ridiculous situation)
If it's like every other smoke detector I've encountered, the bottom is to test the alarm, not to cancel it.
It may vary by country I suppose.
They do at least have a snooze button on them, but I was pretty close to buying a set of the Nest alarms out of my own pocket and replacing them until I saw the price tag.
Whether you have the ability to differentiate your particular tone of horn from all the others is a whole different story.
However, in some heavy-traffic routes like the Straight of Dover or Gibraltar the proximity alarm goes off constantly. So the commanding officers usually disable them, leading to a good number of accidents with fishermen boats.
Absolutely this. Our team is having more problems with this issue than anything else. However, there are two points which seem to contradict:
- Pages should be [...] actionable
- Symptoms should be monitored, not causes
The problem is that can't act on symptoms, only research them and then act on the causes. If you get an alert that says the DB is down, that's an actionable page - start the DB back up. Whereas, being paged that the connections to the DB are failing is far less actionable - you have to spend precious downtime researching the actual cause first. It could be the network, it could be an intermediary proxy, or it could be the DB itself.Now granted, if you're only catching causes, there is the possibility you might miss something with your monitoring, and if you double up on your monitoring (that is, checking symptoms as well as causes), you could get noise. That said, most monitoring solutions (such as Nagios) include dependency chains, so you get alerted on the cause, and the symptom is silenced while the cause is in an error condition. And if you missed a cause, you still get the symptom alert and can fill your monitoring gaps from there.
Leave your research for the RCA and following development to prevent future downtime. When stuff is down, a SA's job is to get it back up.
Have a look at the example at the top of page 5.
Why is a human doing this? Why isn't the tool identifying the cause through a pre-determined dependency tree and only alerting on the salient points instead of relying on a human intelligently parsing through this data when they've been woken up by a page at 3am?
My experience has shown me that this kind of parsing can be avoided by using the right tools.
Knowing immediately that it's the server disk (more realistically a raid array going into recovery mode) will save you a lot of time and effort troubleshooting what would appear as a sporadic slow response issue. There's dozens of potential causes for poor response times, of which a raid array in recovery mode is just one.
And once you know it's a raid array in recovery mode, you can then take immediate action, something you can't do if you are still busy troubleshooting a sporadic response slowness issue.
I feel that ultimately, there's no problem with monitoring for high level symptoms, but they should not be the goal state of monitoring. The goal should be to monitor all possible causes of problems to limit the troubleshooting the SA has to do at 3am when woken by a page.
Plus, you should be using a tool which properly silences high level symptoms if there's a problem with a system which is clearly identified as a parent. That is to say, a server being down will silence "db is not responding" alerts.
If your DB's disk is bad, your real problem isn't that the disk is bad; your real problem is that, for example, customers can't buy products from your site. Fixing the symptoms means making your customers able to buy products from your site, not replacing the DB's disk. If you had, say, a failover slave DB, the point of the alert is to tell you to activate the failover process. Replacing the disk is important, but not urgent in the same way activating the failover is.
(Note the interesting fact that all alerts will then end up being for things the system could do something in response to on its own. Failing over to a slave can be automatic. Alerts are, in effect, the system saying "I need a human to come help me stop this from happening, because I don't know how to stop it myself.")
My point is that alerting on a "cause" may not actually get you to the root cause, and maybe not even all that much closer than a symtpom.
Sorry, I should have been more specific. DB daemon being down is a cause, and should be monitored. Hardware down is a cause, and should be monitored. Network availability is a cause and should be monitored. Power outage... you get my drift.
I think the root cause of my disagreement with this document is the lack of a proper dependency tree in the alerting tool. Their tool appears to want to alert on any and every monitored problem, which necessitates limiting what you monitor for fear of a pager flood. A proper tool can address this problem correctly.
In that context, I find it persuasive based on my own experience. (And I wish I had the kind of dashboard he described available; I'd love to see a presentation or blog post about _that_!)
This omits the implicit "because robotic, scriptable responses should be dealt by robots and scripts".
You should have monitoring for DB is down. But you should page on "product is not working". Having a monitoring system that lets you quickly find what's wrong is extremely important, but you shouldn't be woken up / or distracted for unactionable alerts.
The systems that Rob is responsible for are decoupled from individual pieces of hardware (redundancy/fault-tolerance) and degrade gracefully (if one part fails only a fraction of the working set is affected). Otherwise the huge scale would not be manageable as some piece is always breaking.
So yes: if you run systems that are not somewhat decoupled from their base, cause==symptom and cause-based alerting is indeed the way to go. And you will be woken up by every single piece breaking. The way to improve this is decoupling, and then you'll also switch to symptom-based alerting.
It changes the mindset from "Failure? Just log an error, restore some 'good'-ish state and move on to the next cool feature." towards "New cool feature? What possible failures will it cause? How about improving logging and monitoring on our existing code instead?"
If its worth it to the team to write code with edge conditions that will errantly wake them up in the middle of night occasionally they can make that decision to put that tax on their own lives, not someone else's.
Then the SREs are in every single code review.
Making devs carry pagers certainly helps, but it's a mistake to think that it's a panacea, and that forcing devs to carry pagers will suddenly make them write code with perfect logging and perfect error handling.
That aligned everybody's interests: devs came to SREs for design review and code review, and were incentivized to not just throw code over the wall.
Example you get a server failure which affects a service, and you begin working on replacing that server with a backup, but a switch is also dropping packets and so you are getting alerts on degraded service (symptom) but believe you are fixing that cause (down server) when in fact you will still have a problem after the server is restored. So my challenge is figuring out how to alert on that additional input in a way that folks won't just say "oh yeah, this service, we're working on it already."
The amount of noise that starts coming in when there is some major outage (say some mainframe system fails) is ridiculous.
Right now where I work they solve it by throwing manpower at the problem tbh.
It takes a lot of work by the application owners all working together to really get a coherent picture of how the services are interdependant, but the applications are so large, old code, etc - normal problems I guess a lot of companies face, that its almost impossible to find people who have a complete end to end understanding of most transcations.
Side note: my only monitoring experience is with Wily - anyone have opinions to hare on it?
It's hard to tune them so signal to noise ratio will be high.
Correct me if I wrong, but AFAIK the current state of the art solution to the alerts/log-filtering problem is: "log everything & feed these logs into a real time search engine that produces dashboard/alerts". Like elasticsearch/kibana. No? Curious, is that the approach that is being used internally at Google right now? BTW, the article stated the problem/and desired outcome, but not the solution. (?)
On the other hand, by logging everything as text, and then running intellegent/structurizing real time search engine over logs one can make/modify these decisions at a later time. And it can be done both by devs/ops, without touching the source code!
That's the good thing about this setup, you have all the logs from all your applications (think like custom text logs from your routers, your custom applications, temperature sensors, syslogs, windows servers) aggregated in one place. And when something happens (at a particular moment in time, or with a particular machine, or with a particular key) suddenly you are able to search/drill down and locate the actual cause. And maybe even configure a dashboard or make a plot that would show when this problem was showing up.
Scalable real time search engines with the ability to create trends/dashboards is one powerfull toy ;) It is ridiculuos and silly. But it is an immensely powerfull approach.
You want blackbox monitoring (for close-to-user experience) AND whitebox monitoring (which provide diagnostics of internal state for debugging). True blackbox monitoring is often pretty unreliable so you are usually better of alerting on whitebox reported state of end user perceivable variables, e.g. HTTP error codes, latency and so on.
State of the art is to report a staggering amount of data about the internal state of a server. I mean a lot. 10s to 100s of times the number of parameters you are probably used to seeing.
Both metrics and "logs" can be expressed as events. Those are like points in space. An incident could be like a line; a 2d event with a start and duration.
- is the webserver running - is it responding on port 443 - does it return HTML' and maybe - 'If I submit a search, do I get a result back?'
Nagios scripts are responsible for everything: opening network connections, querying system internals, collecting metrics, interpreting results, and boiling it down to a number between 0 and 3 and an unstructured text output to stdout.
A few of us understand that what we need is a more structured, data driven approach. Collect base metrics first, build a time series, apply a projection, and feed that projection in to a system that understands the actual failure condition.
As an example, imagine you're monitoring /. Nagios runs NRPE, NRPE looks at df /, and if it's 85 percent full (by default), sends a warning page. At 90 percent it sends a critical page. A smarter system collects the df / results, delivers it to a central timeseries database. The new data point is used to create a new projection, and the new projection is used to determine the time to an actual failure. The system above might have an idea of how long it takes to respond, repair, and resolve and issue a page when the disk will fill up if not responded to within 4 hours. That's the ideal solution, IMO.
It doesn't exist, AFAIK. There's a massive backlog of scripts that were written in the monolothic Nagios model that need to be rewritten, and thus this newer better version is always imaginary.
We've actually implemented exactly this at my current workplace. We have a nagios check that queries graphite and calculates a "days until full" value, and alert based on that. We have similar checks for monitoring other infrastructure. These checks take a lot of work to get the right calculation and threshold values, but once they work, it's pretty great.
Generally speaking, it works out pretty well for us. If you're a Splunk query guru you can also correlate and/or combine multiple disparate logs in elaborate ways to create more complex alert conditions.
The same can presumably be done with Elasticsearch/Logstash/Kibana.
We're actually security incident response, not reliability incident response, so our goals and methods differ a bit but the core concepts are all the same.
Doing this has taken us, in the past, from 400 pages a day to under 100 a day, over the course of a week's worth of effort.
This is an excellent point that is missed in most monitoring setups I've seen. A classic example is some request that kills your service process. You get paged for that so you wrap the service in a supervisor like daemon. The immediate issue is fixed and, typically, any future causes of the service process dying are hidden unless someone happens to be looking at the logs one day.
I would love to see smart ways to surface "this will be a problem soon" on alerting systems.
If you (or anyone else reading) ever want to talk about what "this will be a problem soon" might look like in the future drop me a line: dave@pagerduty.com
The subcritical alerts I think of are more things like "Well, the database is _getting_ full, but it's not full yet." Or to borrow someone else's example, "we put in this daemon restarter when it was dying once a week; now it's dying every few minutes and we're only surviving because our proxy is masking the problem but soon it's going to take the whole site down."
The other useful tip I have is to put URLs to internal wikis and/or tickets in the alert body. We write documentation for these to a 3AM standard: if I can't understand it immediately after being woken up at 3AM, it's not clear or actionable enough.
There is no reason anyone should ever run out of disk space if they alert on the trending rate of disk space [rather than the actual amount of disk space used]. But this applies to so, SO many things other than simple resource exhaustion. Seeing the trends is useful to alerts, but it's also useful to humans who can review them weekly and plan for the future.
In a previous position, we had a custom ticketing system that was designed to also be our monitoring dashboard. Alerts that were duplicates would become part of a thread, and each was either it's own ticket or part of a parent ticket. Custom rules would highlight or reassign parts of the dashboard, so critical recurrent alerts were promoted higher than urgent recurrent alerts, and none would go away until they had been addressed and closed with a specific resolution log. The whole thing was designed so a single noc engineer at 3am could close thousands of alerts per minute while logging the reason why, and keep them from recurring if it was a known issue. The noc guys even created a realtime console version so they could use a keyboard to close tickets with predefined responses just seconds after they were opened.
The only paging we had was when our end-to-end tests showed downtime for a user, which were alerts generated by some paid service providers who test your site around the globe. We caught issues before they happened by having rigorous metric trending tools.
It certainly shares some things with end-to-end testing, and blackbox monitoring is very useful for finding high level problems with any complex networked system.
Blackbox monitoring is (imho) only appropriate for 3rd parties. If it's part of your company, it shouldn't be a black box; that means someone got lazy and didn't demand the devs provide an API.
Also, i'm sorry but this really gets to me: at what point are we talking about 'at scale' ? I think it's whenever tons of money is riding on your site's availability and an unexpected failure causes customers to complain. Immediately VPs start screaming "WE NEED TO SCALE UP!!" and then they mandate some half-assed implementation of the comprehensive monitoring solution they claimed was unnecessary just a month before. But maybe i'm just jaded.
I can totally understand why SMBs have rotations. They have less staff. But a monster corporation? This seems like lame penny pinching. Heck for the amount of effort they're clearly putting into automating these alerts, they could likely use the same wage-hours to just hire someone else for a shift. Heck with an international company like Google they could have UK-based staff monitoring US-based sites overnight and visa-versa. Keep everyone on 9-5 and still get 24 hour engineers at their desks.
Having the team feel direct pain is a great motivator for building robust applications - if you know that that quick hack could lead to you getting paged at 3 in the morning you are far more likely to seek out additional solutions(anecdotally speaking) - it means you also have weight within the team to make the right engineering calls.
The team can bond over operations - it sucks being oncall, everyone knows that so everyone tries to make it less sucky, by clearing ops queues before handing off or making ticket messages a bit nicer.
When you start having ticket-free weeks or months it is an awesome feeling, the service works and is robust and your team can spend that time writing new stuff.
Additionally, when you have a small team building/maintaining a service it makes far more sense to rotate the oncall responsibility between them rather than an external engineer.
But the OP is a dedicated reliability engineer, they don't build the actual applications and couldn't given the size of the company. Essentially you're being punished for other team's generated issues.
You getting punished when your own stuff breaks is more a small business issue, not something corporations have.
why do you think that should be the case? I'd argue quite the opposite having done this in a large corporation with pretty decent success.
to be successful, it requires that your monitoring isn't noisy; most issues are repaired in an automated way; then when you get a "page" it is something that someone deeply familiar with the code / service is the best person to react to it. it will drive down the MTTR. and it typically gets the root cause fixed sooner in the code. these folks need to be engineers, but they don't all need to be just the devs, since internet scale services just are noisy and you need to have some randomization buffer. but those engineers need to be dedicated to the service and not a central org, where they aren't going to know how to look at logs, have depth on the intra-service interactions, the latest changes, etc. to me, this is also the best definition of "devops".
For a more detailed look at Google's SRE operations, watch Ben Traynor's excellent talk "Keys to SRE": https://www.usenix.org/conference/srecon14/technical-session...
As soon as you take the view of "This is crappy, but keeping it running is someone else's problem" then everything suffers (product quality, engineering quality, and reliability).
It was exhausting as the people on call didn't have the power to actually fix the system. Instead, they would have to walk someone else through the steps over the phone. If it took a code change to fix the system, too bad that was at least two weeks of red tape, and every single time the error occurred they had to page the person on call to walk the person through over the phone to verify it was the same error and nothing could be done.
Think "big business," "division of responsibilities," "accountability," inept management who never had to suffer under the policies they demanded, and a toxic culture which was proud to "give everything they have for the product."
Luckily, I don't work for those shitty companies anymore.
If you are interested you can also get my point of view from my Velocity talk on Monitoring without alerts. https://www.youtube.com/watch?v=Gqqb8zEU66s. If you are interested also check out www.ruxit.com and let me know what you think of our approach.
On the other hand, if you set up an IMAPS account for each person's alert address, and then have their smartphone use that account to check email, reliability is quite high. (Strangely, sysadmins tend to have wifi in their houses and good data plans, especially when the company pays for it.)
Unacknowledged pages escalate to a secondary oncaller (e.g. if the oncall is out of range/in a tunnel, under a bus) and tertiary depending on configuration (and then it loops, or falls to another rotation, again depending on configuration). The code and services that do escalations is deliberately and carefully vetted to have minimal overlap with production systems (who's failure they might be alerting people to).
Responders can customize their notification methods (push, SMS, phone, and email) and rules, so you can do things like get a lightweight push notification when an alert happens, and then a phone call 2 minutes later if you haven't acknowledged the incident. Teams get escalation timeouts that forward alerts up the chain if the primary hasn't responded after a period of time.
-- Marcin, former Google SRE