My Philosophy on Alerting (2014)
docs.google.com
docs.google.com
Today I'd update my thoughts in a few directions.
Operations is far more about risk management, mitigation, preemption, and response than other elements of tech work. Monitoring is part of that, but so is drilling, analysis, and proceduralising response. Selling ops as a risk-management strategy might also address the comprehension failures management often reveals in response to concerns raised, or cost-based objections. See: https://joindiaspora.com/posts/647a6300a2220139344a002590d8e...
Another piece I've written on problem resolution applies: https://old.reddit.com/r/dredmorbius/comments/2fsr0g/hierarc...
On "cause-based alerts": there are states and there are levels, and sometimes what you want to track are levels. Available storage, filehandles, sockets, or other consumable resources would be well-worth logging and alerting on. Key is that the alerts should trigger some action, which is often not the case.
"Normalization of Deviance" is a term used by Diane Vaughan (I also associate it with Charles Perrow of Normal Accidents), and poses a potential ... risk ... of Ewaschuk's approach: alerts which don't lead to an actionable event might be either ignored or muted, though they do in fact represent risks. Normalising deviance leads to eventual catastrophic failures: the Challenger and Columbia disasters, Three Mile Island and Chernobyl, Champlain Towers South, ...