I believe this is called an 'alert that cried wolf'.
I believe this is called an 'alert that cried wolf'.
It comes up everywhere:
- On-Call alerts received by my engineering team for our microservices that usually self-resolve result in the first action being taken by the engineers to be "just wait and see if it's a false alarm." We work to reduce the number alerts overall to reduce the noise.
- I'm reminded of "Cigna saves millions by having its doctors reject claims without reading them" https://news.ycombinator.com/item?id=35304017 In order to keep up with the claims, they just auto-rejected them to see if they were appealed in a way to filter out the "noise."
- We hear so much about Tesla's "FSD (Supervised)" and it's request that drivers "don't become complacent" but it happens anyway. After enough time behind the wheel with FSD enabled, we are swayed by the string of successes to become overly trusting of the tech.
> On Tuesday, the National Transportation Safety Board said that a crash last year on the Washington subway system that killed nine people had happened partly because train dispatchers had been ignoring 9,000 alarms per week. Air traffic controllers, nuclear plant operators, nurses in intensive-care units and others do the same.
I've been involved in some systems that routinely spit out dozens of warnings per day. After 'fixing' some of them - maybe getting a system down to a few per week, bug reports started coming in that the system was 'broken'. Because people noticed the boxes and messages they routinely ignored were gone, and this must be a problem. Lots of re-education may need to happen on a large scale to actually 'fix' alarm fatigue across the board.
Several years went by... then because we finally had the time our ops team decided to audit and improve this application, including rewriting significant portions of its code as was pretty typical when we took over an application built by an external vendor. Long story short, we resolved almost all of the errors by the expediency of simply correctly interacting with the database behind the application instead of relying on whatever horrific dumpster fire output to the DB that Hibernate was doing. That's when things got super interesting, because the external vendor team that was receiving these monitoring alerts (unbeknownst to my team) had more than 100% turnover in the intervening period of time, having lost all knowledge about this, but suddenly starting getting a bunch of critical urgency alerts going off because the application was no longer emitting the amount of errors expected for the alarm thresholds. This ended up triggering a massive company-wide incident call, even though there was no external customer impact. I'm leaving a lot of details out, but let's just say it sucked up days of my life and was a massive waste of time for all of that.
Fun example of fixing something and causing an "issue" because behavior changed from expectations/baseline.
I tell them to do two things to tell them how to check system health.
1. Is it working normally? if so, then nothing is wrong.
2. Walk by the hardware once a week and check for idiot lights, check the system dashboard if one of the idiot lights is on. If no error is found, its a hardware problem (failed supply, drive, fan), if an error is found it's a failed server - either way, call us.
There are a handful of cases not fixed by this, but they're limited to time sync failures - which are critical, but generates no user facing alarms.
The last time I was in hospital, I came out the next day and my ears were ringing for a couple of days from the _constant_ ringing of alarms.
My wife was in the hospital because she’d essentially stressed herself to the point of a heart arrhythmia. The irregular rhythm meant the equipment was basically useless at accurately determining her heart rate and the readings fluctuated wildly.
So what did she do for several hours? Laid in a bed anxious about the medical concerns, anxious about life, and listened to the monitoring equipment go off multiple times a minute as she crossed various alert thresholds, went back under them, crossed them again, repeat the entire time she was there.
I asked the staff to turn the alarms off or at least adjust the thresholds.
They just short of flat out said they wouldn’t because if they disabled or adjusted them and anything happened it would be on their head.
So my wife spent half a day being told to try and relax while the machine monitoring her heart and respiration and blood pressure constantly screamed at her that everything was wrong.
The better solution is to redesign / engineer systems which automatically solve the alarm condition so no person needs to know about the alarm. That usually requires many people to work together and requires far more complexity to solve, both initially and in ongoing upkeep.
If someone dies because they ignored an alarm you did alarm, that's on them.
When something unexpected or out of the ordinary happens, what do you do about it?
- Nothing? Why did you do nothing? This was clearly a problem. This is your fault. - Figure out appropriate error handling? That’s hard work. - Just throw up an alert on everything and make it someone else’s job to sort through it all? Perfect, job done.
"But the A/B tests show consumers prefer it!". Oh? So you have a degree in stats and experience running scientific tests? No? You're just a programmer who only knows how to enable the A/B test feature of your favorite javascript library?
Your A/B tests aren't testing what you think they are testing.
it's extremely hard to find the point where it becomes negative.
Think of it like emails, or texts, or calls: What's the exact number per day where you stop paying attention to them?
[1] https://99percentinvisible.org/episode/sound-and-health-hosp...