Not saying it's right, just an observation
Until I decided to no longer do that and my life improved considerably.
I've seen status pages and error counts tied to bonuses, which only caused a giant mess of bad incentive alignment and internal lies, customers are unhappy, developers are unhappy, management are lying to upper management, it's so much easier to focus efforts on real problems and just be honest and improve. Thank goodness I dont work there anymore (cough cough Google)
This could have been caught with a cron job and some curl requests :\
Building infra I have to work around all sorts of 3rd party services going out or having blips throughout the day, docker registries, caches, bgp, etc., it's totally an expected part of infra design but not every team has the time or need to build in the resiliency. I see tons of outages that never get reported or IMO aren't reported adequately enough.
With that said, I'm no angel, I get all my service down notifications through slack, so when slacks down..
I've caught a big cloud provider not reporting a degraded service, I assume they knew but politics and $ come in and it's easier to just gaslight everyone. I get it, but my frustration is worth loosing a trailing 9.
I think there should be some 3rd party continuously testing APIs. Degraded states are downtime!
For example status.digitalocean.com is _not_ real time, it's manually updated.
And it's irritating as fuck.
Foobar: for when you are too polite to say FUBAR (Fucked Up Beyond All Recognition).
Named, of course, by the Army when they built the Alaska Highway.