Now 4 out of 10 services are marked as "Incident", yet most of the others are also completely dead.
Now 4 out of 10 services are marked as "Incident", yet most of the others are also completely dead.
This defeats the purpose of a status dashboard and is effectively useless in practice most of the time from a consumers point of view.
If your reliability metrics have lots of false positives, that's on you and you'll have to write down some reason why those false positives exist every time.
Then that company could decide for itself whether to update manually with "not a reliability issue because X".
This lets consumers avoid being gaslighted and businesses don't technically have to call it downtime.
Declare an incident first, investigate later.
Cheating SLAs by delaying the incident is a good way to erode trust within and without.
If that would be the best way to deal with it- why is literally no one doing it this way and what does that tell you?