That's not necessarily true. In kubernetes, for example, you have Liveness and Readiness probes, which can both have a period, timeout, initial delay and importantly a number of failures before you kill the service and spawn a new one.
This allows you to have less frequent checks that are more in-depth, and basic ones that are just the `return true` type.
I guess you're correct that you can have a widespread db outage that makes many of them fail, but then there should be new ones coming very fast, as you can set the deployment minimum for services taking requests. I think you can get very close to stability even in this circumstance.
https://kubernetes.io/docs/tasks/configure-pod-container/con...