The former is basically a "return true" endpoint, which can tell you if the service is alive and reachable. The latter will usually do something like "select 1;" from any attached databases and only succeed if everything is OK.
The former is basically a "return true" endpoint, which can tell you if the service is alive and reachable. The latter will usually do something like "select 1;" from any attached databases and only succeed if everything is OK.
Then if all 100/100 nodes get taken down by some shared problem, the system simply degenerates into picking the idle-est of the 100.
This allows you to have less frequent checks that are more in-depth, and basic ones that are just the `return true` type.
I guess you're correct that you can have a widespread db outage that makes many of them fail, but then there should be new ones coming very fast, as you can set the deployment minimum for services taking requests. I think you can get very close to stability even in this circumstance.
https://kubernetes.io/docs/tasks/configure-pod-container/con...