Now you've got a host that can't access a database or can't do much of anything- but it can return true!
Health check APIs should always do some level of checks on dependencies and complex internal logic (maybe a few specific unit/integration tests?) to ensure things are truly healthy.
10 500s in a row? Go to the timeout chair until you get better.
The former is basically a "return true" endpoint, which can tell you if the service is alive and reachable. The latter will usually do something like "select 1;" from any attached databases and only succeed if everything is OK.
Then if all 100/100 nodes get taken down by some shared problem, the system simply degenerates into picking the idle-est of the 100.
This allows you to have less frequent checks that are more in-depth, and basic ones that are just the `return true` type.
I guess you're correct that you can have a widespread db outage that makes many of them fail, but then there should be new ones coming very fast, as you can set the deployment minimum for services taking requests. I think you can get very close to stability even in this circumstance.
https://kubernetes.io/docs/tasks/configure-pod-container/con...
It can check the database connection, redis, cache, email, up-to-date migrations, and S3 credentials.
There are so many types of load balancing, and so many different ways to guard, simply with best practices, against what this story explains.
I also found the 'quadruple hump' in a graph reference difficult to follow without more of a backstory of what the graph was representing.
I also don't understand why it would be so difficult to grep through the logs of that load balancer (or load balancers) and find that common fault on the backend.