The load-balanced capture effect
rachelbythebay.com
rachelbythebay.com
https://www.usenix.org/conference/lisa14/conference-program/...
Specifically the section about 20mins in where he talks about KS windowing.
For those not already familiar with K-S tests, I'll save you from a google query: http://en.wikipedia.org/wiki/Kolmogorov%E2%80%93Smirnov_test
http://docs.aws.amazon.com/ElasticLoadBalancing/latest/Devel...
Health checks also don't work if your health check ping path is healthy but some other part of your app is broken.
If the load balancer is monitoring this, it's quite capable of knowing what the average number of 500's returned is (knowing it automatically, via monitoring over time), and noticing that one machine is returning 500's at 10x the rate every other machine is, and then reducing it in the rotation.
No? It would not surprise me if this would lead to yet _other_ undesired side effects in certain conditions though -- I'd expect the same of the solution OP proposes though, or anything automated like this.
One of the issues is making your active health check more like a doctor's physical than "'tis but a scratch" self-reporting. But also ensuring you're not dealing with a whole bunch of hypochondriacs.
Passive health checks at least have the property that they fail servers when the servers are unable to serve, even if the active health check does not consider some subsystem in its response. But alone they can easily be fooled by really fast non-error responses.
Anyway, saying "name of brand of load balancers" solves this problem is only covering the most basic cases. General solutions are at best only the first step of the full solution. You need to think about the edges - which I suspect is what Rachel is advocating.
You get a set of basic VM level metrics, and you can feed it custom metrics from your app, or log files. All of which can be configured to alarm. I don't think it's possible to run advanced statistics on the metrics for alarming (eg, standard deviation from 30 minute exceeds N), but it may be. Usually it's just an event count, like more than N 500 errors over X time.
I do agree you need to think deeper than basic health checks though, 'broken server' is always a hard boolean to nail down.
But as others also pointed out, you can still get some "captured" servers until the LB realizes the machine is not super fast but simply superbad.
Also your health check URL might return 200 and the rest of your app 404 / 500, which makes me think
Perhaps the LB should be aware that if it gets back way more 404 and 500 than average for known URLs, then this should be considered as a bad server... I assume advanced LBs support that?
You'd typically learn this after probably six months of running a large-scale continuously-deployed dynamic website and it breaks from poorly-tested configuration changes + hardware issues. Sysadmins know this stuff. That's why there is a job title of sysadmin and not "developer who does sysadmin stuff sometimes".
No load balancer I know of will remove a web server that returns a single 5xx from its healthy pool. It will need to use some heuristic as Rachel points out, some percentage that is based on the statistical norm. Otherwise it'll fail out too many hosts and cause a problem.
I think you're severely underselling developers. I've met people who have only had the Software Engineer title and never installed a Linux distribution who get this stuff at least as well as anyone who calls themself a systems administrator. Sysadmins don't have a monopoly on understanding failure cases and failure handling - false positives, false negatives, outliers and outlier detection, metrics to look at. It's a skill that comes from experience, which can happen (or not) whatever you call yourself or others call you.
I'm lucky enough that I get to focus on this type of problem, and while there's definitely an aptitude portion, it is also a teachable skill - I see my job partly as getting the team I'm working with to be able to do this stuff when I'm not around. That usually means finding two or more people in the team and cultivating their interest in it.
For "random" requests, if you have a 500 response, requests of the same "type" should not longer be sent to that host. This can be changed based on scoreboard settings. Depending on the context, you may choose to serve cached content on 500s. This is one of the reasons multiple layers of cache and application intelligence is so handy.
I'm not underselling anything. Domain-specific knowledge comes with experience. If you ask a mechanical engineer 'What's wrong with my car if it makes the noise "bang-sputz-sputz-screech-screech-screech?"', the engineer will start making you lists of what parts can make each of those noises and begin cross-referencing to see maybe in what conditions a combination of those might happen. The mechanic will immediately tell you that for your 1991 Mercury Sable, the A/F mixture is off, the MAF sensor needs cleaning, the radiator has a crack and the accessory belt needs replacing. Sysadmin is a trade, not a skill.
For a non-healthcheck 5xx response, it almost never is clear that this host in itself is responsible for the 5xx response. 5xx is the correct response when there is an error on the server side (ie, not an error on the client side, not a correct response), but it doesn't mean the server is a problem - it just means that the server experienced a problem in serving the request. That failure itself may be from one of many RPCs that server made to other services. As such, all web servers behind the load balancer for that request type will exhibit the 5xx response (at some rate, and depending on any state in connection sharing/reuse between the server and their upstream service), and all would subsequently be removed. Which isn't the correct response at all.
As someone who has had the job title "Systems Administrator" and the job title "Software Engineer", and currently has neither but still does exactly what he's always done - solving problems by understanding systems and, among other things, by writing code - I wouldn't consider load balancing and failure domains/types/handling as the sole or even primary purview of a systems administrator - especially in the case of large installations.
This is not a solved problem, but it's also not a novel problem. It's a problem that I've faced before, and it requires decent monitoring to catch.
I'm surprised the OP doesn't suggest having your load balancer pay attention to returned error codes.
If the load balancer knows what average amount of 500 or non-200 responses is, and one unit is returning way more than an average rate of non-succesful responses, it would make sense for the load balancer to back off sending to that machine. But maybe still send an occasional request there, so it can notice when/if it's error rate returns to normal.
Do any load balancers work like this?
The reverse of this issue can also cause problems where you get a bad node that accepts the request but never returns (or takes 2-3 mins to respond). As a rule the timeouts will cause queuing, thread pool starvation and general not working all the way back up your request chain to whatever is facing the internet where your site will either hang or give back a 503 page.
Just goes to show you that thinking only of the happy path can often lead you astray.
More than a few real-world front-end implementations lack the kind of rigorous instrumentation that's necessary to identify problems of the kinds mentioned in the article. While proper configuration and experienced ops folks tuning said infrastructure can solve most of the stated problems, very fine-grained monitoring is sometimes the only thing that allows for effective troubleshooting when things really go L-shaped.