Zero Downtime with HAProxy
medium.com
medium.com
As an additional thing I do rolling deploys where I just create an entirely new VM and add it before removing the old one. This just means that I don't have to recycle a node that might potentially have some state on it.
See http://engineeringblog.yelp.com/2015/04/true-zero-downtime-h... or http://inside.unbounce.com/product-dev/haproxy-reloads/ for solutions that are even better than this one.
The challenge is that whether HAProxy's own support for graceful restarts works depends on which OS and kernel version you're running and what network socket features are supported, so ensuring that's all kosher is a bit of a challenge. I took the approach that it's simpler to do something that takes a bit longer without having to worry about all that.
To have 2 servers with the same IP, we use VRRP (Virtual Router Redundancy Protocol) and keepalived. HAProxy is setup on 2 separate servers/instances, and using VRRP/keepalive they both share an IP address (which HAProxy binds to). The servers also have their own unique IP address(es) (on top of the shared one), so the shared address doesn't really "belong" to any machine.
If one server goes down, VRRP gives the IP to the other server and that HAProxy takes over.
Workflow is to just mark the node down for maintenance via socket cat, upgrade it, then mark it back on-line.
# mark node off-line
echo "disable server episode46/web1" | socat stdio /var/lib/haproxy/stats
# update app code
git pull tagged release version
# mark node on-line
echo "enable server episode46/web1" | socat stdio /var/lib/haproxy/statsThe original blog post wraps a restart command in two iptables invocations (and relies on a hacky sleep interval which may or may not work sometimes). The SYN delay method wraps a restart command in two tc invocations. The concepts are more or less identical in complexity as one is telling the kernel "drop SYNs now please" and the other is saying "delay SYNs for a bit please".
All the complexity in the qdisc solution is in the one time setup of the queuing disciplines. I think the largest drawback of the delaying SYN solution is not complexity but that getting it to work with external load balancers is more tricky than getting it to work with internal load balancers. Honestly, you're right that if an org doesn't have to restart HAProxy a ton, then it doesn't make a lot of sense to invest in solving this problem; although if it were me I'd just make sure I was on the latest Linux kernel so that the period during which HAProxy can cause RSTs is as small as possible and not bother with either the iptables or tc solutions.
HAProxy only has this issue on Linux because Linux's SO_REUSEPORT implementation unfortunately introduces a race condition between accept and close. While I haven't personally tested it, HAProxy on one of the BSDs should not have this small window of downtime.
At a high level, you often see programmers sprinkle sleeps into their code to "fix" race conditions or deadlocks. That doesn't really fix the problem, it just moves it around and it's usually done because they don't know how to reason about the underlying problem and fix it properly.
You need to sleep long enough that whatever you're waiting for will definitely have finished. Most of the time you have no exact guarantee of that, so you have to pick some N that is relatively large. Inevitably, no matter what N you pick, sooner or later, the thing you're waiting for will take N + 1 and things break. To make it worse, the N + 1 situation often happens because you're getting an unusually large amount of traffic or because something else in the system is already in a failure state. So the breakage tends to come at the worst possible time and exacerbate things.
Meanwhile, if you sleep for N ms somewhere, one thing you can guarantee is that whatever you're doing will take at least N ms. There's no way to make it faster, even if it may have been unnecessary to wait that long. Often not a big deal, but the more developers sprinkle sleeps into their code, the more often you run into bizarre performance bottlenecks as a result.
Network timeouts and similar are kind of a fact of life, so there's no perfect solution. But if you find yourself trying to solve a problem by sleeping for some arbitrary period of time, a little alarm should go off in your head telling you that there's probably a better solution.
Adding jitter to avoid dogpiling is another case where sleeping is perfectly reasonable.
What I get wary about is the common pattern of: Make a call to some external service. Sleep for some amount of time (to "let it finish"). Then continue under the assumption that it has completed.