Gracefully Restarting a Go Program Without Downtime
gravitational.com
gravitational.com
Sure, unless you're the one running the load balancer or the service discovery system that enables other people to just kill their instances without worrying about graceful restarts.
In the case of this article, it's discussing teleport, a drop in replacement SSH server. For many users, this is their administrative connection to a fleet of servers (teleport can operate as a transparent bastion, as well as on the end host). So in our case, it's a desirable feature to not interrupt existing shell sessions, while replacing the software serving those connections. This could mean the admin's own connection to the server doesn't fail mid-upgrade, or that hundreds of internal employee's don't groan because whatever they were doing that used a shell just got dropped and they need to restart.
Disclaimer: I work with the author of the original article, and contributed to the article.
What I do take issue with is the suggestion that handling such special cases is "table stakes." A little humility goes a long way in bolstering credibility.
Being able to replace vulnerable singletons that require process and TCP socket gymnastics with multiple instances having good fault isolation properties is the real "table stakes" for modern highly-available service architectures.
I worked on a plan to have live migration of multiparty phone calls between data centres in realtime. The tech is available. But your planning outcome may be: it's cheaper to make single node more stable and deal with the occasional failure.
"Stateful" only means you need to care about the state being moved. Things like Java hotswap, Erlang live code reload, and your custom solutions can handle that.
See http://erlang.org/doc/design_principles/release_handling.htm... for the basics. With no shared state, and no global variables, and no side effects, you can update functions so the next time they're called they're something else.
Edit: I also remember hearing some out-there stuff about handing off your IP address and open connection state to another machine so the client doesn't even notice…
"Truly Seamless Reloads with HAProxy – No More Hacks!"
https://www.haproxy.com/blog/truly-seamless-reloads-with-hap...
- Refuse all new connections. If you can change the program, this can be done by closing the listening socket (I think) , but I just installed an iptables rule.
- use the load balancer's API to offline the server(at the time it was F5).
I liked the first solution because it was load balancer agnostic, but required retries by the load balancer or client to avoid disruptions. Curious to hear what other people did.
https://landing.google.com/sre/book/chapters/load-balancing-...
Health checks are part of GCPs load balancer. https://cloud.google.com/compute/docs/load-balancing/health-...
"Commonly, graceful restarts are performed by the active process (dark blue) closing its listeners and passing these matching listening socket files (green) over to a newly started process. This restart causes any foreground process monitoring to incorrectly detect a program crash. overseer attempts to solve this by using a small process to perform this socket file exchange and proxying signals and exit code from the active process."
However this approach need the binary to be started from socketmaster (as child process). I'm not sure what is the real cons tho, except the fact that it alter SIGHUP behavior - which is fine in this case imo.
If you don't have multiple instances of your load balancer, that's weird.
When you remove the host from the routing table, the ECMP algorithm will redistribute the traffic to the other hosts still in the pool. For TCP connections, this will generate reset's, and drop any active connections moved between hosts. Also, at times I've not gotten clear answers from router vendors on whether they support a consistent hash in their ECMP implementation, which means that connections through the load balancers not being reloaded may also be rebalanced to other hosts and be affected by the routing table change.
Even if we consider going down the stack to a layer 4 load balancer, outside the scope of this article, it may still be desirable to combine approaches and offer hitless reloads (a common case) and not just recover from a node failure (a less common case).
The reason is, that you would need state replication of the load balancing tables to support removing the load balancer from the routing tables, and having another load balancer take over it's work while maintaining the connections through the load balancer. This is under the assumption that the service being load balanced requires an uninterrupted connection of course.
Disclaimer: I work with the author of the article and contributed to it.
Lol, no. Modern, professional software development uses a load balancer to manage a canary/red-black deployment strategy. Zero downtime deployments to a live server, outside of telecom/finance, is a sign of an amateurish team.