Grace: Graceful restart for Go servers
github.com
github.com
For all other cases[1], I think graceful restarts are a bad practice. When you use that pattern, you come to rely on your server closing down cleanly. It makes you consider abrupt terminations (pulling a powerplug, kill -9, network partitions) an edge case.
Because you separate 'shutdown cleanly' from 'shutdown abruptly' and that in most deployments clean shutdowns are much more frequent than abrupt ones, you end up not exercising the abrupt termination case very often. This means that your system's reaction to abrupt terminations is likely buggy or neglected. If it works fine because you took great care of handling that case, the effect is that your codebase is now more complex. You have two code paths to handle termination: a clean path and a dirty path.
If instead, you shutdown only one way (abruptly), those two cases become the same. You reduced the problem space and now your code shuts down only 1 way. You can be confident that your system is resilient to abrupt failures, since you exercise that shutdown path every single time you stop your server. Some call this crash-only software.
A caveat with this is that clients must be able to handle abrupt terminations. This usually means that clients must have the ability to retry (with exp-backoff) their operations, and that each operation they perform has a timeout.
How To Know If You're Infected By The Graceful Shutdown Fever:
Assume that your servers are `kill -9` every 2s at
random, are you confident that your system will keep
working properly?
[1]: another valid case I found was in distributed systems that require a certain part of their members to be up at any time. If the nodes coordinate their restarts, they can ensure that a minimum number of nodes are kept healthy.
[2]: I've spent many hours arguing this with very smart colleagues who disagree, so this is a bit of an opinionated view (still I think I'm right =])However there are times you want a graceful restart because an abrupt restart affects your client's UX - such as during a file upload, or where a client may be running a long running synchronous operation (like a credit card charge). In both cases the system may be resilient against a -9 and retries, but annoying if "the world stops" because of a change to prod.
If the server dies because the dev team decides to push to prod 4x in an hour and the user has to retry every time, a graceful restart lets you do that without pissing off anyone who was using that live connection.
I think graceful restarts are fine for servers that offer
web pages or other things where a client will not retry a
request.
+ when retrying a request is not practical.* https://github.com/zimbatm/socketmaster - very simple
* https://github.com/pusher/crank - adds a control socket, dynamic config and more fine-grained controlled restarts
Both are designed to sit on top of your process manager of choice. Unlike self-forking process solutions they don't change PID so it works great with systemd and upstart.
There's a few that are drop-in replacements for http.ListenAndServe, which is great, until you want to do something custom, then it gets a lot more kludgy. Grace is not kludgy.
And, most importantly, it definitely works. Was able to hammer the bejeezus out of my daemon, restarting it willy nilly, without issue or fanfare.
Couldn't immediately tell from reading the source - but is this to say that it uses the sd-* APIs directly, that it implements behavior compatible with sd_listen_fds, or some other form of "socket activation" like superserver or fd-holding?
a := newApp(servers)
https://github.com/facebookgo/grace/blob/master/gracehttp/ht...
EDIT: read my comment below where I elaborate
If you are going to be creating these complex structures why not using a factory function for it? If there were many creations of this type it could severely increase the bulk of the code....
Are you saying you dislike single-character variable names, and that you see them frequently in Go?
In the context of reading someone else's source code, the variable name "a" is inadequate, and I've seen it in many Golang codebases. It's like intentionally obtuse code.
Maybe I haven't read enough Go code in the wild, but it really feels like the C days. Yes, it's GC'd, and that makes a difference. Yet I can't help but wonder how much easier it is to read code from the Python or Ruby communities. I'd even go as far to say that it's easier to read code from the Javascript community.
Your "complaint" of go actually have nothing to do with Go itself...
In my opinion, a is an great variable name. Especially when it is used after calling a function like newApp(...). I may have picked a different name for newApp but it was still clear to me what it did.
Personally I think single character variable names are over used in Go as well, but at least they're generally restricted to the smallest of functions (eg struct methods), or instances where a longer name wouldn't be more descriptive (eg inside for loops (for i := 0), or b[], byte arrays.).
However in this specific example, I think "a := NewApp(server)" is both readable and overly terse. So I do find myself agreeing with the OP in their example.
I have two metrics for variable names. First, is it clear what the variable is? You say yes. Second, is it not obnoxious to type? Clearly not.
So, at least by my metrics, it's a great variable name.
I've looked at a lot of other people's Go code and I have had no issue with 1 character variable names, but maybe that is just me. The methods/functions is generally small enough that it is obvious to me what is happening.
It allows for graceful restarts after replacing your binary. I'm not sure if that is a feature of grace?
Perhaps there's a kernel-level API that could be added to allow sockets to be snatched or handed over to a new process. That is, honestly, probably the more apropos solution. The fact that sockets act as a kind of lock is an implementation detail.
* SO_REUSEPORT - not reusing the socket, but allowing multiple sockets to bind to a single port. [1]
* Fork / Exec - have a parent process signal the child process to terminate after it has spun up a new version and the new version is accepting connections [2]
[1]: https://lwn.net/Articles/542629/
[2]: https://news.ycombinator.com/item?id=9694339