1. After a SIGTERM, the shutdownHook should keep the HTTP server running. Future /@status requests must return an error, but user requests must still succeed!
2. The shutdownHook sleeps for a minimum of the load balancer health check’s checkIntervalSec * unhealthyThreshold + timeoutSec (which by the way must be less than the instanceGroupManager’s health check’s checkIntervalSec * unhealthyThreshold if it uses the same endpoint)
3. Now the load balancer should not be sending new requests. The shutdownHook then waits for any existing requests to drain.
4. After requests drain, the shutdownHook can finally exit gracefully.
It is annoying to have to wait for the health check’s delay (rather than simply draining existing requests as in AWS), but it seems to be necessary for Google-designed load balancers and instance group managers.