I'm not sure it is super critical in the age of containerized workloads with rolling deploys but at the very least the connection draining is a good pattern to implement to prevent deploy/scaling related error spikes.
I'm not sure it is super critical in the age of containerized workloads with rolling deploys but at the very least the connection draining is a good pattern to implement to prevent deploy/scaling related error spikes.
Not sure how the cloud providers do it though, maybe combination of low DNS TTL and rolling restart since they often have huge fleets of servers which handle ingress?
With ingress you'd have a load balancer in front or have it routed in the network layer using BGP.
But I agree it isn't as easy as a in place upgrade.
Multihomed IP <-> Loadbalancer <-> Application
By having the same setup running on multiple locations you can replace the load banacers by taking one location offline (stop announcing the corresponding route). Application instances can be replaced by taking the application instance out of the load balancer.
At least some I am familiar with operate at the packet level and can hand off live "connections" to a peer or hot standby, along with full session state.
Remember that with TCP or anything else, the abstract session is an illusion.