> Interested to hear examples of cases where extremely long-lived connections are worth a lot of engineering pain to achieve.
When I was at WhatsApp, I regularly observed connections from mobile phones (S60 mostly had the longest connection times, but all clients without a working platform push service will try to remain connected) in excess of 30 days. We weren't going against the framework to keep those though, we were all in on hot loading, and didn't do any container rigamarole until the move to Facebook. Keeping alive connections alive reduces work for client and server, and some networks are terrible and getting a connection takes many attempts, but often a working connection will continue to work; of course, some networks are terrible and timeout connections after 10 seconds of idle, but usually that's temporary.
If your 'long' connections are usually around a minute, like is common in http use cases, it's not too bad to keep old servers around while they drain, but if it's an hour or more, you'd need 2x the instances while you're rolling out, if you want to rollout in less than an hour... if you rollout quickly, and want to rollout a follow on, you're going to see even more. Depending on how you manage things, maybe the old and new instances can share machines, resource usage on the old should drop off quickly, but you might not be able to actually fit many instances on one machine, or resource allocation might not allow for it.
We would still have to kill connections in use sometimes; hot loading the BEAM could be possible, but it's not built for that and I haven't met anyone crazy enough to do it. Hot loading the FreeBSD kernel hasn't been done yet either (although, I have seen it for Linux). If there was a real need to keep TCP state while updating BEAM or the OS kernel, it would probably make more sense to build a way to transfer the states (tcp and corresponding application) among machines rather than to hot load all the things; it would probably be more approachable to develop that than to add hotloading to things that didn't plan for it. I don't know how WhatsApp manages today, I left sometime ago, and I avoided getting involved with the FB hosted servers; I know we had the ability to do hot loading, but it didn't really fit in the FB deployment model, and some people didn't like having multiple deployment methods.