I've built that once and it was hell to get right. But probably 90+% of programs I've ever worked on do not achieve any business advantage from graceful restarts (as opposed to drain and restart i.e. the stateless model).
I've built that once and it was hell to get right. But probably 90+% of programs I've ever worked on do not achieve any business advantage from graceful restarts (as opposed to drain and restart i.e. the stateless model).
Kubernetes is great for easily sharing and executing recipes, but it can make it tough for people that want to become chefs that are able to use reason and intuition to build and understand complex systems.
Acknowledging that I'm changing the goalposts a bit, I would consider a lack of understanding about the lifecycle of an application that causes development delays an abstract form of "noticing when it fails"
Kubernetes does a really good job at hiding this.
When it hides this, it also hides a couple other things with software — like failing containers, or apps that require churn or fail. I mean this is literally the reason why things like monit were built in the older days — things broke if you didn’t restart them when they broke.
In kubernetes, it also is easy to accidentally to change the system and make that behavior change (at a global level), that results in stuff like your app not being restarted as often, which results in elevated latency, and subsequent failures.
In a way, you're better off with a 99% reliable dependency than 99.9999%. If you expect it to be down some of the time, you'll plan for it in your architecture and (in theory) gracefully handle the failure whether it's due to an insufficiency warmed container, a network issue, a bad deploy, or whatever else comes along.
It's not great for protocols with significant handshaking and potential for long connections. Things like chat services, shoutcast streaming, calling, imap, really anything realtime or notification driven would do better without forcing clients to reconnect. Especially since many will reconnect multiple times, as they may reconnect to servers which are running old code while the drain and restarts happen. If you need near instant cutover (which is pronably rare), in a drain and restart scenario, you basically need double the capacity, where hotload doesn't need anything extra.
Changing the instance means having to re-transfer all of the data and re-establish all of that state, on all of the nodes. You could easily see draining of connections take 15-20 minutes, and booting back and scaling up to be taking 15-20 minutes as well, if you can do it for _all_ the instances at once (which may not be a guarantee, and you could need to stagger things to be more cost-effective).
You start with each deploy taking easily over an hour. If you deploy 2-3 times a day and that your peak times line up with these, you can more than double your operating cost just to deploy, and that can take more than 4 figures to count.
Some of the systems we maintained (not those we necessarily live deployed to, but still required rolling restarts) required over 5,000 instances and could not just be doubled in size without running into limits in specific regions for given instance types.
If a blue/green deploy takes a couple minutes, you're probably not having a workload where this is worth thinking about that much.
Did you look at any of the node-based options like checkpointing the state to a file on the node, and loading that into your newly started pods? Or using read-many persistent volumes? (Not sure if you needed to write to the state file from every process too?)
(This doesn’t help with connections of course, that’s a bit more thorny.)
For some types of my nodes, the majority of the state was tcp connections and associated processes. I don't think there's a generally available system that is capable of transferring tcp connection state (although, I'd love to build one! if you've got a need, funding, and a flexible time table), which would be a prerequisite to moving the process state. All of those connections need to be ended, and clients reconnect to another server (where they might need to do it again, if they don't get lucky and get a new server to begin with).
The other nodes with more traditional state had up to half a terrabyte of state in memory, and potentially more on disk, and a good deal of writes. That's seven minutes to transfer state on 10G ethernet, assuming you can use the whole bandwidth and sending and receiving that data is faster than the network.
Although, in my experience, we didn't tend to explicitly replicate disk based storage for new nodes, all of our disk based data was transient, so replacing nodes meant writing to new nodes, read from new and old, and retiring the old nodes when their data had all been fetched and deleted, or the data retention cap was missed.
I/O Volume meant networked filesystems would be a big stretch. You could probably do something with dual-ported SAS drives, and redundant pairs of machines on the same rack, but then that pair with both go down when that rack has an unforseen problem, plus good luck getting dual-ported SAS drives hooked up properly when you're in someone else's bare metal managed hosting.
(Yeah OK, maybe we had big performance requirements, but hotloading works just as well for stuff that fits on a single redundant pair, or even a single server in a pinch)
The best policy is for a deployment to lower the health check on the service, wait a few minutes per box, then bounce the process.
This is assuming the service is 100% stateless.
Complexity costs money.
I'd advance the theory that picking an off-the-shelf popular solution is going to be beneficial in that case because you externalize the costs of maintaining your expertise and knowledge to the rest of the ecosystem or industry, to other companies, and just never develop that fiber within your organization. It is, however, worthwhile to develop it for more things than just your tech stack. Everything having to do with on-boarding and dealing with legacy code is improved, along with broader dynamics if you try to be a learning organization that develops expertise in its people.
2. You always design your system in accordance to its deploy system, regardless of whether you realize it or not. If you ship binary artifacts and signed packages to customers who install them on their own devices, you will have a different development practice than if you do CI/CD with a single pipeline that always goes forwards. This will also be somewhat different if you work with open-source components that require paying attention to version schemes rather than just pushing a hash in a container image. If you use feature flags to help merging but also to control deployment, adopt A/B testing, and all these practices, you're intimately adapting your development approaches to deployment mechanisms that are available to you. It's not an extra cost, it's a cost you already pay today.