* State in old system isn’t representable in new system (not a backwards compatible upgrade or more likely a bug exists in handling the new state)
* There’s state outside of the program that’s impossible to transition gracefully (e.g. dirtied IO socket where you don’t know what it’s state is & it’s a resource owned outside of your program)
* Transitioning state means there’s a possibility of failure because the program never reaches a graceful transition point to snapshot the state. So you either have to choose between running the old program forever or abandoning the graceful state transition anyway.
Distributed systems I’ve observed pick one of two strategies:
1. Using the load balancer strategy of migrating off the old version & then terminating it after some grace period.
2. Use a formal distributed state system like CockroachDB, Yuggabyte, DynamoDB, S3, etc etc.
This is probably a big reason why most programs use external storage solutions even if they’re less efficient - it centralizes maintenance of state onto a system that has well defined semantics and can handle repair transparently.