Towards zero-downtime upgrades of stateful systems
stevana.github.io
stevana.github.io
* State in old system isn’t representable in new system (not a backwards compatible upgrade or more likely a bug exists in handling the new state)
* There’s state outside of the program that’s impossible to transition gracefully (e.g. dirtied IO socket where you don’t know what it’s state is & it’s a resource owned outside of your program)
* Transitioning state means there’s a possibility of failure because the program never reaches a graceful transition point to snapshot the state. So you either have to choose between running the old program forever or abandoning the graceful state transition anyway.
Distributed systems I’ve observed pick one of two strategies:
1. Using the load balancer strategy of migrating off the old version & then terminating it after some grace period.
2. Use a formal distributed state system like CockroachDB, Yuggabyte, DynamoDB, S3, etc etc.
This is probably a big reason why most programs use external storage solutions even if they’re less efficient - it centralizes maintenance of state onto a system that has well defined semantics and can handle repair transparently.
Indeed, now contrast with the article on HN discussing orthogonal persistence.
That shouldn't stop us from solving the problem in the cases where it's possible though? We can tackle the corner cases separately with manual overrides.
> This is probably a big reason why most programs use external storage solutions even if they’re less efficient - it centralizes maintenance of state onto a system that has well defined semantics and can handle repair transparently.
This is certainly the case today, what I'm asking is: does it always have to be like that in the future?
As I understand it, Erlang/OTP captures the entire state of the program and it’s a feature of the language and VM to accomplish this. It’s not something you can retrofit into any arbitrary language. For example, your JS app or your Python app or your Rust app won’t be able to do the same easily which means it won’t be robust and it will be error prone. Thus I stand by that there’s no “generic” solution you can bolt onto an arbitrary language.
I believe Erlang supports two versions running along each other. They capped it at two because back when this was developed there wasn't enough RAM. Joe Armstrong gave at least one talk where he says if he'd have liked to support arbitrary number of versions and garbage collect them as old sessions complete.
> Thus I stand by that there’s no “generic” solution you can bolt onto an arbitrary language.
The main point of the post is centered around Barbara Liskov saying "maybe we need languages that are a little bit more complete now". I'm not interested in the limitations of current languages, I'm interested in the future possibilities.
I say you can do hotload in any language that supports dlsym/dlopen or eval. I've done it (rather poorly) in Perl and C, and I'm sure others have done it in other languages.
It's a lot nicer in Erlang, so IMHO, if your use case includes long running processes with expensive to construct or transfer state (such as long running sockets), it's worth considering Erlang or something than can do hot loading.
New system starts in a backwards compatible mode where it accepted all state that was representable in old system. Transition is achieved after the upgrade, with flag variables.
Sometimes the tradeoffs of the application make it worth to spend QA resources to validate a staggered upgrade path.
From the post:
> In a world where software systems are expected to evolve over time, wouldn’t it be neat if programming languages provided some notion of upgrade and could typecheck our code across versions, as opposed to merely typechecking a version of the code in isolation from the next?
There's different ways to handle it, but the 'easy' way is to write your new version that updates the old state to new state on first touch, and make sure either you have a no-op message you can send so the state gets updates or you have some periodic thing that means every state will be updated in X time.
There's also a way to have a 'try_update' that fails without changing anything if there are two versions active already. (Or you can just YOLO and anything with the old old version in the stack gets killed).
I'm not sure if there's better tooling for it now, but there wasn't anything to help you test transitions when I was doing it. For automated tests, you'd need to build a state with the old version, load the new version, run a test, etc. It's the same hole in testing if you do mixed version frontends against a shared database, or mixed version frontend vs backend; it's just more apparent because it's on a single system.
Java OSGi. Great tech, didn't take off, probably to complicated with not enough gain.