You Can’t Have a Rollback Button
blog.skyliner.io
blog.skyliner.io
Sure, rolling back isn't trivial. If the code has side effects (e.g. database, disk or cache state), you need to account for those side effects when rolling back.
The example given in the article could be easily fixed by adding a cache.clear() to the rollback procedure (assuming cache performance isn't considered panic-critical)
For database state, use a migration system and make sure each change you make can be reversed (and make sure that's tested!)
> If developers incorrectly believe that their mistakes can be quickly reversed, they will tend to take more foolish risks. It might be hard to talk them out of it.
I doubt it. Either encourage rapid iteration and deployment, or encourage a more stable, well-tested production environment. Do that via feedback, code reviews, post-mortems, internal docs. The presence or absence of a rollback button is not going to a major contributor to this.
Isn't this debunked clearly in the article? If you add a new feature / column and users begin using it (adding data, making transactions) rolling back is a catastrophic data loss. Sure, the code and database schema will go back to a consistent state, but the data never can.
That's what the is pointing out - that deployed applications + data are a state machine without cycles - i.e., even if you 'rollback', it's akin to adding a revert commit on top of your existing state. The arrow of state, as with time, only moves forward.
Adding a new column to a database table (for example) is a single distinct deployment. You do that one without touching your application code. After the column has been added, you deploy the code that uses it. If the new column is a replacement for an existing column, your new code should continue to write to the existing column as well in the same way it always has. If you have to roll back, you roll back the application code change, but the database change remains, so there's no problem. Meanwhile, the new code that was rolled back was still writing to the existing column, so there's no data missing that the old code expects to be there.
After you are comfortable with the change, you can update your code to stop using the obsolete column, test that change, and then later drop the column from the database table.
If you've added a new column or database table for a new type of data, and have to roll back the application code that uses it, you don't lose any data. The customer is of course unable to use the new feature and new data post-rollback, but that's okay: you can roll out the new feature again when you've figured out what the problem is.
Yes, I'm sure we can all come up with examples where this approach won't work, but they are not the norm. The are rare. If they are not, then you are doing infrastructure wrong.
If everything is tested, the rollback button is useless!
Every CM at every large cloud vendor/distributed infrastructure shop has a field asking the operator to detail the rollback procedure which is usually a set of steps that include both turning off new code and fixing up state. Moreover, virtually all large scale distributed systems need to be able to run in mixed mode with old versions and new versions co-existing in different parts of the fleet for hours or days at a time. It's simply not possible to flip from one version to the next in any way that even loosely approximates atomicity at the fleet scale of Google, Amazon, Facebook, etc. Any competent distributed systems engineer will think through the upgrade implications... "How does Vnext RPC to/from Vcurrent?", "How will Vcurrent handle changes to central database state made by Vnext?", "How will Vcurrent handle local on-disk state if I need to flip back to it from Vnext on some set of nodes?", and so on. Often such upgrades entail multiple deployments where backward compatibility code is pushed ahead of RPC/state changing code.
Does all that mean there aren't occasional mistakes rendering a planned rollback impossible and the only path forward is to soldier through to V(original + 2)? Of course not. But the author's idea that you should generally plan to soldier through to newer code if things go wrong during deployment as a standard operating procedure betrays a lack of experience and sends a dangerous message. When deployments go off the rails is exactly the time when follow-on mistakes are made - at that point you've got a tense conference call going, the managers/PMs/VPs are asking for ETA for the restored system and the only acceptable reply is "very soon", the engineers are trying to diagnose quickly and may overlook some symptoms of the botched deployment, etc. Even the best engineer is not going to do a great job thinking through all the implications of the remediation they are attempting at that point and the code they're now about to push has almost certainly cut a bunch of corners in QA. That's not a situation I willingly put myself in.
The author says "reverting smaller diffs as a roll-forward is more verifiable" near the end of the article. I agree the title makes this a bit confusing, but I don't think he's arguing that the only way to recover is to write a patch under pressure.
By partial revert, I'm imagining that three people have changesets (A, B, C, in that order) that have been deployed. You notice that A broke and you make A' to revert it. I think the author is arguing that it is easier to review A' to see if it is a safe change than it is to verify that A', B', and C' (the full revert) are safe to revert.
In other words, even if you don't use version control to record that you reverted A, B, and C, you still effectively do that by reverting in full. You just know that the combination of A', B', and C' was safe when it was deployed.
Is that what you're imagining or are we talking about different things? (I don't have strong opinions about this, I just want to make sure I understand your perspective (: )
In addition to the strategies mentioned, having backward-and-forward migrations for all database changes is essential. You need to have a plausible path for how to restore state to where it was before the broken change.
If you've done the necessary engineering work, this can ultimately be packaged into a "rollback" button which does things like ramping down dark deploys.
> Skyliner is an AWS platform for continuous delivery. We’re trying to build a straight jacket that you can wear to stop hitting yourself in the face.
I'm usually one of those guys with an instinctive distaste for marketing, but it seems like this is one startup that badly needs some professional help in the way they talk to potential customers...
a *straight jacket* that you can wear
You shouldn't be allowed to sell a metaphor if you can't even spell it properly.In my experience, assuming you're doing a decent job of separating your application code and application state, and are avoiding backward-incompatible changes to state storage, rollbacks are almost always safe, and almost always the right first step when a deployment causes customer-visible errors.
I don't find their trivial example of rollback failure compelling. Sure, you can always find an instance when a particular tool does more harm than good, or just doesn't do good. A single example, or even many examples of such, need not form a basis for discrediting that tool.
Rollbacks are a tool in your operations toolbox. Sometimes you should reach for them, and sometimes you should reach for another tool. Claiming you "can't have a rollback button" is counterproductive and needlessly discards a tool that can help your customers when your process fails and you put them in a tight spot.
In my experience, it's relatively easy and clean to work with `if(feature) this else that` in a simple backend application (a service that pulls from a queue and maintain a state, for example).
The problem arise when you start working with more complex applications. Specifically application with complex UI where you have to check the state of the flag at multiple different places in the code. You end up adding branches all over the code, branches that will be removed once the feature is fully deployed and known to be correct.
It adds a lot of overhead, the code is less readable.
There must be a way to do implement feature flags in a smoother way.
- Only a handful feature flags at a time.
- When the feature is stable, delete from everywhere!
- Easily turn on and off.
- No branching makes it much easier.
How do you handle that? You need a consistent permissions framework.
Or maybe they should be and out implementation of feature flags should always take the current customer into consideration...
It's easy to take out,too
We are building similar capability for Pragma (https://pragma.build), if anyone else is interested in having a similar workflow.
In the process of rolling out a CQRS-based system now though so maybe we'll give it a go when time allows.
A good pattern ive seen for this is having 2 stages to your test environment, a single host and then the rest of the fleet.
First you deploy to the single host. Then you run "compatibility" tests that essentially write to the new code, and read from the old code and vice versa.
Thats the ideal way, but also a lot of work. A simpler way it just to roll out changes very slowly, and hopefully detect a deployment with compatibility issues before it effects too many customers.
Rollback is often a suitable first response because the rollback is not a button, it is a procedure, and it can be practiced in a test environment and on a small subset of production without significant customer impact.
Like the other commenter suggested, gradual rollout is a great way to spot bugs or customer feedback. You can accomplish this by building deployment tools that allow for targeted deployment, and by maintaining services that can be targeted by release. For example, you can use your Configuration Management engine to snapshot a DB, deploy schema changes to it and then your code to the app servers using only that DB. If there's a problem you can roll back the code, the data, the schema, whatever changed at that release version, and reload all the services. For things that affect customer assets you'll need extra procedures to work around potential failures there, but that's not complicated (usually).
Edit: note that you also have to take into account how long it takes to notice the problem. If you're rolling out and notice something bad when you hit 1% or 10% of users, it's one thing, but if you hit 100% rollout and notice bad things happening a week later it's a different situation. For one thing, hopefully this indicates the problem is not spectacularly fatal, and for another you've already got users used to the new behaviors and who knows how much accumulated damage. So then it's worthwhile to step back a bit and figure out how to recover.
For example, if you add a column to a database table, then both the old code should be able to work seamlessly both with and without this column. In practice, this would mean not making the column a required one, unless you can invent some kind of neutral default value.
Additionally, it would have been nice to see some mention of patterns that solve this issue more completely, like CQRS, where state is disposable.
And since it's often hard to generalize, I work today with Elixir and Postgres. Anything specific around this stack would be exceptional.
It's important to remember, appups and relups are themselves code that needs separate testing. Ericsson engineers spend as much time (if not more!) testing the relups as they do writing new code.
but anyone already upgraded would be left to deal with the bugs till they are fixed and the there's a further upgrade scheduled.
Note that a rollback could be useful even in the example in this post. Machines that have hit the bad codepath are corrupted, but machines that haven't hit it yet still have good data. Sometimes you need to stop the bleeding even if there's more work to be done for a fix.
> If developers incorrectly believe that their mistakes can be quickly reversed, they will tend to take more foolish risks.
That may be true, but by the same token, if developers believe their code can't be rolled back, they're writing code that's broken in a different way.
My position would be that you should push your own code even/especially on large websites.
We've internalized writing code that deploys well both going forwards and going back. 99% of all changes don't even need to worry about that, because they are UI/rules/bug changes, not base data schema.
We actually don't roll back database schemata; schemata are applied before code that relies on the changes can be written, and thus must be backwards compatible! If you think you can't do it that way, think harder.
It's not cheap to build all the necessary infrastructure and keep training all new engineers in the necessary skills, but it can absolutely be done and pay off handsomely. We never run stale code, and we never diverge far from master.